biondizzle/vllm - vllm - Gitea: Git with a cup of tea

Author	SHA1	Message	Date
Maral	2e9034c998	[W8A8 Block Linear Refactor][2/N] Remove W8A8Fp8BlockLinearOp and adopt Fp8 block linear kernel selections. (#33892 ) Signed-off-by: maral <maralbahari.98@gmail.com> Signed-off-by: Maral <maralbahari.98@gmail.com>	2026-04-09 08:50:39 +08:00
Lucas Wilkinson	70406eb1dc	[Attention][V0 Deprecation] Deprecate accept output buffer (#39125 ) Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>	2026-04-07 17:14:58 -04:00
shunting314	8b141ed8c3	full cudagraph for flex-attn (#36298 ) Signed-off-by: shunting314 <shunting@meta.com>	2026-04-02 21:15:01 -07:00
Carl Y	1f5ec2889c	[mla] Support fused FP8/NVFP4 output quantization in MLA attention (#35792 ) (#36205 ) Signed-off-by: Carl You <4531192+carlyou@users.noreply.github.com> Signed-off-by: Carl Y <4531192+carlyou@users.noreply.github.com> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>	2026-04-02 21:16:11 -04:00
Monishver	c09ad767cd	Feature/silu block quant fusion v1 (#32996 ) Signed-off-by: Monishver Chandrasekaran <monishverchandrasekaran@gmail.com>	2026-04-01 18:50:43 +00:00
Luka Govedič	40bb175027	[vLLM IR] 1/N Implement IR skeleton and rms_norm op (#33825 ) Signed-off-by: Luka Govedič <lgovedic@redhat.com> Signed-off-by: Xinyu Chen <xinyu1.chen@intel.com> Signed-off-by: chzhang <chaojun.zhang@intel.com> Signed-off-by: Luka Govedic <luka.govedic@gmail.com> Co-authored-by: Xinyu Chen <xinyu1.chen@intel.com> Co-authored-by: Chaojun Zhang <chaojun.zhang@intel.com> Co-authored-by: Luka Govedič <ProExpertProg@h100-01.nemg-001.lab.rdu2.dc.redhat.com>	2026-03-31 22:15:05 -04:00
BadrBasowid	077a9a8e37	[torch.compile] Refactor Attention Quant Fusion Pass and Remove Boilerplate (#37373 ) Signed-off-by: BadrBasowid <badr.basowid@gmail.com> Co-authored-by: vllmellm <vllm.ellm@embeddedllm.com>	2026-03-31 14:15:50 -04:00
wliao2	4dfad17ed1	replace cuda_device_count_stateless() to current_platform.device_count() (#37841 ) Signed-off-by: Liao, Wei <wei.liao@intel.com> Signed-off-by: wliao2 <wei.liao@intel.com> Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com> Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com> Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>	2026-03-31 22:32:54 +08:00
Nicolò Lucchesi	cc06b4e86b	[Mamba][Bugfix] Raise on insufficient cache blocks instead of silently capping cudagraph sizes (#38270 ) Signed-off-by: NickLucche <nlucches@redhat.com>	2026-03-30 09:41:50 +00:00
Andreas Karatzas	f2d16207c7	[ROCm][CI] Fix flaky GPTQ compile correctness test (#38161 ) Signed-off-by: Andreas Karatzas <akaratza@amd.com>	2026-03-26 19:57:00 +08:00
Richard Zou	6e37c46b35	[compile] Add some more startup tests for top models (#38046 ) Signed-off-by: Richard Zou <zou3519@gmail.com>	2026-03-25 12:02:22 -04:00
Harry Mellor	d215d1efca	[Mypy] Better fixes for the `mypy` issues in `vllm/config` (#37902 ) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2026-03-25 06:14:43 -07:00
vllmellm	42e9547976	[ROCm][Test] Fix ROCM_AITER_UNIFIED_ATTN attn+quant fusion test (#37640 ) Signed-off-by: vllmellm <vllm.ellm@embeddedllm.com>	2026-03-25 05:06:15 +00:00
Terry Gao	82580b10ac	[Perf] Disable inductor runtime asserts by default for serving perfor… (#37485 ) Signed-off-by: tianrengao <terrygao87@gmail.com> Co-authored-by: Tianren Gao <tianren@fb.com>	2026-03-24 19:37:51 -04:00
Richard Zou	89f572dbc0	[BugFix] fix VLLM_USE_STANDALONE_COMPILE=0 (#38015 ) Signed-off-by: Richard Zou <zou3519@gmail.com>	2026-03-24 19:08:26 +00:00
Wentao Ye	c59a132f96	[V0 Deprecation] Refactor kv cache from list to element (#37487 ) Signed-off-by: yewentao256 <zhyanwentao@126.com>	2026-03-23 20:10:11 -07:00
Andreas Karatzas	3b06c55c78	[ROCm][CI] Fix MEGA_AOT_ARTIFACT fallback when PyTorch < 2.10.0 lacks AOT support (#37763 ) Signed-off-by: Andreas Karatzas <akaratza@amd.com>	2026-03-22 16:02:03 +08:00
Yongye Zhu	87bd91892f	[MoE Refactor] Mxfp4 oracle rebased (#37128 ) Signed-off-by: Yongye Zhu <zyy1102000@gmail.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-03-21 03:37:04 +00:00
Zhengxu Chen	2e089b96a8	[compile] Add compiled artifact counter for VLLM_USE_MEGA_AOT_ARTIFACT=1. (#37589 ) Signed-off-by: zhxchen17 <zhxchen17@fb.com>	2026-03-20 16:22:46 +00:00
Zhengxu Chen	c0f5fae601	[compile] Fix aot test failures with torch 2.12. (#37604 ) Signed-off-by: zhxchen17 <zhxchen17@fb.com>	2026-03-20 16:06:29 +00:00
Xiao	ea2c148fa7	[compile][graph_partition]Add tensor size handling (#36038 ) Signed-off-by: Xiao Fu <xiaofu@meta.com>	2026-03-19 19:55:25 -07:00
Laith Sakka	112944fab9	test Qwen/Qwen3-4B-Instruct-2507 for unbacked (#36064 ) Signed-off-by: Laith Sakka <lsakka@meta.com>	2026-03-19 17:28:45 -04:00
Wentao Ye	0d81a1fe61	[V0 Deprecation] Deprecate virtual engine (#37195 ) Signed-off-by: yewentao256 <zhyanwentao@126.com>	2026-03-18 14:30:14 -07:00
elvischenv	296839a1b0	[Perf] Eliminate padding and slicing op for GPT-OSS with Flashinfer MXFP4 MXFP8 MoE (#30647 ) Signed-off-by: elvischenv <219235043+elvischenv@users.noreply.github.com>	2026-03-18 15:01:26 +00:00
Terry Gao	3e6a1e1686	[Custom Ops] Add functional + out variant for scaled_fp4_quant (#34389 ) Signed-off-by: tianrengao <terrygao87@gmail.com>	2026-03-16 18:51:46 -04:00
Rohan Potdar	a4ad9db541	Enable RoPE+KV cache fusion for ROCm AITER FA (non-shuffle layout) (#35786 ) Signed-off-by: Rohan138 <rohanpotdar138@gmail.com>	2026-03-13 07:33:22 +00:00
Kunshang Ji	53ec16a705	[Hardware] Replace torch.cuda.device_count/current_device/set_device API (#36145 ) Signed-off-by: Kunshang Ji <jikunshang95@gmail.com> Signed-off-by: Kunshang Ji <kunshang.ji@intel.com>	2026-03-12 07:57:47 -07:00
Luka Govedič	9556af87d5	[torch.compile] Add support for non-contiguous fused RMSNorm + group quant (#36551 ) Signed-off-by: Luka Govedič <lgovedic@redhat.com> Signed-off-by: Luka Govedič <ProExpertProg@users.noreply.github.com> Co-authored-by: Copilot <198982749+Copilot@users.noreply.github.com> Co-authored-by: ProExpertProg <11367180+ProExpertProg@users.noreply.github.com>	2026-03-11 10:56:55 -07:00
Richard Zou	822e250ab7	[torch.compile] Use FakeTensors instead of real GPU tensors for single-size compilation (#36093 ) Signed-off-by: Richard Zou <zou3519@gmail.com>	2026-03-11 16:07:09 +00:00
Richard Zou	09b6f99852	[compile] aot_compile should respect VLLM_DISABLE_COMPILE_CACHE (#36358 ) Signed-off-by: Richard Zou <zou3519@gmail.com>	2026-03-11 03:12:03 -07:00
Jiangyun Zhu	ca5fb4bbd8	[Bugfix] Avoid merging empty-only partitions into splitting-op subgraphs (#36595 ) Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>	2026-03-10 07:39:01 -07:00
Copilot	4b87ffbefb	[torch.compile] Rename `compile_ranges_split_points` to `compile_ranges_endpoints` (#36027 ) Signed-off-by: Luka Govedič <ProExpertProg@users.noreply.github.com> Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com> Co-authored-by: ProExpertProg <11367180+ProExpertProg@users.noreply.github.com> Co-authored-by: Luka Govedič <ProExpertProg@users.noreply.github.com>	2026-03-09 18:04:40 +00:00
Jiangyun Zhu	e5ff140216	[cudagraph] fix cudagraph warning in deepseekv32 (#28044 ) Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>	2026-03-08 20:27:41 -04:00
Zhengxu Chen	a97954b6a8	[compile] Consistent compiler config for saved/loaded vllm backends. (#35810 ) Signed-off-by: zhxchen17 <zhxchen17@fb.com>	2026-03-05 15:08:12 -05:00
Jiayi Yan	6a895197fa	[Bugfix][CI] fix typos (#34934 ) Signed-off-by: 1195343015 <1195343015@qq.com> Signed-off-by: Jiayi Yan <66017932+1195343015@users.noreply.github.com> Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2026-03-05 17:05:46 +00:00
Kunshang Ji	66a2209645	[Hardware] Replace `torch.cuda.synchronize()` api with `torch.accelerator.synchronize` (#36085 ) Signed-off-by: Kunshang Ji <kunshang.ji@intel.com>	2026-03-05 10:36:39 +00:00
Zhengxu Chen	dd6dbd93f8	[compile] Fix extra cache save on warm start. (#35921 ) Signed-off-by: zhxchen17 <zhxchen17@fb.com>	2026-03-05 12:56:30 +08:00
Richard Zou	5569f5218d	[torch.compile] Stop lazily compiling (#35472 ) Signed-off-by: Richard Zou <zou3519@gmail.com>	2026-03-04 12:13:17 -08:00
Stefano Castagnetta	d7166e74c1	[CI] Add Blackwell AsyncTP correctness test (#35871 ) Signed-off-by: Stefano Castagnetta <scastagnetta@nvidia.com>	2026-03-04 19:41:21 +00:00
Bhuminjay Soni	fb3e78ab09	[Feature][CI]: compare `func` & `no_func` outputs in test_functionalization.py (#35481 ) Signed-off-by: Bhuminjay <bhuminjaysoni@gmail.com> Signed-off-by: Bhuminjay Soni <Soni5Happy@gmail.com> Co-authored-by: Luka Govedič <ProExpertProg@users.noreply.github.com>	2026-03-04 18:01:16 +00:00
haosdent	d6e04f4c43	[Bugfix] Cap FULL decode cudagraph sizes for Mamba/hybrid models (#34094 ) (#34571 ) Signed-off-by: haosdent <haosdent@gmail.com> Co-authored-by: zjy0516 <riverclouds.zhu@qq.com>	2026-03-04 11:56:22 +01:00
Kunshang Ji	16d2ad1d38	[Hardware] Replace `torch.cuda.empty_cache` with `torch.accelerator.empty_cache` (#30681 ) Signed-off-by: Kunshang Ji <kunshang.ji@intel.com> Signed-off-by: Kunshang Ji <jikunshang95@gmail.com> Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2026-03-04 09:49:47 +00:00
TJian	fb7fdc49c4	[ROCm] [CI] Add new fusion test cases that are relevant to vLLM IR Ops (#34307 ) Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com> Signed-off-by: vllmellm <vllm.ellm@embeddedllm.com> Co-authored-by: vllmellm <vllm.ellm@embeddedllm.com>	2026-03-03 06:24:21 -08:00
Richard Zou	d1a6e96d9e	[torch.compile] Improve cold and warm start compile tests (#35709 ) Signed-off-by: Richard Zou <zou3519@gmail.com>	2026-03-02 19:27:06 +00:00
Itay Alroy	dea268336f	[1/N] Elastic EP Milestone 2 (#34861 ) Signed-off-by: Yongji Wu <wuyongji317@gmail.com> Signed-off-by: Itay Alroy <ialroy@nvidia.com> Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com> Signed-off-by: Ron Tourgeman <rtourgeman@nvidia.com> Co-authored-by: Yongji Wu <wuyongji317@gmail.com> Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com> Co-authored-by: Ron Tourgeman <rtourgeman@nvidia.com>	2026-02-28 04:46:42 +00:00
Zhengxu Chen	29b35477b0	[compile] Fix caching error over pytree slice node. (#35308 ) Signed-off-by: zhxchen17 <zhxchen17@fb.com>	2026-02-27 19:34:16 +00:00
Jason Li	9d37941017	[torch.compile] Sequence Parallelism threshold compile ranges (#28672 ) Signed-off-by: jasonlizhengjian <jasonlizhengjian@gmail.com> Signed-off-by: Jason Li <jasonlizhengjian@gmail.com> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> Co-authored-by: Luka Govedič <ProExpertProg@users.noreply.github.com>	2026-02-26 05:00:12 +00:00
Hanjie Qiu	71dfce6aa6	[Kernel] Refactor FlashInfer allreduce for mnnvl backend (#34109 ) Signed-off-by: hjjq <50634613+hjjq@users.noreply.github.com> Signed-off-by: wzhao18 <wzhao18.sz@gmail.com> Co-authored-by: wzhao18 <wzhao18.sz@gmail.com> Co-authored-by: Wei Zhao <51183510+wzhao18@users.noreply.github.com>	2026-02-26 03:17:20 +00:00
Rohan Potdar	f38f8c9742	[ROCm]: Enable customop and rope+kvcache fusion for AITER RoPE (#35180 ) Signed-off-by: Rohan138 <rohanpotdar138@gmail.com>	2026-02-25 04:36:40 +00:00
BadrBasowid	6af03f2394	[Refactor] [1/N] Reorganize kernel abstraction directory (#34055 ) Signed-off-by: BadrBasowid <badr.basowid@gmail.com> Co-authored-by: vllmellm <vllm.ellm@embeddedllm.com> Co-authored-by: TJian <tunjian.tan@embeddedllm.com>	2026-02-24 06:47:22 +00:00

1 2 3 4 5 ...

304 Commits