biondizzle/vllm - vllm - Gitea: Git with a cup of tea

Author	SHA1	Message	Date
Benjamin Bartels	0e5a9382af	[Bugfix] accept redacted thinking blocks in Anthropic messages (#36992 ) Signed-off-by: Benjamin Bartels <benjaminba@tiglab-ubuntu.ilab.local> Signed-off-by: bbartels <benjamin@bartels.dev> Co-authored-by: Benjamin Bartels <benjaminba@tiglab-ubuntu.ilab.local>	2026-03-16 22:01:57 +08:00
Fynn Schmitt-Ulms	04bf5a35fa	[Spec Decode] Update extract_hidden_states to use deferred kv_connector clear (#37013 )	2026-03-16 14:53:45 +01:00
Robin Nabel	bf9a185395	GLM4 tool parser: fix streaming mode (#35208 ) Signed-off-by: Robin Nabel <opensource@nabel.co> Co-authored-by: Chauncey <chaunceyjiang@gmail.com>	2026-03-16 18:48:52 +08:00
Kunshang Ji	747b068136	[Hardware] Replace memory related torch.cuda APIs (#37031 ) Signed-off-by: Kunshang Ji <jikunshang95@gmail.com>	2026-03-16 10:24:48 +00:00
haosdent	116ed130f4	[Bugfix] Fix GDN attention crash with mixed decode/spec-decode batches (#34871 ) Signed-off-by: haosdent <haosdent@gmail.com>	2026-03-16 10:30:23 +01:00
Isotr0py	912fbe9555	[Bugfix] Fix Qwen2.5-Omni/Qwen3-Omni use_audio_in_video with multi-video inputs (#37147 ) Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn>	2026-03-16 08:56:06 +00:00
Roy Wang	821eb80c0d	[Performance][Model Loader] Skip non-local expert weights during EP model loading (#37136 ) Signed-off-by: esmeetu <jasonailu87@gmail.com>	2026-03-16 01:33:36 -07:00
Andreas Karatzas	a2956a0f8e	[ROCm][CI] Retrying in case of batch variance effects and reducing flakiness (#36442 ) Signed-off-by: Andreas Karatzas <akaratza@amd.com>	2026-03-16 16:08:51 +08:00
Andreas Karatzas	911355e216	[ROCm] Fix KV copy methods and auto-select attention backend for ROCm (#36845 ) Signed-off-by: Andreas Karatzas <akaratza@amd.com>	2026-03-16 16:07:27 +08:00
leo-cf-tian	2754231ba3	[Kernel] Add FlashInfer MoE A2A Kernel (#36022 ) Signed-off-by: wzhao18 <wzhao18.sz@gmail.com> Signed-off-by: Leo Tian <lctian@nvidia.com> Co-authored-by: wzhao18 <wzhao18.sz@gmail.com> Co-authored-by: Stefano Castagnetta <scastagnetta@nvidia.com> Co-authored-by: root <root@lyris0267.lyris.clusters.nvidia.com>	2026-03-15 23:45:32 -07:00
bigshanedogg	2390d44209	[Model] Add HyperCLOVAX-SEED-Think-14B language model support (#37107 ) Signed-off-by: bigshanedogg <bigshane319@gmail.com>	2026-03-16 06:40:05 +00:00
Andreas Karatzas	d4c57863f7	[ROCm][CI] Fix engine teardown and text normalization to stabilize voxtral test (#37138 ) Signed-off-by: Andreas Karatzas <akaratza@amd.com>	2026-03-16 04:49:31 +00:00
Andrew Xia	e9163b536e	[responsesAPI][ez] add a unit test for SimpleContext logprobs (#37126 ) Signed-off-by: Andrew Xia <axia@meta.com>	2026-03-15 17:12:26 -07:00
Lalithnarayan C	7acaea634c	In-Tree AMD Zen CPU Backend via zentorch [1/N] (#35970 ) Signed-off-by: Lalithnarayan C <Lalithnarayan.C@amd.com> Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com> Co-authored-by: Chinmay-Kulkarni-AMD <Chinmay.Kulkarni@amd.com> Co-authored-by: Tyler Michael Smith <tyler@neuralmagic.com> Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-03-15 23:35:35 +00:00
Hari	a3e2e250f0	[Feature] Add Azure Blob Storage support for RunAI Model Streamer (#34614 ) Signed-off-by: hasethuraman <hsethuraman@microsoft.com>	2026-03-15 19:38:21 +08:00
Isotr0py	143e4dccdf	[Misc] Add online audio_in_video test (#36775 ) Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn>	2026-03-15 00:14:11 -07:00
arlo	8c29042bb9	[Feature] Add InstantTensor weight loader (#36139 )	2026-03-14 18:05:23 +01:00
Sergey Zinchenko	4a718e770d	[Bug] Fix Failure in /v1/chat/completions/render for Multimodal Requests (https://github.com/vllm-project/vllm/issues/35665 ) (#35684 )	2026-03-14 14:10:11 +00:00
Flora Feng	bcfdadb1bc	[Refactor] Relocate chat completion and anthropic tests (#36919 ) Signed-off-by: sfeng33 <4florafeng@gmail.com>	2026-03-14 12:16:16 +08:00
Andrew Xia	f680dc1b39	[responsesAPI] prioritize content over summary in reasoning item input (#36516 ) Signed-off-by: Andrew Xia <axia@meta.com> Signed-off-by: Andrew Xia <mitandrewxia@gmail.com> Signed-off-by: Andrew Xia <axia@fb.com> Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> Co-authored-by: Andrew Xia <axia@fb.com>	2026-03-14 09:20:30 +08:00
Giulio Leone	b41aa264f9	fix: resolve chat template names before kwargs detection (#36937 ) Co-authored-by: giulio-leone <giulio.leone@users.noreply.github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>	2026-03-14 00:20:16 +00:00
Benjamin Chislett	8b346309a5	[Refactor] Consolidate SupportsEagle (#36063 ) Signed-off-by: Benjamin Chislett <bchislett@nvidia.com>	2026-03-13 23:22:40 +00:00
Kevin H. Luu	f1816fb192	[CI] Split V1 e2e + engine (1 GPU) into separate jobs (#36945 ) Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>	2026-03-13 14:16:02 -07:00
Harry Mellor	0005d2a3c9	Use Transformers v5 `WeightRenaming` for Transformers modeling backend (#31545 ) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2026-03-13 20:49:08 +00:00
Mark McLoughlin	7afe0faab1	[Frontend][Core] Re-add shutdown timeout - allowing in-flight requests to finish (#36666 ) Signed-off-by: Mark McLoughlin <markmc@redhat.com> Signed-off-by: Nick Hill <nickhill123@gmail.com> Co-authored-by: Claude Sonnet 4.5 <noreply@anthropic.com> Co-authored-by: Nick Hill <nickhill123@gmail.com>	2026-03-13 12:10:06 -07:00
yugong333	b3ce711b93	Fp8 lora dense kernel (#35242 ) Signed-off-by: Yu Gong <yu3.gong@gmail.com>	2026-03-13 19:05:08 +00:00
Isotr0py	abf61aaa8e	[Bugfix] Fix Qwen2.5-omni/Qwen3-omni mm_processor cache for audio_in_video request (#36800 ) Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn>	2026-03-13 18:16:05 +00:00
Itay Alroy	d5af196c18	[2/N] Elastic EP Milestone 2: Integrating NIXL-EP (#35627 ) Signed-off-by: Itay Alroy <ialroy@nvidia.com> Co-authored-by: Yongji Wu <wuyongji317@gmail.com> Co-authored-by: Ron Tourgeman <rtourgeman@nvidia.com>	2026-03-13 09:25:33 -04:00
Or Ozeri	cfaf4668f7	[kv_offload+HMA][1/N]: Support multiple KV groups in OffloadingSpec (#36610 ) Signed-off-by: Or Ozeri <oro@il.ibm.com>	2026-03-13 08:04:21 +00:00
Sage	a2268617cf	[Frontend] Delegate preprocessing to `OpenAIServingRender` (#36483 ) Signed-off-by: Sage Ahrac <sagiahrak@gmail.com>	2026-03-13 00:39:43 -07:00
Rohan Potdar	a4ad9db541	Enable RoPE+KV cache fusion for ROCm AITER FA (non-shuffle layout) (#35786 ) Signed-off-by: Rohan138 <rohanpotdar138@gmail.com>	2026-03-13 07:33:22 +00:00
Nick Hill	b373b5102a	[Tests] Shutdown test `RemoteVLLMServer` cleanly (#36950 ) Recent PR #33949 changed the teardown logic of the RemoteVLLMServer test utility class to send SIGTERM to all vllm (sub)processes at once, which breaks the clean/coordinated shutdown logic that assumes only the top-level process will receive a signal (for example when running in a container that's shut down). This caused a bunch of errors and stacktraces in some test logs, even though those tests still pass. We should still attempt a normal shutdown and only kill other procs if they are still running after a few seconds. Example: tests/v1/distributed/test_external_lb_dp.py::test_external_lb_completion_streaming Signed-off-by: Nick Hill <nickhill123@gmail.com>	2026-03-13 07:32:55 +00:00
Csrayz	bc2c0c86ef	[Frontend] Fix usage incorrectly returned with empty stream_options` (#36379 ) Signed-off-by: Csrayz <33659823+Csrayz@users.noreply.github.com>	2026-03-13 03:33:04 +00:00
whyiug	1ce13cf992	[Model] Add support for BERT-like Chinese ERNIE pooling models (#36385 ) Signed-off-by: whyiug <whyiug@hotmail.com> Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io>	2026-03-13 03:23:53 +00:00
Nikita	10f08dedfa	[Model] Add ColPali late interaction model for multi-modal retrieval (#36818 ) Signed-off-by: Nikita Sukharev <kaonael@gmail.com> Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>	2026-03-13 02:18:57 +00:00
Xinan Miao	2cdf92228c	[Feature]: Remove Chunking From FusedMoE (#34086 ) Signed-off-by: SouthWest7 <am1ao@qq.com> Signed-off-by: Southwest <1403572259@qq.com> Signed-off-by: southwest <am1ao@qq.com> Signed-off-by: Xinan Miao <1403572259@qq.com> Co-authored-by: SouthWest7 <am1ao@qq.com>	2026-03-12 14:24:38 -04:00
Marc Sun	c973ecdead	[bnb] Skip moe + bnb test (#36896 ) Signed-off-by: Marc Sun <marc@huggingface.co>	2026-03-12 18:03:25 +00:00
Eunkwang Jeon	bdc2343454	[Bugfix] Fix KeyError in parse_response_input for reasoning items with optional content (#34499 ) Signed-off-by: jeonsworld <jeonsworld@gmail.com>	2026-03-13 00:13:36 +08:00
SoluMilken	85199f9681	[Bugfix] fix main branch pre-commit error (1 line change) (#36897 ) Signed-off-by: SoluMilken <ypiheyn.imm02g@g2.nctu.edu.tw>	2026-03-12 09:08:37 -07:00
grimulkan	a1257fd1ea	[Kernel] Add FP8 KV cache support to Triton MLA decode attention (#34597 ) Signed-off-by: grimulkan <grimulkan@gmail.com>	2026-03-12 08:32:34 -07:00
Kunshang Ji	53ec16a705	[Hardware] Replace torch.cuda.device_count/current_device/set_device API (#36145 ) Signed-off-by: Kunshang Ji <jikunshang95@gmail.com> Signed-off-by: Kunshang Ji <kunshang.ji@intel.com>	2026-03-12 07:57:47 -07:00
Wei Zhao	2e693f48e7	[Perf] Add TRTLLM FP8 MoE Modular Kernel (#36307 ) Signed-off-by: wzhao18 <wzhao18.sz@gmail.com> Co-authored-by: Michael Goin <mgoin64@gmail.com>	2026-03-12 07:32:31 -07:00
Martin Hickey	7f1f36bf91	[CI] Fix mypy for vllm/reasoning (#35742 ) Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com> Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com> Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2026-03-12 12:21:33 +00:00
caozuoba	9e19f8338b	[Perf] add packed recurrent fast path for decode (#36596 ) Signed-off-by: hdj <1293066020@qq.com> Co-authored-by: Roger Wang <hey@rogerw.io>	2026-03-12 04:01:57 -07:00
Chauncey	5a71cdd76e	[Bugfix] Fix crash when tool_choice=required exceeds max_tokens (#36841 ) Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com>	2026-03-12 03:28:45 -07:00
Shanshan Shen	f0d3658c0f	[MM][OOT] Support CPU `seq_lens` for OOT MMEncoderAttention kernels (#36605 ) Signed-off-by: shen-shanshan <467638484@qq.com> Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn> Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn>	2026-03-12 03:28:23 -07:00
sfeiqiang	8cb24d3aed	[KV Connector] Support using FlexKV as KV Cache Offloading option. (#34328 ) Signed-off-by: phaedonsun <phaedonsun@tencent.com> Co-authored-by: phaedonsun <phaedonsun@tencent.com>	2026-03-12 00:46:20 -07:00
István Ketykó	00726c74c9	[Bugfix][Model] Fix DeepSeek-OCR TensorSchema crash on empty images_crop (#36670 ) Signed-off-by: István Ketykó <istvan.ketyko@gmail.com>	2026-03-12 15:35:54 +08:00
Chauncey	9fe404ed04	[Frontend] OpenAI Responses API supports Tool/Function calling with streaming (#29947 ) Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com>	2026-03-12 15:03:50 +08:00
Sage	802f306cd1	[Tests] Skip model weight download for render-only test server (#36813 ) Signed-off-by: Sage Ahrac <sagiahrak@gmail.com>	2026-03-12 06:24:42 +00:00

1 2 3 4 5 ...

4808 Commits