biondizzle/vllm - vllm - Gitea: Git with a cup of tea

Author	SHA1	Message	Date
d.transposed	922d3b401b	[Bugfix] Handle the edge case in detokenizer where processed tokens contain both `stop` str and `eos` token (#23938 ) Signed-off-by: dtransposed <damian.bogunowicz@gmail.com>	2025-09-09 07:30:24 -07:00
wang.yuqi	19332c0479	[Model] Systematic support for fp32 head, pooling models part (#23810 ) Signed-off-by: wang.yuqi <noooop@126.com>	2025-09-09 07:29:50 -07:00
Didier Durand	46876dff32	[Doc]: fixing typos to improve docs (#24480 ) Signed-off-by: Didier Durand <durand.didier@gmail.com>	2025-09-08 23:06:04 -07:00
Ming Yang	1823a00d67	[Misc] Support bench serve long context (#24373 ) Signed-off-by: Ming Yang <minos.future@gmail.com>	2025-09-08 22:53:10 -07:00
Cyrus Leung	948dd3443b	[Bugfix] Fix Apertus HF repo name (#24447 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>	2025-09-08 21:40:29 -07:00
Zebing Lin	82dfb12e52	[Core] Use sha256 bytes instead of BlockHash to reduce GC overhead (#23673 ) Signed-off-by: linzebing <linzebing1995@gmail.com>	2025-09-08 21:34:37 -07:00
elvischenv	bba1042c6f	[Flashinfer] Support Flashinfer TRTLLM FP8-qkv BF16/FP16-out Attention Kernel (#23647 ) Signed-off-by: elvischenv <219235043+elvischenv@users.noreply.github.com>	2025-09-08 20:53:07 -07:00
Matthew Bonanni	620db1fc58	[Attention] FlashAttention MLA cudagraph support (#23958 ) Signed-off-by: Matthew Bonanni <mbonanni001@gmail.com> Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>	2025-09-08 22:05:26 +00:00
Jiangyun Zhu	7be141b2c5	[CI] Enable encoder model compilation test (#24442 ) Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>	2025-09-08 11:48:06 -07:00
Jee Jee Li	8d7f39b48c	[Model] Remove quantized mixtral (#24437 ) Signed-off-by: Jee Jee Li <pandaleefree@gmail.com>	2025-09-08 11:02:14 -07:00
Chenheli Hua	01dfb5e982	[Frontend] User-provided uuids for medias in chat. (RFC #22044 ) (#23449 ) Signed-off-by: Roger Wang <hey@rogerw.io> Signed-off-by: Chenheli Hua <huachenheli@outlook.com> Signed-off-by: Roger Wang <hey@rogerw.me> Signed-off-by: Cyrus Leung <cyrus.tl.leung@gmail.com> Co-authored-by: Roger Wang <hey@rogerw.io> Co-authored-by: Roger Wang <hey@rogerw.me> Co-authored-by: Cyrus Leung <cyrus.tl.leung@gmail.com>	2025-09-08 06:42:20 -07:00
Harry Mellor	03dd652c16	Move `KVEventsConfig` from `config/__init__.py` to `config/kv_events.py` (#24433 ) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2025-09-08 06:41:27 -07:00
Christian Pinto	9cd76b71ab	[Misc] Terratorch related fixes (#24337 ) Signed-off-by: Christian Pinto <christian.pinto@ibm.com> Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>	2025-09-08 06:40:26 -07:00
tomeras91	e041314184	[Bugfix] Fix mamba2 prefill chunking (#23279 ) Signed-off-by: Tomer Asida <57313761+tomeras91@users.noreply.github.com> Signed-off-by: tomeras91 <57313761+tomeras91@users.noreply.github.com> Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>	2025-09-08 11:42:41 +00:00
Chenheli Hua	3bca396f79	[CI/Build] Fix local image inputs in test_pixtral.py (#24401 ) Signed-off-by: Chenheli Hua <huachenheli@outlook.com> Co-authored-by: Roger Wang <hey@rogerw.io>	2025-09-08 03:31:35 +00:00
22quinn	3a3e91bdfe	[CI/Build] Disable flaky test_structured_output tests (#24404 ) Signed-off-by: 22quinn <33176974+22quinn@users.noreply.github.com>	2025-09-08 02:51:59 +00:00
Xingyu Liu	b3d7e3c845	[Sampler] Support returning all prompt logprobs (#23868 ) Signed-off-by: Xingyu Liu <charlotteliu12x@gmail.com> Co-authored-by: 22quinn <33176974+22quinn@users.noreply.github.com> Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>	2025-09-07 19:34:31 -07:00
Ming Yang	86173ad593	[Kernel] Support decode context parallelism on Blackwell with CUTLASS MLA (#24385 ) Signed-off-by: Ming Yang <minos.future@gmail.com> Signed-off-by: youkaichao <youkaichao@gmail.com> Co-authored-by: youkaichao <youkaichao@gmail.com>	2025-09-08 09:27:12 +08:00
Flora Feng	0661cb9df3	Add renderer-based prompt processing for embedding and classification endpoints (#24356 ) Signed-off-by: sfeng33 <4florafeng@gmail.com>	2025-09-07 08:26:48 +00:00
Woosuk Kwon	105d3d62ef	[TPU] Remove TopKTopPSampler dependency for TPU sampler (#24391 ) Signed-off-by: Woosuk Kwon <woosuk.kwon@berkeley.edu>	2025-09-07 01:12:36 -07:00
Aaron Pham	e67597545b	[CI][Fix] deterministic seed for flaky CI runs on structured outputs (#24380 ) Signed-off-by: Aaron Pham <contact@aarnphm.xyz>	2025-09-07 11:10:40 +08:00
Woosuk Kwon	4172235ab7	[V0 deprecation] Deprecate V0 Neuron backend (#21159 ) Signed-off-by: Woosuk Kwon <woosuk.kwon@berkeley.edu>	2025-09-06 16:15:18 -07:00
elvischenv	e68dc2f014	[Bugfix] Fix unstable silu_mul+nvfp4 quant fusion test (#24370 ) Signed-off-by: elvischenv <219235043+elvischenv@users.noreply.github.com>	2025-09-06 20:39:34 +00:00
Ye (Charlotte) Qi	a3645ed94d	[Frontend][Responses API] Support reporting tool output tokens and fix reasoning token count (#24285 ) Signed-off-by: Ye (Charlotte) Qi <yeq@meta.com>	2025-09-06 13:27:15 -07:00
Aaron Pham	fb691ee4e7	[Fix] [gpt-oss] fix non-tool calling path for chat completion (#24324 )	2025-09-06 19:10:32 +00:00
Jee Jee Li	7555d6b34a	[Bugfix] Fix test_mixtral_moe (#24371 )	2025-09-06 09:32:03 -07:00
Roger Wang	b121ca22ad	[CI] Disable flaky structured output test from CI (#24366 ) Signed-off-by: Roger Wang <hey@rogerw.io>	2025-09-06 13:31:56 +00:00
wang.yuqi	6d6c6b05d3	[New Model]: google/embeddinggemma-300m (#24318 ) Signed-off-by: wang.yuqi <noooop@126.com>	2025-09-05 22:58:36 -07:00
yzds	ac201a0eaf	[Feature] Support Decode Context Parallel (DCP) for MLA (#23734 ) Signed-off-by: hongchao <hongchao@msh.team> Signed-off-by: youkaichao <youkaichao@gmail.com> Co-authored-by: hongchao <hongchao@msh.team> Co-authored-by: youkaichao <youkaichao@gmail.com>	2025-09-06 13:24:05 +08:00
Didier Durand	35bf193864	[Doc]: fix typos in Python comments (#24294 ) Signed-off-by: Didier Durand <durand.didier@gmail.com> Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>	2025-09-05 19:41:12 -07:00
elvischenv	eedb2a2a10	[Bugfix] Fix silu_mul+quant fusion test (#24341 ) Signed-off-by: elvischenv <219235043+elvischenv@users.noreply.github.com>	2025-09-05 20:13:42 +00:00
Aaron Pham	c29fb540ff	[gpt-oss] tool parser supports for /chat/completions [1/n] (#22386 ) Signed-off-by: Aaron Pham <contact@aarnphm.xyz> Co-authored-by: Simon Mo <simon.mo@hey.com>	2025-09-04 20:39:12 -07:00
Zhuohan Li	886ccbe5ba	[CI/Build] Reduce the number of redundant cases to test for LoRA (#24276 ) Signed-off-by: Zhuohan Li <zhuohan123@gmail.com>	2025-09-04 21:58:44 +00:00
elvischenv	adc3ddb430	[Bugfix][Misc] Fix silu_and_mul_nvfp4_quant issue and extract common utils for nvfp4 kernel source files (#23727 ) Signed-off-by: elvischenv <219235043+elvischenv@users.noreply.github.com> Signed-off-by: Luka Govedič <ProExpertProg@users.noreply.github.com> Co-authored-by: Luka Govedič <ProExpertProg@users.noreply.github.com>	2025-09-04 14:25:45 -07:00
Seiji Eicher	60b755cbcb	[Misc] Have AsyncLLM `custom_stat_loggers` extend default logger list (#20952 ) Signed-off-by: Seiji Eicher <seiji@anyscale.com> Signed-off-by: Seiji Eicher <58963096+eicherseiji@users.noreply.github.com> Co-authored-by: Nick Hill <nhill@redhat.com>	2025-09-04 14:25:30 -07:00
Didier Durand	83609ca91d	[Doc]: fix typos in Python comments (#24173 ) Signed-off-by: Didier Durand <durand.didier@gmail.com> Co-authored-by: Russell Bryant <rbryant@redhat.com> Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>	2025-09-04 08:52:17 -07:00
nvjullin	37241077d5	[Misc] Removed force_fp8_e4m3fnuz from FP8LinearOp (#23725 ) Signed-off-by: Julien Lin <jullin@nvidia.com> Signed-off-by: Luka Govedič <ProExpertProg@users.noreply.github.com> Co-authored-by: Luka Govedič <ProExpertProg@users.noreply.github.com>	2025-09-04 09:25:40 -04:00
Kebe	8f423e5f43	[Feature][Response API] Add streaming support for non-harmony (#23741 ) Signed-off-by: Kebe <mail@kebe7jun.com>	2025-09-04 17:49:06 +08:00
Lucas Wilkinson	402759d472	[Attention] FlashAttn MLA (#14258 ) Signed-off-by: Lucas Wilkinson <lwilkinson@neuralmagic.com> Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com> Signed-off-by: Matthew Bonanni <mbonanni001@gmail.com> Co-authored-by: Matthew Bonanni <mbonanni001@gmail.com> Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>	2025-09-04 02:47:59 -07:00
mgazz	51d5e9be7d	[Core][Model] Terratorch backend integration (#23513 ) Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com> Signed-off-by: Christian Pinto <christian.pinto@ibm.com> Co-authored-by: Christian Pinto <christian.pinto@ibm.com> Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>	2025-09-04 00:22:41 -07:00
bingchen-mi	e7fc70016f	[Model] Add MiDashengLM model support (#23652 ) Signed-off-by: chenbing8 <chenbing8@xiaomi.com> Signed-off-by: bingchen-mi <chenbing8@xiaomi.com> Co-authored-by: Jee Jee Li <pandaleefree@gmail.com> Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn>	2025-09-04 00:08:09 -07:00
Li, Jiang	57b1ce94f7	[CPU] Refactor CPU unquantized linear (#24150 ) Signed-off-by: jiang1.li <jiang1.li@intel.com>	2025-09-04 14:28:45 +08:00
Flora Feng	712b273f65	[Refactor] Introduce basic Renderer for completion-style request (#24010 ) Signed-off-by: sfeng33 <4florafeng@gmail.com>	2025-09-04 05:21:12 +00:00
wuhang	a38f8bd54c	[Feature][Responses API]Support MCP tools with streaming mode + background mode (#23927 ) Signed-off-by: wuhang <wuhang6@huawei.com>	2025-09-04 04:05:10 +00:00
Peter Pan	b5ee1e3261	Remove deprecated `PyNcclConnector` (#24151 ) Signed-off-by: Peter Pan <Peter.Pan@daocloud.io>	2025-09-03 22:49:16 +00:00
Matthew Bonanni	a742322092	[Attention] Blackwell FP8 MLA support with CUTLASS_MLA backend (#23289 ) Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>	2025-09-03 14:05:24 -04:00
bnellnm	e9b92dcd89	[Kernels] Overlap shared experts with send/recv (#23273 ) Signed-off-by: Bill Nell <bnell@redhat.com>	2025-09-03 12:35:18 -04:00
nopperl	fa4311d85f	[V1] v1 engine + full CUDA graph support for PLaMo2 (#23998 ) Signed-off-by: Hemmi Shinichi <shemmi@preferred.jp> Signed-off-by: nopperl <54780682+nopperl@users.noreply.github.com> Co-authored-by: Hemmi Shinichi <shemmi@preferred.jp> Co-authored-by: Thomas Parnell <tom.parnell@gmail.com>	2025-09-03 08:24:02 -07:00
wang.yuqi	51383bd472	[CI] Accelerate mteb test by setting SentenceTransformers mteb score to a constant (#24088 ) Signed-off-by: wang.yuqi <noooop@126.com>	2025-09-03 17:23:56 +08:00
Isotr0py	9c99e4871f	[Misc] Clean up deadcode for legacy processing pipeline (#24153 ) Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn>	2025-09-03 08:34:29 +00:00

1 2 3 4 5 ...

2828 Commits