Cyrus Leung
|
82551ad616
|
[Core] Don't use cache during multi-modal profiling (#14336)
|
2025-03-06 08:03:31 -08:00 |
|
courage17340
|
caac5c2e59
|
[Bugfix][Core] fix abort_seq_group and memory leak when n>1 (#14326)
Signed-off-by: courage17340 <courage17340@163.com>
|
2025-03-06 23:59:32 +08:00 |
|
Thomas Parnell
|
6bd1dd9d26
|
[Kernel] [V1] Improved performance for V1 Triton (ROCm) backend (#14152)
|
2025-03-06 07:39:16 -08:00 |
|
Nicolò Lucchesi
|
fa82b93853
|
[Frontend][Docs] Transcription API streaming (#13301)
Signed-off-by: NickLucche <nlucches@redhat.com>
|
2025-03-06 10:39:35 +00:00 |
|
Nicolò Lucchesi
|
69ff99fdcd
|
[Core] Optimizing cross-attention QKVParallelLinear computation (#12325)
Signed-off-by: NickLucche <nlucches@redhat.com>
Signed-off-by: NickLucche <nick@nlucches-4xa100.c.openshift-330514.internal>
Co-authored-by: NickLucche <nick@nlucches-4xa100.c.openshift-330514.internal>
|
2025-03-06 09:37:26 +00:00 |
|
lkchen
|
5d802522a7
|
[V1][VLM][Pixtral-HF] Support Pixtral-HF on V1 (#14275)
Signed-off-by: Linkun Chen <github@lkchen.net>
|
2025-03-06 08:58:41 +00:00 |
|
kYLe
|
1769928079
|
[Model] Update Paligemma multimodal processing with PromptUpdate (#14015)
Signed-off-by: Kyle Huang <kylhuang@nvidia.com>
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>
|
2025-03-06 08:31:38 +00:00 |
|
Pavani Majety
|
ed6ea06577
|
[Hardware] Update the flash attn tag to support Blackwell (#14244)
|
2025-03-05 22:01:37 -08:00 |
|
Nicolò Lucchesi
|
5ee10e990d
|
[Bugfix][CI] ALiBi test case in xformers multi_query_kv_attention (#11301)
|
2025-03-05 20:00:53 -08:00 |
|
Ce Gao
|
f5f7f00cd9
|
[Bugfix][Structured Output] Support outlines engine with reasoning outputs for DeepSeek R1 (#14114)
|
2025-03-06 03:49:20 +00:00 |
|
Rui Qiao
|
abcc61e0af
|
[misc] Mention ray list nodes command to troubleshoot ray issues (#14318)
Signed-off-by: Rui Qiao <ruisearch42@gmail.com>
|
2025-03-06 02:00:36 +00:00 |
|
Lucas Wilkinson
|
f6bb18fd9a
|
[BugFix] MLA + V1, illegal memory access and accuracy issues (#14253)
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
|
2025-03-05 17:10:13 -08:00 |
|
Lucas Wilkinson
|
4dacaa4a83
|
[BugFix] Fix prefix caching V0 MLA (#14255)
Signed-off-by: Lucas Wilkinson <lwilkinson@neuralmagic.com>
Co-authored-by: Ying Zhong <zhongyingmatrix@gmail.com>
|
2025-03-05 17:07:42 -08:00 |
|
Tyler Michael Smith
|
a7ea35aa67
|
[Bugfix] Remove num_tokens_across_dp (#14302)
Signed-off-by: Tyler Michael Smith <tyler@neuralmagic.com>
|
2025-03-05 23:55:55 +00:00 |
|
pyc96
|
1e3e76b6cc
|
[Bugfix] Fix DeepSeek MTP crash when using TP1ModelRunner with CUDA graph due to shape mismatch (#14237)
Signed-off-by: pyc96 <pychen96@gmail.com>
|
2025-03-05 22:22:40 +00:00 |
|
Serena
|
1b7624bf5c
|
[misc] Add FlashMLA as a new option of VLLM_ATTENTION_BACKEND env (#14267)
|
2025-03-05 21:28:50 +00:00 |
|
Nick Hill
|
ac60dc7fe1
|
[V1][BugFix] Fix for mixed top_k batch (#14301)
Signed-off-by: Nick Hill <nhill@redhat.com>
Co-authored-by: Ye Cao <caoye.cao@alibaba-inc.com>
|
2025-03-05 20:43:04 +00:00 |
|
Vincent
|
a4f1ee35d6
|
Deprecate best_of Sampling Parameter in anticipation for vLLM V1 (#13997)
Signed-off-by: vincent-4 <vincentzhongy+githubvincent4@gmail.com>
Signed-off-by: Brayden Zhong <b8zhong@uwaterloo.ca>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2025-03-05 20:22:43 +00:00 |
|
Nick Hill
|
a32c8669ca
|
[V1][Minor] Remove obsolete FIXME comment (#14304)
Signed-off-by: Nick Hill <nhill@redhat.com>
|
2025-03-05 11:59:23 -08:00 |
|
Isotr0py
|
e17e4488bd
|
[LoRA] Remove linear hack outside transformers backend (#14177)
Signed-off-by: Isotr0py <2037008807@qq.com>
|
2025-03-05 15:06:28 +00:00 |
|
Robert Shaw
|
257e200a25
|
[V1][Frontend] Add Testing For V1 Runtime Parameters (#14159)
Signed-off-by: rshaw@neuralmagic.com <rshaw@neuralmagic.com>
|
2025-03-05 14:18:55 +00:00 |
|
Zhe Zhang
|
47d4a7e004
|
Small update for external_launcher backend docs (#14288)
|
2025-03-05 21:30:00 +08:00 |
|
Lu Fang
|
8d6cd32b7b
|
[Bugfix][V1] Fix allowed_token_ids for v1 Sampler (#14169)
Signed-off-by: Lu Fang <lufang@fb.com>
|
2025-03-05 08:49:44 +00:00 |
|
Roger Wang
|
ec79b67c77
|
[Misc][V1] Avoid using envs.VLLM_USE_V1 in mm processing (#14256)
Signed-off-by: Roger Wang <ywang@roblox.com>
|
2025-03-05 07:37:16 +00:00 |
|
Benjamin Chislett
|
32985bed7c
|
[Frontend] Allow return_tokens_as_token_ids to be passed as a request param (#14066)
Signed-off-by: Benjamin Chislett <benjamin.chislett@centml.ai>
|
2025-03-05 06:30:40 +00:00 |
|
youkaichao
|
6eaf93020d
|
[platforms] improve rocm debugging info (#14257)
|
2025-03-04 21:32:18 -08:00 |
|
Tyler Michael Smith
|
72c62eae5f
|
[V1] EP/TP MoE + DP Attention (#13931)
|
2025-03-04 21:27:26 -08:00 |
|
Congcong Chen
|
0a995d5434
|
[Model] New model support for Phi-4-multimodal-instruct (#14119)
|
2025-03-04 20:57:01 -08:00 |
|
Cody Yu
|
ade3f7d988
|
[V1][Bugfix] Do not reset prefix caching metrics (#14235)
|
2025-03-05 04:39:13 +00:00 |
|
rainkert
|
0df25101d6
|
[Bugfix] Fix gptq_marlin for deepseek-v3 (#13750)
Signed-off-by: dangshunya <dangshunya@baichuan-inc.com>
Co-authored-by: dangshunya <dangshunya@baichuan-inc.com>
|
2025-03-05 12:25:53 +08:00 |
|
Michael Goin
|
fbfc3ee37e
|
[V1][TPU] TPU multimodal model support for ragged attention (#14158)
Signed-off-by: Michael Goin <mgoin64@gmail.com>
|
2025-03-04 19:58:48 -05:00 |
|
Tyler Michael Smith
|
4f5b059f14
|
Clean up unused padding_idx variables across many model definitions (#13240)
Signed-off-by: Tyler Michael Smith <tyler@neuralmagic.com>
|
2025-03-04 21:27:00 +00:00 |
|
Kuntai Du
|
288ca110f6
|
[Security] Serialize using safetensors instead of pickle in Mooncake Pipe (#14228)
Signed-off-by: KuntaiDu <kuntai@uchicago.edu>
|
2025-03-04 21:10:32 +00:00 |
|
Harry Mellor
|
e5b2f1601a
|
[Frontend] Do prompt_logprobs clamping for chat as well as completions (#14225)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2025-03-04 20:13:06 +00:00 |
|
Harry Mellor
|
9badee53de
|
Fix performance when --generation-config is not None (#14223)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2025-03-04 20:59:22 +01:00 |
|
Siyuan Liu
|
beebf4742a
|
[TPU][Profiler] Support start_profile/stop_profile in TPU worker (#13988)
Signed-off-by: Siyuan Liu <lsiyuan@google.com>
Co-authored-by: mgoin <mgoin64@gmail.com>
|
2025-03-04 14:40:06 -05:00 |
|
lkchen
|
b3cf368d79
|
[V1][Molmo] Fix get_multimodal_embeddings() in molmo.py (#14161)
|
2025-03-04 15:43:59 +00:00 |
|
Mark McLoughlin
|
c8525f06fc
|
[V0][Metrics] Deprecate some questionable request time metrics (#14135)
Signed-off-by: Mark McLoughlin <markmc@redhat.com>
|
2025-03-04 15:11:33 +00:00 |
|
Nick Hill
|
5db6b2c961
|
[V1][BugFix] Fix remaining sync engine client shutdown errors/hangs (#13869)
Signed-off-by: Nick Hill <nhill@redhat.com>
|
2025-03-04 15:06:47 +00:00 |
|
Michael Goin
|
6247bae6c6
|
[Bugfix] Restrict MacOS CPU detection (#14210)
Signed-off-by: mgoin <mgoin64@gmail.com>
|
2025-03-04 22:25:27 +08:00 |
|
youkaichao
|
71c4b40562
|
[sleep mode] error out with expandable_segments (#14189)
Signed-off-by: youkaichao <youkaichao@gmail.com>
|
2025-03-04 18:54:19 +08:00 |
|
youkaichao
|
ac65bc92df
|
[platform] add debug logging during inferring the device type (#14195)
Signed-off-by: youkaichao <youkaichao@gmail.com>
|
2025-03-04 18:39:16 +08:00 |
|
Zhanwen Chen
|
66233af7b6
|
Use math.prod instead of np.prod for trivial ops (#14142)
|
2025-03-03 21:09:22 -08:00 |
|
Rui Qiao
|
bf13d40972
|
[core] Pass all driver env vars to ray workers unless excluded (#14099)
Signed-off-by: Rui Qiao <ruisearch42@gmail.com>
|
2025-03-04 11:44:17 +08:00 |
|
Cody Yu
|
989f4f430c
|
[Misc] Remove lru_cache in NvmlCudaPlatform (#14156)
Signed-off-by: Cody Yu <hao.yu.cody@gmail.com>
|
2025-03-04 11:09:34 +08:00 |
|
Divakar Verma
|
bb5b640359
|
[core] moe fp8 block quant tuning support (#14068)
Signed-off-by: Divakar Verma <divakar.verma@amd.com>
|
2025-03-04 01:30:23 +00:00 |
|
Travis Johnson
|
c060b71408
|
[Model] Add support for GraniteMoeShared models (#13313)
Signed-off-by: Travis Johnson <tsjohnso@us.ibm.com>
Co-authored-by: Cyrus Leung <cyrus.tl.leung@gmail.com>
|
2025-03-04 08:04:52 +08:00 |
|
iefgnoix
|
79e4937c65
|
[v1] Add comments to the new ragged paged attention Pallas kernel (#14155)
Signed-off-by: Xiongfei Wei <isaacwxf23@gmail.com>
Co-authored-by: Michael Goin <mgoin64@gmail.com>
|
2025-03-03 23:00:55 +00:00 |
|
Michael Goin
|
19d98e0c7d
|
[Kernel] Optimize moe intermediate_cache usage (#13625)
Signed-off-by: mgoin <mgoin64@gmail.com>
|
2025-03-03 16:29:53 -05:00 |
|
Michael Goin
|
2b04c209ee
|
[Bugfix] Allow shared_experts skip quantization for DeepSeekV2/V3 (#14100)
Signed-off-by: mgoin <mgoin64@gmail.com>
|
2025-03-03 14:20:24 -07:00 |
|