biondizzle/vllm - vllm - Gitea: Git with a cup of tea

Author	SHA1	Message	Date
Harry Mellor	a0f44bb616	Allow `markdownlint` to run locally (#36398 ) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2026-03-08 20:05:24 -07:00
Xiang Shi	e68de8adc0	docs: fix wrong cc in int8.md (#36209 ) Signed-off-by: Xiang Shi <realkevin@tutanota.com>	2026-03-06 06:01:02 +00:00
BADAOUI Abdennacer	8dc8a99b56	[ROCm] Enable bitsandbytes quantization support on ROCm (#34688 ) Signed-off-by: badaoui <abdennacerbadaoui0@gmail.com>	2026-02-21 00:34:55 -08:00
Harry Mellor	a21cedf4ff	Bump `lm-eval` version for Transformers v5 compatibility (#33994 ) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2026-02-16 05:24:35 -08:00
Tianqi Ren	786806dd44	[Doc] Update Marlin support matrix for Turing (#34319 ) Signed-off-by: Tianqi Ren <tianqi.r@outlook.com>	2026-02-11 09:03:41 +00:00
danisereb	084aa19f02	Add support for ModelOpt MXFP8 dense models (#33786 ) Signed-off-by: Daniel Serebrenik <daserebrenik@nvidia.com>	2026-02-08 11:16:48 -08:00
Michael Goin	29fba76781	[UX] Use gguf `repo_id:quant_type` syntax for examples and docs (#33371 ) Signed-off-by: mgoin <mgoin64@gmail.com>	2026-01-31 12:14:54 +08:00
Aidan Reilly	133765760b	[Docs] Adding links and intro to Speculators and LLM Compressor (#32849 ) Signed-off-by: Aidan Reilly <aireilly@redhat.com> Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com> Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2026-01-29 14:12:35 -08:00
Robert Shaw	247d1a32ea	[Quantization][Deprecation] Remove BitBlas (#32683 ) Signed-off-by: Robert Shaw <robshaw@redhat.com> Signed-off-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com> Co-authored-by: Robert Shaw <robshaw@redhat.com>	2026-01-28 11:06:22 +00:00
Eldar Kurtić	44f08af3a7	Add llmcompressor fp8 kv-cache quant (per-tensor and per-attn_head) (#30141 ) Signed-off-by: Eldar Kurtic <8884008+eldarkurtic@users.noreply.github.com> Signed-off-by: eldarkurtic <8884008+eldarkurtic@users.noreply.github.com>	2026-01-22 13:29:57 -07:00
Michael Goin	6388b50058	[Docs] Add docs about OOT Quantization Plugins (#32035 ) Signed-off-by: mgoin <mgoin64@gmail.com>	2026-01-14 15:25:45 +08:00
Yi Liu	50632adc58	Consolidate Intel Quantization Toolkit Integration in vLLM (#31716 ) Signed-off-by: yiliu30 <yi4.liu@intel.com>	2026-01-14 07:11:30 +00:00
Andrew Bennett	f243abc92d	Fix various typos found in `docs` (#32212 ) Signed-off-by: Andrew Bennett <potatosaladx@meta.com>	2026-01-13 03:41:47 +00:00
Cyrus Leung	d201807339	[Chore] Bump `lm-eval` version (#31264 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>	2025-12-24 05:39:13 -08:00
CedricHuang	19cc9468fd	[Feature]: Support NVIDIA ModelOpt HF FP8 variants FP8_PER_CHANNEL_PER_TOKEN and FP8_PB_WO in vLLM (#30957 )	2025-12-21 22:34:49 -05:00
Steve Westerhouse	9d701e90d8	[Doc] Clarify FP8 KV cache computation workflow (#31071 ) Signed-off-by: westers <steve.westerhouse@origami-analytics.com>	2025-12-22 08:41:37 +08:00
Zhiyu	cd00c443d2	[Misc] Rename TensorRT Model Optimizer to Model Optimizer (#30091 ) Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>	2025-12-08 07:05:27 +00:00
Tyler Michael Smith	4dd42db566	Remove VLLM_SKIP_WARMUP tip (#29331 ) Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com>	2025-11-24 22:16:05 +00:00
Rob Mulla	dd39f91edb	[Doc] cleanup TPU documentation and remove outdated examples (#29048 ) Signed-off-by: Rob Mulla <rob.mulla@gmail.com> Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com> Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2025-11-21 00:05:59 +00:00
Didier Durand	7ed27f3cb5	[Doc]: fix typos in various files (#28945 ) Signed-off-by: Didier Durand <durand.didier@gmail.com>	2025-11-18 22:52:30 -08:00
Uranus	6a25ea5f0e	[Docs] Update oneshot imports (#28188 ) Signed-off-by: UranusSeven <109661872+UranusSeven@users.noreply.github.com>	2025-11-19 05:30:08 +00:00
Didier Durand	2bb4435cb7	[Doc]: fix typos in various files (#28567 ) Signed-off-by: Didier Durand <durand.didier@gmail.com>	2025-11-15 19:27:50 +00:00
xuebwang-amd	05576df85c	[ROCm][Quantization] extend AMD Quark to support mixed-precision quantized model (#24239 ) Signed-off-by: xuebwang-amd <xuebwang@amd.com> Co-authored-by: fxmarty-amd <felmarty@amd.com> Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>	2025-11-11 12:05:22 -05:00
Harry Mellor	4ffd6e8942	[Docs] Reduce custom syntax used in docs (#27009 ) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2025-10-16 20:05:34 -07:00
wangxiyuan	8f4b313c37	[Misc] rename torch_dtype to dtype (#26695 ) Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>	2025-10-15 12:11:48 +00:00
Cyrus Leung	6256697997	[Doc] ruff format remaining Python examples (#26795 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>	2025-10-15 01:25:49 -07:00
HDCharles	b92ab3deda	Notice for deprecation of AutoAWQ (#26820 ) Signed-off-by: HDCharles <39544797+HDCharles@users.noreply.github.com> Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>	2025-10-14 13:39:59 -07:00
fxmarty-amd	41f1cf38f2	[Feature][OCP MX] Support mxfp6 and mixed mxfp6-mxfp4 (#21166 )	2025-10-07 09:35:26 -04:00
Param	99028fda44	Fix INT8 quantization error on Blackwell GPUs (SM100+) (#25935 ) Signed-off-by: padg9912 <phone.and.desktop@gmail.com>	2025-09-30 19:19:53 -07:00
Harry Mellor	abc7989adc	[Docs] Remove Neuron install doc as backend no longer exists (#24396 ) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2025-09-13 00:15:03 -07:00
Harry Mellor	6421b66bf4	[Docs] Move quant supported hardware table to README (#23663 ) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2025-08-26 22:26:46 +00:00
Didier Durand	47455c424f	[Doc: ]fix various typos in multiple files (#23487 ) Signed-off-by: Didier Durand <durand.didier@gmail.com> Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com> Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2025-08-25 00:04:04 +00:00
Cyrus Leung	8896eb72eb	[Deprecation] Remove `prompt_token_ids` arg fallback in `LLM.generate` and `LLM.embed` (#18800 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>	2025-08-22 10:56:57 +08:00
Michael Goin	4fc722eca4	[Kernel/Quant] Remove AQLM (#22943 ) Signed-off-by: mgoin <mgoin64@gmail.com> Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>	2025-08-16 19:38:21 +00:00
Harry Mellor	7be7f3824a	[Docs] Improve API docs (+small tweaks) (#22459 ) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2025-08-08 03:02:51 -07:00
Harry Mellor	ba5c5e5404	[Docs] Switch to better markdown linting pre-commit hook (#21851 ) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2025-07-29 19:45:08 -07:00
Michael Goin	947e982ede	[Docs] Minimize spacing for supported_hardware.md table (#21779 )	2025-07-28 18:46:39 -07:00
Wenhua Cheng	5ac3168ee3	[Docs] add auto-round quantization readme (#21600 ) Signed-off-by: Wenhua Cheng <wenhua.cheng@intel.com> Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2025-07-25 08:52:42 -07:00
Ning Xie	d97841078b	[Misc] unify variable for LLM instance (#20996 ) Signed-off-by: Andy Xie <andy.xning@gmail.com>	2025-07-21 12:18:33 +01:00
Harry Mellor	be54a951a3	[Docs] Fix hardcoded links in docs (#21287 ) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2025-07-21 02:23:57 -07:00
Nir David	01513a334a	Support FP8 Quantization and Inference Run on Intel Gaudi (HPU) using INC (Intel Neural Compressor) (#12010 ) Signed-off-by: Nir David <ndavid@habana.ai> Signed-off-by: Uri Livne <ulivne@habana.ai> Co-authored-by: Uri Livne <ulivne@habana.ai>	2025-07-16 15:33:41 -04:00
fxmarty-amd	332d4cb17b	[Feature][Quantization] MXFP4 support for MOE models (#17888 ) Signed-off-by: Felix Marty <felmarty@amd.com> Signed-off-by: Bowen Bao <bowenbao@amd.com> Signed-off-by: Felix Marty <Felix.Marty@amd.com> Co-authored-by: Bowen Bao <bowenbao@amd.com>	2025-07-09 13:19:02 -07:00
Harry Mellor	b942c094e3	Stop using title frontmatter and fix doc that can only be reached by search (#20623 ) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2025-07-08 03:27:40 -07:00
Harry Mellor	b4bab81660	Remove unnecessary explicit title anchors and use relative links instead (#20620 ) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2025-07-08 02:49:13 -07:00
Harry Mellor	af107d5a0e	Make distinct `code` and `console` admonitions so readers are less likely to miss them (#20585 ) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2025-07-07 19:55:28 -07:00
Jee Jee Li	1819fbda63	[Quantization] Bump to use latest bitsandbytes (#20424 ) Signed-off-by: Jee Jee Li <pandaleefree@gmail.com>	2025-07-03 21:58:46 +08:00
Lukas Geiger	c3649e4fee	[Docs] Fix syntax highlighting of shell commands (#19870 ) Signed-off-by: Lukas Geiger <lukas.geiger94@gmail.com>	2025-06-23 17:59:09 +00:00
Reid	f17aec0d63	[doc] Fold long code blocks to improve readability (#19926 ) Signed-off-by: reidliu41 <reid201711@gmail.com> Co-authored-by: reidliu41 <reid201711@gmail.com>	2025-06-23 05:24:23 +00:00
Cyrus Leung	a5115f4ff5	[Doc] Fix quantization link titles (#19478 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>	2025-06-11 01:27:22 -07:00
aws-elaineyz	1661a9c28f	[Doc][Neuron] Update documentation for Neuron (#18868 ) Signed-off-by: Elaine Zhao <elaineyz@amazon.com>	2025-05-28 19:44:01 -07:00

1 2

53 Commits