biondizzle/vllm - vllm - Gitea: Git with a cup of tea

Author	SHA1	Message	Date
Harry Mellor	cdc1fa12eb	Remove unused kwargs from model definitions (#13555 )	2025-02-24 17:13:52 -08:00
Jee Jee Li	105b8ce4c0	[Misc] Reduce LoRA-related static variable (#13166 )	2025-02-22 00:21:30 -08:00
Patrick von Platen	d366ccc4e3	[RFC] [Mistral] FP8 format (#10130 ) Signed-off-by: mgoin <mgoin64@gmail.com> Co-authored-by: mgoin <mgoin64@gmail.com>	2025-02-08 14:12:53 -07:00
Amit Garg	538fab93cd	PR #12718 (#12718 )	2025-02-07 06:22:37 -08:00
Russell Bryant	e489ad7a21	[Misc] Add SPDX-License-Identifier headers to python source files (#12628 ) - Add SPDX license headers to python source files - Check for SPDX headers using pre-commit commit 9d7ef44c3cfb72ca4c32e1c677d99259d10d4745 Author: Russell Bryant <rbryant@redhat.com> Date: Fri Jan 31 14:18:24 2025 -0500 Add SPDX license headers to python source files This commit adds SPDX license headers to python source files as recommended to the project by the Linux Foundation. These headers provide a concise way that is both human and machine readable for communicating license information for each source file. It helps avoid any ambiguity about the license of the code and can also be easily used by tools to help manage license compliance. The Linux Foundation runs license scans against the codebase to help ensure we are in compliance with the licenses of the code we use, including dependencies. Having these headers in place helps that tool do its job. More information can be found on the SPDX site: - https://spdx.dev/learn/handling-license-info/ Signed-off-by: Russell Bryant <rbryant@redhat.com> commit 5a1cf1cb3b80759131c73f6a9dddebccac039dea Author: Russell Bryant <rbryant@redhat.com> Date: Fri Jan 31 14:36:32 2025 -0500 Check for SPDX headers using pre-commit Signed-off-by: Russell Bryant <rbryant@redhat.com> --------- Signed-off-by: Russell Bryant <rbryant@redhat.com>	2025-02-02 11:58:18 -08:00
Pavani Majety	b02fd288b2	[Hardware][NV] Fix Modelopt model loading for k-v-scales for Llama models. (#11787 ) Signed-off-by: Pavani Majety <pmajety@nvidia.com> Co-authored-by: mgoin <michael@neuralmagic.com>	2025-01-29 01:46:12 -08:00
Gregory Shtrasberg	e97f802b2d	[FP8][Kernel] Dynamic kv cache scaling factors computation (#11906 ) Signed-off-by: Gregory Shtrasberg <Gregory.Shtrasberg@amd.com> Co-authored-by: Micah Williamson <micah.williamson@amd.com>	2025-01-23 18:04:03 +00:00
Michael Goin	9aa1519f08	Various cosmetic/comment fixes (#12089 ) Signed-off-by: mgoin <michael@neuralmagic.com>	2025-01-16 09:59:06 +00:00
kewang-xlnx	de0526f668	[Misc][Quark] Upstream Quark format to VLLM (#10765 ) Signed-off-by: kewang-xlnx <kewang@xilinx.com> Signed-off-by: kewang2 <kewang2@amd.com> Co-authored-by: kewang2 <kewang2@amd.com> Co-authored-by: Michael Goin <michael@neuralmagic.com>	2025-01-15 11:05:15 -05:00
RunningLeon	97eb97b5a4	[Model]: Support internlm3 (#12037 )	2025-01-15 11:35:17 +00:00
Jee Jee Li	a3a3ee4e6f	[Misc] Merge bitsandbytes_stacked_params_mapping and packed_modules_mapping (#11924 ) Signed-off-by: Jee Jee Li <pandaleefree@gmail.com>	2025-01-15 07:49:49 +08:00
Isotr0py	d14e98d924	[Model] Support GGUF models newly added in `transformers` 4.46.0 (#9685 ) Signed-off-by: Isotr0py <2037008807@qq.com> Co-authored-by: Cyrus Leung <cyrus.tl.leung@gmail.com>	2025-01-13 00:13:44 +00:00
bjmsong	187e32997c	[Bugfix] Change kv scaling factor by param json on nvidia gpu (#11688 ) Signed-off-by: bjmsong <bjmsong@126.com> Co-authored-by: bjmsong <bjmsong@126.com>	2025-01-02 21:11:39 +00:00
Cyrus Leung	96d673e0f8	[Bugfix] Fix error handling of unsupported sliding window (#11213 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>	2024-12-15 10:59:42 -07:00
Cyrus Leung	c889d5888b	[Doc] Explicitly state that PP isn't compatible with speculative decoding yet (#10975 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>	2024-12-07 17:20:49 +00:00
Cyrus Leung	133707123e	[Model] Replace embedding models with pooling adapter (#10769 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>	2024-12-01 08:02:54 +08:00
Patrick von Platen	e7cfc4ef4c	[Interleaved ATTN] Support for Mistral-8B (#10591 ) Signed-off-by: youkaichao <youkaichao@gmail.com> Co-authored-by: youkaichao <youkaichao@gmail.com>	2024-11-30 07:45:50 +00:00
Cyrus Leung	9b4b150395	[Bugfix] Ignore `lm_head` when loading embedding models (#10719 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>	2024-11-27 19:05:29 +00:00
shunxing12345	1209261e93	[Model] Support telechat2 (#10311 ) Signed-off-by: Isotr0py <2037008807@qq.com> Co-authored-by: xiangw2 <xiangw2@chinatelecom.cn> Co-authored-by: Isotr0py <2037008807@qq.com>	2024-11-27 11:32:35 +00:00
Jee Jee Li	15cc2a9f1a	[Misc]Further reduce BNB static variable (#10597 ) Signed-off-by: Jee Jee Li <pandaleefree@gmail.com>	2024-11-26 22:54:12 -08:00
youkaichao	45ac4ff270	[bugfix] fix aria model and add torch.compile (#10645 ) Signed-off-by: youkaichao <youkaichao@gmail.com>	2024-11-25 18:32:09 -08:00
Cyrus Leung	cf73f0c95e	[Model] Enable optional prefix when loading embedding models (#10639 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>	2024-11-25 18:14:33 +00:00
youkaichao	eebad39f26	[torch.compile] support all attention backends (#10558 ) Signed-off-by: youkaichao <youkaichao@gmail.com>	2024-11-22 14:04:42 -08:00
Isotr0py	c4e464333e	[Misc] Add uninitialized params tracking for `AutoWeightsLoader` (#10327 ) Signed-off-by: Isotr0py <2037008807@qq.com>	2024-11-18 09:07:46 +08:00
Murali Andoorveedu	b2e0ad3b59	[Perf] Reduce peak memory usage of llama (#10339 ) Signed-off-by: andoorve <37849411+andoorve@users.noreply.github.com>	2024-11-15 00:38:20 +00:00
Woosuk Kwon	bbd3e86926	[V1] Support VLMs with fine-grained scheduling (#9871 ) Signed-off-by: Woosuk Kwon <woosuk.kwon@berkeley.edu> Co-authored-by: Roger Wang <ywang@roblox.com>	2024-11-13 04:53:13 +00:00
youkaichao	f89d18ff74	[6/N] pass whole config to inner model (#10205 ) Signed-off-by: youkaichao <youkaichao@gmail.com>	2024-11-11 06:41:46 +00:00
youkaichao	1a95f10ee7	[5/N] pass the whole config to model (#9983 ) Signed-off-by: youkaichao <youkaichao@gmail.com>	2024-11-09 14:17:28 +08:00
Joe Runde	d58268c56a	[V1] Make v1 more testable (#9888 ) Signed-off-by: Joe Runde <Joseph.Runde@ibm.com>	2024-11-06 11:57:35 -08:00
Jee Jee Li	2003cc3513	[Model][LoRA]LoRA support added for LlamaEmbeddingModel (#10071 ) Signed-off-by: Jee Jee Li <pandaleefree@gmail.com>	2024-11-06 09:49:19 +00:00
Aaron Pham	21063c11c7	[CI/Build] drop support for Python 3.8 EOL (#8464 ) Signed-off-by: Aaron Pham <contact@aarnphm.xyz>	2024-11-06 07:11:55 +00:00
Jee Jee Li	fb2716d641	[Misc]Reduce BNB static variable (#9987 ) Signed-off-by: Jee Jee Li <pandaleefree@gmail.com>	2024-11-04 17:04:40 +00:00
Went-Liang	81f09cfd80	[Model] Support math-shepherd-mistral-7b-prm model (#9697 ) Signed-off-by: Went-Liang <wenteng_liang@163.com>	2024-10-30 09:33:42 -07:00
Michael Goin	bc73e9821c	[Bugfix] Fix prefix strings for quantized VLMs (#9772 )	2024-10-29 16:02:59 -07:00
wangshuai09	4e2d95e372	[Hardware][ROCM] using current_platform.is_rocm (#9642 ) Signed-off-by: wangshuai09 <391746016@qq.com>	2024-10-28 04:07:00 +00:00
youkaichao	17c79f3c36	[torch.compile] auto infer dynamic_arg_dims from type annotation (#9589 )	2024-10-22 13:43:37 -07:00
Cyrus Leung	7abba39ee6	[Model] VLM2Vec, the first multimodal embedding model in vLLM (#9303 )	2024-10-16 14:31:00 +08:00
youkaichao	e00c094f15	[torch.compile] generic decorators (#9258 )	2024-10-10 15:54:23 -07:00
youkaichao	e4d652ea3e	[torch.compile] integration with compilation control (#9058 )	2024-10-10 12:39:36 -07:00
Isotr0py	07c11cf4d4	[Bugfix] Fix lm_head weights tying with lora for llama (#9227 )	2024-10-10 21:11:56 +08:00
Cyrus Leung	8bfaa4e31e	[Bugfix] fix composite weight loading and EAGLE weight loading (#9160 )	2024-10-09 00:36:55 -07:00
chenqianfzh	2f4117c38e	support bitsandbytes quantization with more models (#9148 )	2024-10-08 19:52:19 -06:00
Isotr0py	f19da64871	[Core] Refactor GGUF parameters packing and forwarding (#8859 )	2024-10-07 10:01:46 +00:00
Cyrus Leung	b22b798471	[Model] PP support for embedding models and update docs (#9090 ) Co-authored-by: Roger Wang <136131678+ywang96@users.noreply.github.com>	2024-10-06 16:35:27 +08:00
Murali Andoorveedu	0f6d7a9a34	[Models] Add remaining model PP support (#7168 ) Signed-off-by: Muralidhar Andoorveedu <muralidhar.andoorveedu@centml.ai> Signed-off-by: Murali Andoorveedu <muralidhar.andoorveedu@centml.ai> Co-authored-by: DarkLight1337 <tlleungac@connect.ust.hk>	2024-10-04 10:56:58 +08:00
Patrick von Platen	29f49cd6e3	[Model] Allow loading from original Mistral format (#8168 ) Co-authored-by: Michael Goin <michael@neuralmagic.com>	2024-09-06 17:02:05 -06:00
afeldman-nm	428dd1445e	[Core] Logprobs support in Multi-step (#7652 )	2024-08-29 19:19:08 -07:00
Cyrus Leung	7025b11d94	[Bugfix] Fix weight loading for Chameleon when TP>1 (#7410 )	2024-08-13 05:33:41 +00:00
Isotr0py	360bd67cf0	[Core] Support loading GGUF model (#5191 ) Co-authored-by: Michael Goin <michael@neuralmagic.com>	2024-08-05 17:54:23 -06:00
Travis Johnson	630dd9e0ae	[Bugfix][Model] Skip loading lm_head weights if using tie_word_embeddings (#6758 ) Signed-off-by: Travis Johnson <tsjohnso@us.ibm.com>	2024-07-31 19:49:11 -07:00

1 2 3 4

166 Commits