Logo
Explore Help
Register Sign In
biondizzle/vllm
1
0
Fork 0
You've already forked vllm
Code Issues Pull Requests Actions 2 Packages Projects Releases Wiki Activity
Files
927975ead8ea8f2c818844680fd7b29d0f62d0e9
vllm/vllm/model_executor
History
Andrey Talman 2111997f96 [release 2.11] Update to torch 2.11 (#34644)
2026-04-07 18:55:48 -07:00
..
kernels
[ROCm][Quantization] Add asymmetric INT8 quantization support to TritonInt8ScaledMMLinearKernel (#38501)
2026-04-06 09:42:10 +08:00
layers
[release 2.11] Update to torch 2.11 (#34644)
2026-04-07 18:55:48 -07:00
model_loader
[Frontend] new online quantization frontend (#38138)
2026-04-03 11:58:39 -04:00
models
[Attention][V0 Deprecation] Deprecate accept output buffer (#39125)
2026-04-07 17:14:58 -04:00
offloader
Bugfix for offloading+prefetch for GLM-4.7-FP8 (#37178)
2026-03-17 21:22:09 +08:00
warmup
[Perf] Change Trtllm fp8 MoE to use Shuffled Weights and BlockMajorK Layout (#38993)
2026-04-05 10:54:31 -04:00
__init__.py
[Platform] Deprecate seed_everything (#31659)
2026-01-04 18:34:04 -08:00
custom_op.py
Add ability to replace oot ops when using lora (#37181)
2026-03-16 18:04:15 -07:00
parameter.py
[Mypy] Fix mypy for vllm/model_executor (except vllm/model_executor/layers) (#37904)
2026-03-24 17:14:01 +00:00
utils.py
[BugFix] Fix EPLB fail for MoeFP4 model with Marlin backend (#33262)
2026-01-29 16:52:11 +08:00
Powered by Gitea Version: 1.25.2 Page: 342ms Template: 3ms
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API