Logo
Explore Help
Register Sign In
biondizzle/vllm
1
0
Fork 0
You've already forked vllm
Code Issues Pull Requests Actions 2 Packages Projects Releases Wiki Activity
Files
183dad7a85487dbb351c43c11d2180a6108d5448
vllm/csrc/quantization/gptq_marlin
History
Jinzhen Lin d06ba4ed3f [Kernel] moe wna16 marlin kernel (#14447)
Signed-off-by: Jinzhen Lin <linjinzhen@hotmail.com>
Co-authored-by: Michael Goin <michael@neuralmagic.com>
Co-authored-by: mgoin <mgoin64@gmail.com>
2025-04-14 20:05:22 -07:00
..
awq_marlin_repack.cu
Fix CUDA kernel index data type in vllm/csrc/quantization/gptq_marlin/awq_marlin_repack.cu +10 (#15160)
2025-03-25 15:36:45 +08:00
gptq_marlin_repack.cu
Fix CUDA kernel index data type in vllm/csrc/quantization/gptq_marlin/awq_marlin_repack.cu +10 (#15160)
2025-03-25 15:36:45 +08:00
gptq_marlin.cu
[Bugfix] fix use_atomic_add support of marlin kernel when using v1 engine (#15946)
2025-04-05 20:04:22 -07:00
marlin_dtypes.cuh
[Kernel] moe wna16 marlin kernel (#14447)
2025-04-14 20:05:22 -07:00
marlin.cuh
[Kernel] moe wna16 marlin kernel (#14447)
2025-04-14 20:05:22 -07:00
Powered by Gitea Version: 1.25.2 Page: 71ms Template: 1ms
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API