Logo
Explore Help
Register Sign In
biondizzle/vllm
1
0
Fork 0
You've already forked vllm
Code Issues Pull Requests Actions 2 Packages Projects Releases Wiki Activity
Files
ca90f503041f17ace06a288f06ebfc455402e3bc
vllm/csrc/moe
History
xiangze-arm f32cbc9a0c [CPU]Improve dynamic 4bit moe performance (#27240)
Signed-off-by: Zhang Xiangze <Xiangze.Zhang@arm.com>
2025-11-04 06:33:23 +00:00
..
marlin_moe_wna16
Convert formatting to use ruff instead of yapf + isort (#26247)
2025-10-05 07:06:22 -07:00
permute_unpermute_kernels
Fix CUDA permute/unpermute for use with DeepGemm Moe (#17934)
2025-07-27 07:08:00 -07:00
dynamic_4bit_int_moe_cpu.cpp
[CPU]Improve dynamic 4bit moe performance (#27240)
2025-11-04 06:33:23 +00:00
grouped_topk_kernels.cu
Use macro guard CUDA functions for back compatibility in grouped_topk_kernel.cu (#25346)
2025-09-23 09:45:39 -07:00
moe_align_sum_kernels.cu
[GPTOSS][DP/EP][Marlin] Enable GPTOSS Batched DP/EP using Marlin kernels (#25997)
2025-10-16 12:53:11 -07:00
moe_lora_align_sum_kernels.cu
Early exit for MoE LoRA kernels (#27131)
2025-11-03 20:22:17 +08:00
moe_ops.h
Early exit for MoE LoRA kernels (#27131)
2025-11-03 20:22:17 +08:00
moe_permute_unpermute_op.cu
[Kernel] CUTLASS MoE FP8: Integrate cuda moe permute/unpermute (#23045)
2025-08-20 10:35:26 -04:00
moe_wna16_utils.h
pre-commit autoupdate (#17380)
2025-04-29 06:46:55 -07:00
moe_wna16.cu
…
topk_softmax_kernels.cu
[Kernel][Performance] Fuse float cast and renormalize to topk softmax kernel (#26717)
2025-10-17 07:30:35 +00:00
torch_bindings.cpp
Early exit for MoE LoRA kernels (#27131)
2025-11-03 20:22:17 +08:00
Powered by Gitea Version: 1.25.2 Page: 5623ms Template: 12ms
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API