Logo
Explore Help
Register Sign In
biondizzle/nvfp4-megamoe-kernel
1
0
Fork 0
You've already forked nvfp4-megamoe-kernel
Code Issues Pull Requests Actions Packages Projects Releases Wiki Activity
Files
a468f72a0eb65245655a774ecf7a2eec82f9d796
nvfp4-megamoe-kernel/dsv4/layers
History
biondizzle a468f72a0e CUDA graph: Pre-allocate L1 GEMM output buffers in MoE and SharedExpert
Pass out= parameter to run_fused_swiglu_grouped_gemm to avoid per-step
torch.zeros() allocation during CUDA graph capture.
2026-06-03 23:17:43 +00:00
..
__init__.py
Restructure: cutedsl/ -> dsv4/ with proper layering
2026-05-21 17:30:44 +00:00
grouped_linear.py
Fix grouped_linear GEMM output buffer shape and extraction
2026-06-03 22:26:40 +00:00
linear.py
CUDA graph: Fix _assemble_scales_single_group swizzle size
2026-06-03 17:02:34 +00:00
mhc.py
CUDA graph: Fix per-step allocations in decode loop
2026-06-03 16:38:35 +00:00
moe.py
CUDA graph: Pre-allocate L1 GEMM output buffers in MoE and SharedExpert
2026-06-03 23:17:43 +00:00
router.py
CRITICAL FIX: runtime activation global scale to prevent E4M3 overflow
2026-06-01 14:21:16 +00:00
shared_expert.py
CUDA graph: Pre-allocate L1 GEMM output buffers in MoE and SharedExpert
2026-06-03 23:17:43 +00:00
Powered by Gitea Version: 1.25.2 Page: 106ms Template: 2ms
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API