This website requires JavaScript.
Explore
Help
Register
Sign In
biondizzle
/
nvfp4-megamoe-kernel
Watch
1
Star
0
Fork
0
You've already forked nvfp4-megamoe-kernel
Code
Issues
Pull Requests
Actions
Packages
Projects
Releases
Wiki
Activity
Files
851ec9b4d5fff23a604037f0a4820078cb8c6901
nvfp4-megamoe-kernel
/
dsv4
/
kernels
History
biondizzle
851ec9b4d5
P3 WIP: fused RMSNorm + quantize kernel skeleton (not yet integrated)
2026-06-02 09:02:52 +00:00
..
attention
perf: skip MQA GQA expansion in FMHA (stride=0, no 128x K/V copy)
2026-06-02 03:54:03 +00:00
cache
fix: correct gather.py kernel_dir path
2026-05-30 21:12:09 +00:00
compressor
fix: import torch.utils.cpp_extension explicitly in production_compress
2026-06-01 05:20:44 +00:00
cuda
P3 WIP: fused RMSNorm + quantize kernel skeleton (not yet integrated)
2026-06-02 09:02:52 +00:00
gemm
fix: use cute.where() directly for clamp in fused SwiGLU
2026-06-02 08:16:41 +00:00
indexer
P0 COMPLETE: Eliminate ALL .item() CPU-GPU syncs from NVFP4 activation path
2026-06-01 21:05:03 +00:00
router
Switch router to Nvfp4Linear production GEMM (custom CuTeDSL kernel crashes MLIR)
2026-06-01 11:17:54 +00:00
__init__.py
Restructure: cutedsl/ -> dsv4/ with proper layering
2026-05-21 17:30:44 +00:00