This website requires JavaScript.
Explore
Help
Register
Sign In
biondizzle
/
nvfp4-megamoe-kernel
Watch
1
Star
0
Fork
0
You've already forked nvfp4-megamoe-kernel
Code
Issues
Pull Requests
Actions
Packages
Projects
Releases
Wiki
Activity
Files
44fb9b6c00fc3ccdd5c0500122e4754303bf4550
nvfp4-megamoe-kernel
/
dsv4
History
biondizzle
44fb9b6c00
Fix: pass self.mma_tiler_mnk (full K) to _compute_stages, not self.mma_tiler (K=1 placeholder)
2026-06-01 10:55:43 +00:00
..
cache
E1: Wire LayerCacheHandle gather methods + CUDA gather kernels
2026-05-30 21:09:21 +00:00
kernels
Fix: pass self.mma_tiler_mnk (full K) to _compute_stages, not self.mma_tiler (K=1 placeholder)
2026-06-01 10:55:43 +00:00
layers
Wire NVFP4 fused router kernel into e2e single-shot pipeline
2026-06-01 09:47:48 +00:00
loader
Restructure: cutedsl/ -> dsv4/ with proper layering
2026-05-21 17:30:44 +00:00
model
Fix remaining mHC API references: layer_compare.py, layer.py comment
2026-05-31 18:38:34 +00:00
ops
fix: clamp block_amax to E4M3 max (448) in quantize_activation_nvfp4 — prevents NaN from overflow
2026-06-01 04:59:06 +00:00
reference
Restructure: cutedsl/ -> dsv4/ with proper layering
2026-05-21 17:30:44 +00:00
__init__.py
Restructure: cutedsl/ -> dsv4/ with proper layering
2026-05-21 17:30:44 +00:00