This website requires JavaScript.
Explore
Help
Register
Sign In
biondizzle
/
nvfp4-megamoe-kernel
Watch
1
Star
0
Fork
0
You've already forked nvfp4-megamoe-kernel
Code
Issues
Pull Requests
Actions
Packages
Projects
Releases
Wiki
Activity
Files
911a80e7217c585dd9d4e0ff591c596d37335768
nvfp4-megamoe-kernel
/
dsv4
History
biondizzle
911a80e721
D1.3: Fix tOrP0 for SMEM-P - skip make_tensor when offset is 0
...
CuTeDSL doesn't support OpResult + int. When offset is 0 (SMEM-P), just use tOrP directly.
2026-05-23 21:03:00 +00:00
..
cache
Flush compressor: schema fix, prepare_forward, flush_write kernels, state rotation
2026-05-22 00:25:47 +00:00
kernels
D1.3: Fix tOrP0 for SMEM-P - skip make_tensor when offset is 0
2026-05-23 21:03:00 +00:00
layers
Fix layer construction: match existing API signatures, add RMSNorm impl
2026-05-21 23:31:58 +00:00
loader
Restructure: cutedsl/ -> dsv4/ with proper layering
2026-05-21 17:30:44 +00:00
model
Fix layer construction: match existing API signatures, add RMSNorm impl
2026-05-21 23:31:58 +00:00
ops
fix: import ceil_div in quantize.py (was NameError at runtime)
2026-05-23 08:40:24 +00:00
reference
Restructure: cutedsl/ -> dsv4/ with proper layering
2026-05-21 17:30:44 +00:00
__init__.py
Restructure: cutedsl/ -> dsv4/ with proper layering
2026-05-21 17:30:44 +00:00