Logo
Explore Help
Register Sign In
biondizzle/nvfp4-megamoe-kernel
1
0
Fork 0
You've already forked nvfp4-megamoe-kernel
Code Issues Pull Requests Actions Packages Projects Releases Wiki Activity
1,930 Commits 1 Branch 17 Tags
ca661d32e8e94de838d550da52e934dd7f5c6312
Commit Graph

11 Commits

Author SHA1 Message Date
biondizzle
ff8c677486 fix: SMEM size for MMA test — account for both sQ0 and sK0 2026-05-28 23:06:07 +00:00
biondizzle
fee022a485 test: MMA→4-warp read using proven fmha_common+umma_desc infra 2026-05-28 23:05:29 +00:00
biondizzle
e1a708a187 test: try 16x256b.x1 with column step=4 (4 cols per read) 2026-05-28 23:03:51 +00:00
biondizzle
95003eced2 test: 16x256b.x1 loads with uint32_t regs, matching working pattern 2026-05-28 23:03:10 +00:00
biondizzle
fffb493b0e fix: 16x256b.x1 load syntax — single address operand 2026-05-28 23:02:23 +00:00
biondizzle
44dcd6e8d0 test: 16x256b.x1 multiple LOADS — do they crash like stores? 2026-05-28 23:02:03 +00:00
biondizzle
d54bce6a6d fix: correct SMEM size for MMA 4-warp test 2026-05-28 23:01:12 +00:00
biondizzle
be45e87891 test: MMA→4-warp TMEM read — do warps see different rows? 2026-05-28 23:00:27 +00:00
biondizzle
6b0d57074a test: TMEM cross-warp visibility with different sync strategies 2026-05-28 22:59:31 +00:00
biondizzle
77d190278e test: simpler TMEM 4-warp read — direct store+load 2026-05-28 22:58:48 +00:00
biondizzle
91b03bd6bd test: verify 4-warp TMEM read with 32x32b.x8 after MMA 2026-05-28 22:57:59 +00:00
Powered by Gitea Version: 1.25.2 Page: 103ms Template: 4ms
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API