biondizzle
212fc85627
P6: One-way TMEM→regs→SMEM→TMA store epilogue
- fmha_6warp_multihead.cuh: Rewritten epilogue with proper Blackwell pipeline
1. TMEM → regs (tcgen05.ld, warp-collective)
2. epilogue_op in regs (normalize, FP4 hook via ENABLE_FP4_EPILOGUE)
3. regs → SMEM row-major (sO_epi, for TMA tile format)
4. TMA store SMEM → GMEM (async, enables multi-CTA)
Fallback to direct GMEM write when tma_o is nullptr.
Added FmhaParams.tma_o field and ENABLE_FP4_EPILOGUE template param.
- fmha_6warp_tma_multirow_multitile.cuh: Same epilogue pattern for multi-tile.
Writes normalized output to sO_epi_rowmajor + TMA store (or direct GMEM).
Added tma_o to FmhaTmaMultiRowMultiTileParams.
- fmha_tma.cuh: Added tma_store_2d and tma_store_wait for async GMEM writes.
- fmha_multihead_capi.cu: Added fmha_multihead_decode_tma_launch with
per-(head,batch) TMA descriptors. Updated SMEM size calculation for sO_epi + sMbarStore.
- fmha_multitile_capi.cu: Added tma_o=nullptr (backward compatible), updated SMEM size.
2026-05-30 16:56:07 +00:00
..
2026-05-30 16:56:07 +00:00
2026-05-22 00:08:38 +00:00
2026-05-21 17:30:44 +00:00
2026-05-28 16:17:47 +00:00
2026-05-21 17:30:44 +00:00
2026-05-28 04:59:01 +00:00
2026-05-28 16:17:47 +00:00
2026-05-21 22:04:20 +00:00
2026-05-21 17:30:44 +00:00