P6: One-way TMEM→regs→SMEM→TMA store epilogue
- fmha_6warp_multihead.cuh: Rewritten epilogue with proper Blackwell pipeline
1. TMEM → regs (tcgen05.ld, warp-collective)
2. epilogue_op in regs (normalize, FP4 hook via ENABLE_FP4_EPILOGUE)
3. regs → SMEM row-major (sO_epi, for TMA tile format)
4. TMA store SMEM → GMEM (async, enables multi-CTA)
Fallback to direct GMEM write when tma_o is nullptr.
Added FmhaParams.tma_o field and ENABLE_FP4_EPILOGUE template param.
- fmha_6warp_tma_multirow_multitile.cuh: Same epilogue pattern for multi-tile.
Writes normalized output to sO_epi_rowmajor + TMA store (or direct GMEM).
Added tma_o to FmhaTmaMultiRowMultiTileParams.
- fmha_tma.cuh: Added tma_store_2d and tma_store_wait for async GMEM writes.
- fmha_multihead_capi.cu: Added fmha_multihead_decode_tma_launch with
per-(head,batch) TMA descriptors. Updated SMEM size calculation for sO_epi + sMbarStore.
- fmha_multitile_capi.cu: Added tma_o=nullptr (backward compatible), updated SMEM size.