- One-way: TMEM → registers (normalize) → SMEM → GMEM - Eliminates TMEM round-trip error for O normalization - O rescale (kt>0) still uses old atoms (fix later) - Based on CUTLASS FMHA reference's correction_epilog pattern
- One-way: TMEM → registers (normalize) → SMEM → GMEM - Eliminates TMEM round-trip error for O normalization - O rescale (kt>0) still uses old atoms (fix later) - Based on CUTLASS FMHA reference's correction_epilog pattern