- Remove noop + normalize TMEM round-trips (3% error per trip) - Use epilogue_tmem_copy_and_partition for TMEM→reg (paired atoms) - Use epilogue_smem_copy_and_partition for reg→SMEM (paired atoms) - Apply 1/row_sum normalization in register space (exact) - TMA store from SMEM→GMEM (no TMEM write-back) - Add iter_acc_early_release_in_epilogue attribute - Update SMEM-P comments to reflect coordinate-indexed fallback