Replace NO-OP round-trip + normalize + epilogue_tma_store with: - get_tmem_load_op + get_smem_store_op paired atoms - One-way TMEM→reg (normalize) →SMEM→GMEM - Eliminates ~3% error from TMEM layout mismatch - O rescale disabled (single KV tile only for now) - Pre-computed TMA partitions outside if blocks
1.6 KiB
1.6 KiB