Packed E2M1 output has 2 elements per byte, so block_n elements = block_n/2 bytes. block_n/4 was under-sizing the TMA SMEM row by 2x → OOB write → LAUNCH_FAILED.