The CuTeDSL MLIR optimizer crashes (SIGABRT/core dump) on the combination of exp+log+sqrt in a for-range loop. The kernel now writes raw FP32 logits (with gsa*gsb applied) and sqrt(softplus) is done in PyTorch post-kernel. The GEMM is still pure NVFP4 Blackwell tensor cores.