Add _output_buf_padded for the flat GEMM output, pass as out= parameter to run_nvfp4_grouped_gemm to avoid per-step torch.zeros() allocation.
Add _output_buf_padded for the flat GEMM output, pass as out= parameter to run_nvfp4_grouped_gemm to avoid per-step torch.zeros() allocation.