Apply to_blocked swizzle on entire padded buffer at once instead of per-expert loops. No .item()/.cpu() calls. Fully cudagraph-safe.
Apply to_blocked swizzle on entire padded buffer at once instead of per-expert loops. No .item()/.cpu() calls. Fully cudagraph-safe.