perf: P0/P1/P2 — fused SwiGLU for MoE+SE, eliminate per-call gsa fill
P0: Enable fused SwiGLU for all MoE instances (moe._fused_swiglu = True).
Eliminates ~8 BF16 kernel launches per MoE per token (gate/up split,
SiLU, clamp, elementwise multiply → single fused kernel launch).
P1: Enable fused SwiGLU for shared expert (SE):
- Added set_fused_swiglu() method to Nvfp4SharedExpert
- Added _run_l1_fused() using run_fused_swiglu_grouped_gemm (1-group)
- Interleave L1 weights at finalize time for fused kernel compatibility
- Fused kernel handles SwiGLU + clamp in registers, outputs BF16
P2: Eliminate per-call _gsa_buf.fill_() in Nvfp4Linear:
- _activation_global_scale is set once at warmup, never changes after
- Skip redundant fill_() via _gsa_buf_initialized flag
- Saves 244 CPU→GPU scalar fills per token (4 linears × 61 layers)
P3: Deferred (in-kernel RoPE fusion — kernel-side change, not single_shot)