The warmup was running for every MoE layer (61 layers × 8 ranks = 488 compile attempts). The kernel is cached after the first compile — subsequent calls are instant. But the print spam was insane. Now uses a class-level flag to compile exactly once per process. All 61 layers on a rank share the same compiled kernel.