GPU scalars can't be used for Python indexing (requires sync). Compute padded_expert_offsets on CPU via .cpu().tolist() for the Python loop. This is OK for cudagraph: Python code only runs during capture, not replay. The GPU kernel launches recorded during capture are deterministic.