docs: clarify L1 interleave removal — transpose is still needed
This commit is contained in:
@@ -272,10 +272,10 @@ Additionally, the dest buffer must be zero-initialized before remap because CUTL
|
||||
**Bug:** `logical_widths=[3072, 3072]` caused the function to apply expert 0's scale to gate half and expert 1's scale to up half of ALL experts. All other experts' global scales were discarded.
|
||||
**Fix:** Removed the `logical_widths` branch entirely. The `else` branch correctly broadcasts each expert's own `(E, 1)` global scale across `(E, N, K//16)`.
|
||||
|
||||
### 4. L1 weight interleave removed
|
||||
### 4. L1 weight interleave removed (transpose still needed)
|
||||
**File:** `weight_transform.py`
|
||||
**Bug:** `_interleave_l1_weights` assumed gate/up were pre-interleaved in groups of 16 and that the kernel used 2CTA UMMA layout. vLLM uses plain concat `[gate; up]` along the output dim, and our CUTLASS kernel uses `ClusterShape<1, 1, 1>`.
|
||||
**Fix:** Removed entirely. `l1_weight_out = l1_weight.contiguous()`.
|
||||
**Fix:** Removed the interleave function. Weights still need a transpose from checkpoint layout `(N, K_half)` row-major to CUTLASS layout `(K_half, N)` column-major — this is standard row→column conversion, not interleaving. Both L1 and L2 weights and scales are transposed.
|
||||
|
||||
### 5. SF remap: idx2crd+flatten coordinate extraction
|
||||
**File:** `cutlass_nvfp4_gemm.cu`
|
||||
|
||||
Reference in New Issue
Block a user