2.0 KiB
2.0 KiB
CURRENT_ISSUE.md — PV GEMM for Prefill
Status: HD=16 ✅ (cos 0.9997), HD=64 ✅ (cos 0.931, BF16 accumulation precision)
What works:
- HD=16: Full pipeline QK(SS) → softmax → PV(SS, 8 K-tiles) → epilogue. Cosine 0.9997.
- HD=64: Full pipeline with BLOCK_MN_B=64. Cosine 0.931 — V canonical layout verified correct, error is BF16 accumulation precision in PV MMA (kind::f16 uses BF16 dot products, reference uses FP32).
- Both require
cudaFuncSetAttributefor SMEM >48KB (HD=64 uses 52.9 KB).
Key architectural decisions:
- PV via SS MMA with SMEM-P (NOT TS MMA). TS MMA's A-fragment TMEM layout (Layout A) is incompatible with 32x32b stores. SS MMA with both operands in SMEM avoids this.
- Per-K-tile P fill into reusable (128,16) buffer from shared
s_p_vals[128]. The (128,128) canonical P with K-tile offsets has an accumulation bug. - V canonical layout:
g_mn=d/8, g_k=lr/8, llr=d%8, lc=lr%8where d=HD(MN), lr=seq_pos(K). The original code had MN/K swapped. - PV MMA scale: ~1.0 for both BLOCK_MN_B=16 and 64. QK MMA scale is also 1.0 (NOT 0.5 as assumed earlier).
- SMEM >48KB: Must opt in with
cudaFuncSetAttribute(kernel, cudaFuncAttributeMaxDynamicSharedMemorySize, smem).
Files:
test_fmha_v5.cu— HD=16 full pipeline (cos 0.9997) ✅test_fmha_hd64_smem_p.cu— HD=64 full pipeline (cos 0.931) 🚧test_pv_ss_128.cu— PV SS MMA with (128,128) P (shows accumulation bug)test_pv_ss_b64.cu— PV SS MMA with B=(64,16) BLOCK_MN=64 (works)test_ss_ts_sequence.cu— SS+TS coexistence testtest_pv_ss.cu— Minimal PV SS MMA (A=128×16, B=16×16)
Next steps:
- HD=64 precision: Investigate FP32-accumulation MMA variant (kind::f32) for PV
- HD=128, HD=256: Extend the pipeline (more QK/PV K-tiles, larger V)
- Prefill T>1: Fill all 128 rows of sPk from s_p_vals
- Multi-head: Per-head launch or head-packed M dimension
- Production kernel: Extract into
fmha_sm100.cuhwith proper warp specialization