- test_fmha_v3_per_row.py: Mike's per-row patch with deadlock fix (moved C6 O-rescale after softmax_done_bar, fixed pv_done_bar for kt=0) Still GPU hangs — needs further debugging - test_fmha_v3_fixed_v.py: s_k parameter + acc_pipe consumer fix Same cosine as original (V TMA handles data shape correctly) - Baseline: n=128→0.993, n=256→0.725, n=384→0.620 Key insight: QK TMEM load fragment has 4 rows × 32 cols per thread. Fragment-level row_max/row_sum is wrong for per-row operations. Per-row tracking (4 separate row_max/row_sum per thread) is needed.
32 KiB
32 KiB