|
|
e58980f80e
|
fix: increase test timeout for TMEM kernel
|
2026-05-28 06:41:59 +00:00 |
|
|
|
99b35eb2de
|
test: standalone CUDA test for FMHA SM100 (no PyTorch needed)
|
2026-05-28 05:31:03 +00:00 |
|
|
|
97df02ea07
|
fix: -Xcompiler -fPIC for nvcc shared library
|
2026-05-28 05:22:15 +00:00 |
|
|
|
4dfb71bc20
|
test: nvcc direct compilation test (avoid torch JIT __bf16 ICE)
|
2026-05-28 05:21:41 +00:00 |
|
|
|
09dfd4a41f
|
fix: rename .cpp to .cu for CUDA compilation
|
2026-05-28 05:16:41 +00:00 |
|
|
|
4c194b7254
|
fix: add CUDA include path for host compiler
|
2026-05-28 05:15:48 +00:00 |
|
|
|
f0660d0bd7
|
fix: use C++20 for cuda_bf16.h compat
|
2026-05-28 05:13:18 +00:00 |
|
|
|
6bd3356582
|
fix: include cuda_bf16.h unconditionally, add --expt-relaxed-constexpr
|
2026-05-28 05:13:01 +00:00 |
|
|
|
3eb432d064
|
fix: CUTLASS path /root/cutlass
|
2026-05-28 05:06:48 +00:00 |
|
|
|
66d9f5c60f
|
fix: --x cu for .cuh compilation
|
2026-05-28 05:06:13 +00:00 |
|
|
|
4dcd80ea0d
|
fix: use full nvcc path
|
2026-05-28 05:05:55 +00:00 |
|
|
|
fac7275f2b
|
test: nvcc compilation test for FMHA SM100 kernel
|
2026-05-28 05:05:31 +00:00 |
|
|
|
230c350c77
|
FMHA SM100: Raw CUDA C++ decode kernel — initial skeleton
6-warp specialization using CUTLASS C++ atoms directly:
- tcgen05.mma for QK (SMEM→SMEM→TMEM) and PV (TMEM→SMEM→TMEM)
- TMEM accumulator with one-way correction epilogue (TMEM→regs→SMEM→GMEM)
- In-kernel O rescale via registers (fixes D1.5 TMEM round-trip!)
- D3/D4/D5c masks, NVFP4 quantize helpers, FP8 E4M3 encode
- PyTorch binding with head_dim template dispatch
This bypasses all CuTeDSL limitations: float→int, TMEM round-trip,
multi-CTA, hd=512 MLIR compilation hang.
|
2026-05-28 05:04:44 +00:00 |
|