Fix Blackwell: skip FlashMLA assertion + force CuTeDSL kernel

1. DeepseekV4MLAAttention.__init__ had a hard assertion that the
   attention backend MUST be FlashMLA. On Blackwell, FlashMLA doesn't
   work but we bypass it via _attention_impl_blackwell(). Added
   _is_blackwell flag to skip FlashMLA-specific init (fp8_ds_mla
   cache format conversion).

2. Added VLLM_NVFP4_GEMM_BACKEND=cutedsl env var to docker-compose.yml
   to force CuTeDSL kernel selection for NVFP4 linear layers.

3. Updated register_cutedsl_kernel.py to also register CuTeDSL in
   _NVFP4_BACKEND_TO_KERNEL dict (for the env var override path).
This commit is contained in:
2026-05-19 08:19:23 +00:00
parent 2856323360
commit e1a642452a
3 changed files with 15 additions and 1 deletions

View File

@@ -11,6 +11,7 @@ services:
- PYTHONUNBUFFERED=1
- VLLM_RPC_TIMEOUT_MS=600000
- CLAWMINE_DEBUG=1
- VLLM_NVFP4_GEMM_BACKEND=cutedsl
command:
- /model
- --trust-remote-code