This website requires JavaScript.
Explore
Help
Register
Sign In
biondizzle
/
vllm
Watch
1
Star
0
Fork
0
You've already forked vllm
Code
Issues
Pull Requests
Actions
2
Packages
Projects
Releases
Wiki
Activity
Files
fa183e92713456dec682088a362dd9908100cc03
vllm
/
vllm
/
attention
History
Benjamin Chislett
304419576a
[Perf] Refactor cudagraph_support to enable full CUDA graphs for spec decoding with FlashInfer (
#28479
)
...
Signed-off-by: Benjamin Chislett <
bchislett@nvidia.com
>
2025-11-13 01:56:40 +09:00
..
backends
[CPU] Refactor CPU attention backend (
#27954
)
2025-11-12 09:43:06 +08:00
layers
[Perf] Refactor cudagraph_support to enable full CUDA graphs for spec decoding with FlashInfer (
#28479
)
2025-11-13 01:56:40 +09:00
ops
VLLM_USE_TRITON_FLASH_ATTN
V0 variable deprecation (
#27611
)
2025-11-11 18:34:36 -08:00
utils
[Misc] Refactor Attention kv transfer methods into decorator (
#27816
)
2025-11-12 16:05:44 +00:00
__init__.py
Convert formatting to use
ruff
instead of
yapf
+
isort
(
#26247
)
2025-10-05 07:06:22 -07:00
layer.py
[Misc] Refactor Attention kv transfer methods into decorator (
#27816
)
2025-11-12 16:05:44 +00:00
selector.py
[V0 deprecation] Deprecate use_v1 parameter (
#28112
)
2025-11-12 14:03:52 +00:00