This website requires JavaScript.
Explore
Help
Register
Sign In
biondizzle
/
vllm
Watch
1
Star
0
Fork
0
You've already forked vllm
Code
Issues
Pull Requests
Actions
2
Packages
Projects
Releases
Wiki
Activity
Files
6edd43de3ce2aa9ca93b8ece656af7547526afd3
vllm
/
vllm
/
v1
/
core
History
jaime campos salas
891c60dcd5
fix(kv-cache): increase hybrid attention grouping threshold from 1.25 to 1.5 (
#36684
)
...
Signed-off-by: Jaime Campos Salas <
jaime.campos.salas@gmail.com
>
2026-03-12 23:28:27 -04:00
..
sched
[Refactor] Remove dead code in KV connector (
#36424
)
2026-03-11 19:40:17 +00:00
__init__.py
…
block_pool.py
[feat] Add per-block extra_keys to KV events (
#33304
)
2026-02-20 20:11:40 -08:00
encoder_cache_manager.py
[Refactor] Move profiling methods to MM budget (
#33559
)
2026-02-02 23:27:00 +08:00
kv_cache_coordinator.py
[BugFix] Avoid prefix cache hit in the same schedule step for mamba layers (
#29387
)
2026-02-10 07:41:16 +00:00
kv_cache_manager.py
[BUGFIX][Mamba][Qwen3.5] Zero freed SSM cache blocks on GPU (
#35219
)
2026-03-10 03:32:20 -07:00
kv_cache_metrics.py
[Core][Observability] Add KV cache residency metrics (
#27793
)
2025-12-01 18:27:53 +00:00
kv_cache_utils.py
fix(kv-cache): increase hybrid attention grouping threshold from 1.25 to 1.5 (
#36684
)
2026-03-12 23:28:27 -04:00
single_type_kv_cache_manager.py
[BUGFIX][Mamba][Qwen3.5] Zero freed SSM cache blocks on GPU (
#35219
)
2026-03-10 03:32:20 -07:00