This website requires JavaScript.
Explore
Help
Register
Sign In
biondizzle
/
vllm
Watch
1
Star
0
Fork
0
You've already forked vllm
Code
Issues
Pull Requests
Actions
2
Packages
Projects
Releases
Wiki
Activity
Files
44a6528028ad79951de08b6a7928f6c05788d00d
vllm
/
tests
/
v1
/
spec_decode
History
Giancarlo Delfin
c32e97602d
[Model Runner V2] Enable forcing a specific acceptance rate during rejection sampling (
#38045
)
...
Signed-off-by: Giancarlo Delfin <
gdelfin@inferact.ai
>
2026-03-26 13:38:12 -07:00
..
__init__.py
…
test_acceptance_length.py
[Hardware] Replace torch.cuda.device_count/current_device/set_device API (
#36145
)
2026-03-12 07:57:47 -07:00
test_eagle_step_kernel.py
feat(spec_decode): fuse EAGLE step slot mapping and metadata updates (
#33503
)
2026-03-11 04:35:33 +00:00
test_eagle.py
[Async][Spec Decoding] Zero-bubble async scheduling + spec decoding (
#32951
)
2026-03-23 15:37:22 -04:00
test_extract_hidden_states.py
[Async][Spec Decoding] Zero-bubble async scheduling + spec decoding (
#32951
)
2026-03-23 15:37:22 -04:00
test_max_len.py
…
test_mtp.py
[BugFix] Add support for MTP num_speculative_tokens > 1 with sparse MLA (
#34552
)
2026-03-03 07:21:57 -08:00
test_ngram.py
…
test_speculators_eagle3.py
…
test_synthetic_rejection_sampler_utils.py
[Model Runner V2] Enable forcing a specific acceptance rate during rejection sampling (
#38045
)
2026-03-26 13:38:12 -07:00
test_tree_attention.py
[ROCm][CI] Fix spec decode logprobs flakiness and parametrize tree attention backends (
#34599
)
2026-02-20 20:25:23 -08:00