This website requires JavaScript.
Explore
Help
Register
Sign In
biondizzle
/
vllm
Watch
1
Star
0
Fork
0
You've already forked vllm
Code
Issues
Pull Requests
Actions
2
Packages
Projects
Releases
Wiki
Activity
Files
88016c372a5962eb98f4dfc71243ccd64433710e
vllm
/
vllm
/
utils
History
Li, Jiang
88016c372a
[Bugfix] Fix pooling models on CPU backend (
#23392
)
...
Signed-off-by: jiang1.li <
jiang1.li@intel.com
>
2025-08-22 09:47:17 +00:00
..
__init__.py
[Bugfix] Fix pooling models on CPU backend (
#23392
)
2025-08-22 09:47:17 +00:00
deep_gemm.py
[Feature] Enable DeepGEMM Linear on B200; 1.5% E2E throughput improvement (
#23351
)
2025-08-21 21:01:08 -07:00
flashinfer.py
[NVIDIA] Support Flashinfer TRTLLM FP8-q/kv/out Attention Kernel (
#21716
)
2025-08-19 08:22:15 -04:00
jsontree.py
[Misc] Move jsontree to utils (
#22622
)
2025-08-11 03:49:32 -07:00
tensor_schema.py
[Bugfix] Fix MiniCPMV Image input inference failed (
#22813
)
2025-08-13 09:41:41 -07:00