This website requires JavaScript.
Explore
Help
Register
Sign In
biondizzle
/
vllm
Watch
1
Star
0
Fork
0
You've already forked vllm
Code
Issues
Pull Requests
Actions
2
Packages
Projects
Releases
Wiki
Activity
Files
1bf2dd9df025feb82e27f90f534a3bf829ae75e9
vllm
/
tests
/
quantization
History
Li, Jiang
0b952af458
[Hardware][Intel] Support compressed-tensor W8A8 for CPU backend (
#7257
)
2024-09-11 09:46:46 -07:00
..
__init__.py
…
test_bitsandbytes.py
support bitsandbytes 8-bit and FP4 quantized models (
#7445
)
2024-08-29 19:09:08 -04:00
test_compressed_tensors.py
[Hardware][Intel] Support compressed-tensor W8A8 for CPU backend (
#7257
)
2024-09-11 09:46:46 -07:00
test_configs.py
…
test_cpu_offload.py
[ci][test] adjust max wait time for cpu offloading test (
#7709
)
2024-08-20 17:12:44 -07:00
test_experts_int8.py
[Kernel] W8A16 Int8 inside FusedMoE (
#7415
)
2024-08-16 10:06:51 -07:00
test_fp8.py
…
test_lm_head.py
…
utils.py
…