vllm/csrc/quantization at e69ded7d1c8a4f6ed26e64090bdc050c06cde3b9 - vllm - Gitea: Git with a cup of tea

biondizzle/vllm

Files

History

Dipika Sikka ca3ea51bde [Kernel] Dynamic Per-Token Activation Quantization (#5037 )

Co-authored-by: Varun Sundar Rabindranath <varunsundar08@gmail.com>
Co-authored-by: Varun Sundar Rabindranath <varun@neuralmagic.com>

2024-06-07 09:36:26 -07:00

..

[CI/Build] Enforce style for C++ and CUDA code with clang-format (#4722 )

2024-05-22 07:18:41 +00:00

[CI/Build] Enforce style for C++ and CUDA code with clang-format (#4722 )

2024-05-22 07:18:41 +00:00

compressed_tensors

[Kernel] Dynamic Per-Token Activation Quantization (#5037 )

2024-06-07 09:36:26 -07:00

[Kernel] Add GPU architecture guards to the CUTLASS w8a8 kernels to reduce binary size (#5157 )

2024-06-05 10:44:15 -07:00

[CI/Build] Enforce style for C++ and CUDA code with clang-format (#4722 )

2024-05-22 07:18:41 +00:00

[CI/Build] Enforce style for C++ and CUDA code with clang-format (#4722 )

2024-05-22 07:18:41 +00:00

[CI/Build] Enforce style for C++ and CUDA code with clang-format (#4722 )

2024-05-22 07:18:41 +00:00

Revert "[Kernel] Marlin_24: Ensure the mma.sp instruction is using the ::ordered_metadata modifier (introduced with PTX 8.5)" (#5149 )

2024-05-30 22:00:26 -07:00

[CI/Build] Enforce style for C++ and CUDA code with clang-format (#4722 )

2024-05-22 07:18:41 +00:00