vllm/quantization at 8e192ff967b44b186ea02d30e49fddf656fdfe50 - vllm - 丝路新云-代码仓

xinyun/vllm

mirror of https://git.datalinker.icu/vllm-project/vllm.git synced 2026-06-20 10:47:17 +08:00

History

Dipika Sikka a1242324c9

[Kernel] Initial Activation Quantization Support (#4525 )

Co-authored-by: Varun Sundar Rabindranath <varunsundar08@gmail.com>
Co-authored-by: Varun Sundar Rabindranath <varun@neuralmagic.com>

2024-05-23 21:29:18 +00:00

..

[CI/Build] Enforce style for C++ and CUDA code with clang-format (#4722 )

2024-05-22 07:18:41 +00:00

[CI/Build] Enforce style for C++ and CUDA code with clang-format (#4722 )

2024-05-22 07:18:41 +00:00

compressed_tensors

[Kernel] Initial Activation Quantization Support (#4525 )

2024-05-23 21:29:18 +00:00

[Kernel] Fixup for CUTLASS kernels in CUDA graphs (#4954 )

2024-05-22 14:10:43 +00:00

[CI/Build] Enforce style for C++ and CUDA code with clang-format (#4722 )

2024-05-22 07:18:41 +00:00

[CI/Build] Enforce style for C++ and CUDA code with clang-format (#4722 )

2024-05-22 07:18:41 +00:00

[CI/Build] Enforce style for C++ and CUDA code with clang-format (#4722 )

2024-05-22 07:18:41 +00:00

Marlin 24 prefill performance improvement (about 25% better on average) (#4983 )

2024-05-23 02:39:27 -04:00

[CI/Build] Enforce style for C++ and CUDA code with clang-format (#4722 )

2024-05-22 07:18:41 +00:00