vllm/quantization at a360ff80bb34f9dfcd21cf880c2030daa2d6b3a3 - vllm - 丝路新云-代码仓

xinyun/vllm

mirror of https://git.datalinker.icu/vllm-project/vllm.git synced 2026-07-17 22:47:22 +08:00

History

Tyler Michael Smith 1197e02141

[Build] Guard against older CUDA versions when building CUTLASS 3.x kernels (#5168 )

2024-05-31 17:21:38 -07:00

..

[CI/Build] Enforce style for C++ and CUDA code with clang-format (#4722 )

2024-05-22 07:18:41 +00:00

[CI/Build] Enforce style for C++ and CUDA code with clang-format (#4722 )

2024-05-22 07:18:41 +00:00

compressed_tensors

[Kernel] Initial Activation Quantization Support (#4525 )

2024-05-23 21:29:18 +00:00

[Build] Guard against older CUDA versions when building CUTLASS 3.x kernels (#5168 )

2024-05-31 17:21:38 -07:00

[CI/Build] Enforce style for C++ and CUDA code with clang-format (#4722 )

2024-05-22 07:18:41 +00:00

[CI/Build] Enforce style for C++ and CUDA code with clang-format (#4722 )

2024-05-22 07:18:41 +00:00

[CI/Build] Enforce style for C++ and CUDA code with clang-format (#4722 )

2024-05-22 07:18:41 +00:00

Revert "[Kernel] Marlin_24: Ensure the mma.sp instruction is using the ::ordered_metadata modifier (introduced with PTX 8.5)" (#5149 )

2024-05-30 22:00:26 -07:00

[CI/Build] Enforce style for C++ and CUDA code with clang-format (#4722 )

2024-05-22 07:18:41 +00:00