vllm/quantization at ccd4f129e8ad95191b3c8d6d0e935382b10c5164 - vllm - 丝路新云-代码仓

xinyun/vllm

mirror of https://git.datalinker.icu/vllm-project/vllm.git synced 2026-05-07 16:32:26 +08:00

History

Tyler Michael Smith ccd4f129e8

[Kernel] Add GPU architecture guards to the CUTLASS w8a8 kernels to reduce binary size (#5157 )

Co-authored-by: Cody Yu <hao.yu.cody@gmail.com>

2024-06-05 10:44:15 -07:00

..

[CI/Build] Enforce style for C++ and CUDA code with clang-format (#4722 )

2024-05-22 07:18:41 +00:00

[CI/Build] Enforce style for C++ and CUDA code with clang-format (#4722 )

2024-05-22 07:18:41 +00:00

compressed_tensors

[Kernel] Pass a device pointer into the quantize kernel for the scales (#5159 )

2024-06-03 09:52:30 -07:00

[Kernel] Add GPU architecture guards to the CUTLASS w8a8 kernels to reduce binary size (#5157 )

2024-06-05 10:44:15 -07:00

[CI/Build] Enforce style for C++ and CUDA code with clang-format (#4722 )

2024-05-22 07:18:41 +00:00

[CI/Build] Enforce style for C++ and CUDA code with clang-format (#4722 )

2024-05-22 07:18:41 +00:00

[CI/Build] Enforce style for C++ and CUDA code with clang-format (#4722 )

2024-05-22 07:18:41 +00:00

Revert "[Kernel] Marlin_24: Ensure the mma.sp instruction is using the ::ordered_metadata modifier (introduced with PTX 8.5)" (#5149 )

2024-05-30 22:00:26 -07:00

[CI/Build] Enforce style for C++ and CUDA code with clang-format (#4722 )

2024-05-22 07:18:41 +00:00