mirror of
https://git.datalinker.icu/vllm-project/vllm.git
synced 2026-07-20 18:07:15 +08:00
[Docs] Move quant supported hardware table to README (#23663)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
This commit is contained in:
parent
2f13319f47
commit
6421b66bf4
@ -4,7 +4,6 @@ Quantization trades off model precision for smaller memory footprint, allowing l
|
|||||||
|
|
||||||
Contents:
|
Contents:
|
||||||
|
|
||||||
- [Supported Hardware](supported_hardware.md)
|
|
||||||
- [AutoAWQ](auto_awq.md)
|
- [AutoAWQ](auto_awq.md)
|
||||||
- [AutoRound](auto_round.md)
|
- [AutoRound](auto_round.md)
|
||||||
- [BitsAndBytes](bnb.md)
|
- [BitsAndBytes](bnb.md)
|
||||||
@ -19,3 +18,50 @@ Contents:
|
|||||||
- [AMD Quark](quark.md)
|
- [AMD Quark](quark.md)
|
||||||
- [Quantized KV Cache](quantized_kvcache.md)
|
- [Quantized KV Cache](quantized_kvcache.md)
|
||||||
- [TorchAO](torchao.md)
|
- [TorchAO](torchao.md)
|
||||||
|
|
||||||
|
## Supported Hardware
|
||||||
|
|
||||||
|
The table below shows the compatibility of various quantization implementations with different hardware platforms in vLLM:
|
||||||
|
|
||||||
|
<style>
|
||||||
|
td:not(:first-child) {
|
||||||
|
text-align: center !important;
|
||||||
|
}
|
||||||
|
td {
|
||||||
|
padding: 0.5rem !important;
|
||||||
|
white-space: nowrap;
|
||||||
|
}
|
||||||
|
|
||||||
|
th {
|
||||||
|
padding: 0.5rem !important;
|
||||||
|
min-width: 0 !important;
|
||||||
|
}
|
||||||
|
|
||||||
|
th:not(:first-child) {
|
||||||
|
writing-mode: vertical-lr;
|
||||||
|
transform: rotate(180deg)
|
||||||
|
}
|
||||||
|
</style>
|
||||||
|
|
||||||
|
| Implementation | Volta | Turing | Ampere | Ada | Hopper | AMD GPU | Intel GPU | Intel Gaudi | x86 CPU | AWS Neuron | Google TPU |
|
||||||
|
|-----------------------|---------|----------|----------|-------|----------|-----------|-------------|-------------|-----------|--------------|--------------|
|
||||||
|
| AWQ | ❌ | ✅︎ | ✅︎ | ✅︎ | ✅︎ | ❌ | ✅︎ | ❌ | ✅︎ | ❌ | ❌ |
|
||||||
|
| GPTQ | ✅︎ | ✅︎ | ✅︎ | ✅︎ | ✅︎ | ❌ | ✅︎ | ❌ | ✅︎ | ❌ | ❌ |
|
||||||
|
| Marlin (GPTQ/AWQ/FP8) | ❌ | ❌ | ✅︎ | ✅︎ | ✅︎ | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ |
|
||||||
|
| INT8 (W8A8) | ❌ | ✅︎ | ✅︎ | ✅︎ | ✅︎ | ❌ | ❌ | ❌ | ✅︎ | ✅︎ | ✅︎ |
|
||||||
|
| FP8 (W8A8) | ❌ | ❌ | ❌ | ✅︎ | ✅︎ | ✅︎ | ❌ | ❌ | ❌ | ✅︎ | ❌ |
|
||||||
|
| BitBLAS | ✅︎ | ✅ | ✅︎ | ✅︎ | ✅︎ | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ |
|
||||||
|
| BitBLAS (GPTQ) | ❌ | ❌ | ✅︎ | ✅︎ | ✅︎ | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ |
|
||||||
|
| bitsandbytes | ✅︎ | ✅︎ | ✅︎ | ✅︎ | ✅︎ | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ |
|
||||||
|
| DeepSpeedFP | ✅︎ | ✅︎ | ✅︎ | ✅︎ | ✅︎ | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ |
|
||||||
|
| GGUF | ✅︎ | ✅︎ | ✅︎ | ✅︎ | ✅︎ | ✅︎ | ❌ | ❌ | ❌ | ❌ | ❌ |
|
||||||
|
| INC (W8A8) | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | ✅︎ | ❌ | ❌ | ❌ |
|
||||||
|
|
||||||
|
- Volta refers to SM 7.0, Turing to SM 7.5, Ampere to SM 8.0/8.6, Ada to SM 8.9, and Hopper to SM 9.0.
|
||||||
|
- ✅︎ indicates that the quantization method is supported on the specified hardware.
|
||||||
|
- ❌ indicates that the quantization method is not supported on the specified hardware.
|
||||||
|
|
||||||
|
!!! note
|
||||||
|
This compatibility chart is subject to change as vLLM continues to evolve and expand its support for different hardware platforms and quantization methods.
|
||||||
|
|
||||||
|
For the most up-to-date information on hardware support and quantization methods, please refer to <gh-dir:vllm/model_executor/layers/quantization> or consult with the vLLM development team.
|
||||||
|
|||||||
@ -5,7 +5,7 @@ vLLM now supports [BitBLAS](https://github.com/microsoft/BitBLAS) for more effic
|
|||||||
!!! note
|
!!! note
|
||||||
Ensure your hardware supports the selected `dtype` (`torch.bfloat16` or `torch.float16`).
|
Ensure your hardware supports the selected `dtype` (`torch.bfloat16` or `torch.float16`).
|
||||||
Most recent NVIDIA GPUs support `float16`, while `bfloat16` is more common on newer architectures like Ampere or Hopper.
|
Most recent NVIDIA GPUs support `float16`, while `bfloat16` is more common on newer architectures like Ampere or Hopper.
|
||||||
For details see [supported hardware](supported_hardware.md).
|
For details see [supported hardware](README.md#supported-hardware).
|
||||||
|
|
||||||
Below are the steps to utilize BitBLAS with vLLM.
|
Below are the steps to utilize BitBLAS with vLLM.
|
||||||
|
|
||||||
|
|||||||
@ -1,32 +0,0 @@
|
|||||||
# Supported Hardware
|
|
||||||
|
|
||||||
The table below shows the compatibility of various quantization implementations with different hardware platforms in vLLM:
|
|
||||||
|
|
||||||
<style>
|
|
||||||
th {
|
|
||||||
white-space: nowrap;
|
|
||||||
min-width: 0 !important;
|
|
||||||
}
|
|
||||||
</style>
|
|
||||||
|
|
||||||
| Implementation | Volta | Turing | Ampere | Ada | Hopper | AMD GPU | Intel GPU | Intel Gaudi | x86 CPU | AWS Neuron | Google TPU |
|
|
||||||
|-----------------------|---------|----------|----------|-------|----------|-----------|-------------|-------------|-----------|--------------|--------------|
|
|
||||||
| AWQ | ❌ | ✅︎ | ✅︎ | ✅︎ | ✅︎ | ❌ | ✅︎ | ❌ | ✅︎ | ❌ | ❌ |
|
|
||||||
| GPTQ | ✅︎ | ✅︎ | ✅︎ | ✅︎ | ✅︎ | ❌ | ✅︎ | ❌ | ✅︎ | ❌ | ❌ |
|
|
||||||
| Marlin (GPTQ/AWQ/FP8) | ❌ | ❌ | ✅︎ | ✅︎ | ✅︎ | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ |
|
|
||||||
| INT8 (W8A8) | ❌ | ✅︎ | ✅︎ | ✅︎ | ✅︎ | ❌ | ❌ | ❌ | ✅︎ | ✅︎ | ✅︎ |
|
|
||||||
| FP8 (W8A8) | ❌ | ❌ | ❌ | ✅︎ | ✅︎ | ✅︎ | ❌ | ❌ | ❌ | ✅︎ | ❌ |
|
|
||||||
| BitBLAS (GPTQ) | ✅︎ | ✅︎ | ✅︎ | ✅︎ | ✅︎ | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ |
|
|
||||||
| bitsandbytes | ✅︎ | ✅︎ | ✅︎ | ✅︎ | ✅︎ | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ |
|
|
||||||
| DeepSpeedFP | ✅︎ | ✅︎ | ✅︎ | ✅︎ | ✅︎ | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ |
|
|
||||||
| GGUF | ✅︎ | ✅︎ | ✅︎ | ✅︎ | ✅︎ | ✅︎ | ❌ | ❌ | ❌ | ❌ | ❌ |
|
|
||||||
| INC (W8A8) | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | ✅︎ | ❌ | ❌ | ❌ |
|
|
||||||
|
|
||||||
- Volta refers to SM 7.0, Turing to SM 7.5, Ampere to SM 8.0/8.6, Ada to SM 8.9, and Hopper to SM 9.0.
|
|
||||||
- ✅︎ indicates that the quantization method is supported on the specified hardware.
|
|
||||||
- ❌ indicates that the quantization method is not supported on the specified hardware.
|
|
||||||
|
|
||||||
!!! note
|
|
||||||
This compatibility chart is subject to change as vLLM continues to evolve and expand its support for different hardware platforms and quantization methods.
|
|
||||||
|
|
||||||
For the most up-to-date information on hardware support and quantization methods, please refer to <gh-dir:vllm/model_executor/layers/quantization> or consult with the vLLM development team.
|
|
||||||
Loading…
x
Reference in New Issue
Block a user