Chenheli Hua 7f2ea7074e
[Frontend][Multimodal] Allow skipping media data when UUIDs are provided. (#23950)
Signed-off-by: Roger Wang <hey@rogerw.io>
Signed-off-by: Chenheli Hua <huachenheli@outlook.com>
Signed-off-by: Roger Wang <hey@rogerw.me>
Co-authored-by: Roger Wang <hey@rogerw.io>
Co-authored-by: Roger Wang <hey@rogerw.me>
2025-09-13 02:16:06 +00:00
..

Features

Compatibility Matrix

The tables below show mutually exclusive features and the support on some hardware.

The symbols used have the following meanings:

  • ✅ = Full compatibility
  • 🟠 = Partial compatibility
  • ❌ = No compatibility
  • ❔ = Unknown or TBD

!!! note Check the ❌ or 🟠 with links to see tracking issue for unsupported feature/hardware combination.

Feature x Feature

Feature [CP][chunked-prefill] APC LoRA SD CUDA graph pooling enc-dec logP prmpt logP async output multi-step mm best-of beam-search
[CP][chunked-prefill] ✅
APC ✅ ✅
LoRA ✅ ✅ ✅
SD ✅ ✅ ❌ ✅
CUDA graph ✅ ✅ ✅ ✅ ✅
pooling 🟠* 🟠* ✅ ❌ ✅ ✅
enc-dec ❌ ❌ ❌ ❌ ✅ ✅ ✅
logP ✅ ✅ ✅ ✅ ✅ ❌ ✅ ✅
prmpt logP ✅ ✅ ✅ ✅ ✅ ❌ ✅ ✅ ✅
async output ✅ ✅ ✅ ❌ ✅ ❌ ❌ ✅ ✅ ✅
multi-step ❌ ✅ ❌ ❌ ✅ ❌ ❌ ✅ ✅ ✅ ✅
mm ✅ ✅ 🟠^ ❔ ✅ ✅ ✅ ✅ ✅ ✅ ❔ ✅
best-of ✅ ✅ ✅ ❌ ✅ ❌ ✅ ✅ ✅ ❔ ❌ ✅ ✅
beam-search ✅ ✅ ✅ ❌ ✅ ❌ ✅ ✅ ✅ ❔ ❌ ❔ ✅ ✅

* Chunked prefill and prefix caching are only applicable to last-token pooling.
^ LoRA is only applicable to the language backbone of multimodal models.

{ #feature-x-hardware }

Feature x Hardware

Feature Volta Turing Ampere Ada Hopper CPU AMD TPU
[CP][chunked-prefill] ❌ ✅ ✅ ✅ ✅ ✅ ✅ ✅
APC ❌ ✅ ✅ ✅ ✅ ✅ ✅ ✅
LoRA ✅ ✅ ✅ ✅ ✅ ✅ ✅ ✅
SD ✅ ✅ ✅ ✅ ✅ ✅ ✅ ❌
CUDA graph ✅ ✅ ✅ ✅ ✅ ❌ ✅ ❌
pooling ✅ ✅ ✅ ✅ ✅ ✅ ✅ ❌
enc-dec ✅ ✅ ✅ ✅ ✅ ✅ ❌ ❌
mm ✅ ✅ ✅ ✅ ✅ ✅ ✅ ❌
logP ✅ ✅ ✅ ✅ ✅ ✅ ✅ ❌
prmpt logP ✅ ✅ ✅ ✅ ✅ ✅ ✅ ❌
async output ✅ ✅ ✅ ✅ ✅ ❌ ❌ ❌
multi-step ✅ ✅ ✅ ✅ ✅ ❌ ✅ ❌
best-of ✅ ✅ ✅ ✅ ✅ ✅ ✅ ❌
beam-search ✅ ✅ ✅ ✅ ✅ ✅ ✅ ❌

!!! note Please refer to [Feature support through NxD Inference backend][feature-support-through-nxd-inference-backend] for features supported on AWS Neuron hardware