vllm/kernels at 2fc5d6e0d7596dd93dbf4e1ca776f17449bb2143 - vllm

mirror of https://git.datalinker.icu/vllm-project/vllm.git synced 2026-08-04 00:57:03 +08:00

History

Varun Sundar Rabindranath 19bee6d12d

[Performance][DP/EP] Add silu_mul_per_token_group_quant_fp8_colmajor kernel (#29470 )

Signed-off-by: Varun Sundar Rabindranath <vsundarr@redhat.com>
Co-authored-by: Varun Sundar Rabindranath <vsundarr@redhat.com>
Co-authored-by: Tyler Michael Smith <tyler@neuralmagic.com>

2025-12-03 18:04:59 +00:00

attention

[Misc] Remove redundant attention var constants (#29650 )

2025-11-28 04:35:19 -08:00

core

Update rope_scaling to rope_parameters in preparation for Transformers v5 (#28542 )

2025-11-19 09:06:36 -08:00

mamba

[V1] [Hybrid] Mamba1 Automatic Prefix Caching (#26377 )

2025-11-02 04:16:23 -08:00

moe

[Performance][DP/EP] Add silu_mul_per_token_group_quant_fp8_colmajor kernel (#29470 )

2025-12-03 18:04:59 +00:00

quantization

[Kernel][Quantization] add w4a8 support for marlin kernel (#24722 )

2025-11-29 07:19:33 -08:00

__init__.py

[CI/Build] Move test_utils.py to tests/utils.py (#4425 )

2024-05-13 23:50:09 +09:00

allclose_default.py

Convert formatting to use ruff instead of yapf + isort (#26247 )

2025-10-05 07:06:22 -07:00

quant_utils.py

[Chore]:Extract math and argparse utilities to separate modules (#27188 )

2025-10-26 04:03:32 -07:00

test_apply_repetition_penalties.py

Convert formatting to use ruff instead of yapf + isort (#26247 )

2025-10-05 07:06:22 -07:00

test_cache_kernels.py

[Bugfix][cache_kernels]: Fix OOB in cache_kernels.cu (#28760 )

2025-11-20 02:52:02 -08:00

test_fla_layernorm_guard.py

[PERF] [Qwen3-next] Speed up gated RMSNorm (#26207 )

2025-10-12 08:27:50 +00:00

test_flex_attention.py

[V0 Deprecation] Remove VLLM_USE_V1 from tests (#26341 )

2025-10-07 15:42:31 +00:00

test_fused_quant_activation.py

Convert formatting to use ruff instead of yapf + isort (#26247 )

2025-10-05 07:06:22 -07:00

test_onednn.py

[CPU] Refactor CPU attention backend (#27954 )

2025-11-12 09:43:06 +08:00

test_shuffle_rows.py

Convert formatting to use ruff instead of yapf + isort (#26247 )

2025-10-05 07:06:22 -07:00

test_top_k_per_row.py

[Deepseek v3.2] Remove extra logics in indexer (#26465 )

2025-10-21 23:34:03 +00:00

utils.py

[Feat] Support non-gated activations in NVFP4 modelopt path (#29004 )

2025-11-30 11:02:40 -05:00