mirror of https://git.datalinker.icu/vllm-project/vllm.git synced 2026-06-22 18:47:24 +08:00

History

[Chore] Cleanup guided namespace, move to structured outputs config (#22772 )

Signed-off-by: Aaron Pham <contact@aarnphm.xyz>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>

2025-09-18 09:20:27 +00:00

auto_tune

Add a batched auto tune script (#25076 )

2025-09-17 22:41:18 +00:00

cutlass_benchmarks

[Refactor] Remove duplicate ceil_div (#20023 )

2025-06-25 05:19:09 +00:00

disagg_benchmarks

Remove deprecated PyNcclConnector (#24151 )

2025-09-03 22:49:16 +00:00

fused_kernels

[Misc] Add SPDX-FileCopyrightText (#19100 )

2025-06-03 11:20:17 -07:00

kernels

[Kernel] Delegate construction of FusedMoEQuantConfig to FusedMoEMethodBase subclasses (#22537 )

2025-09-17 17:43:31 -06:00

multi_turn

Add more documentation and improve usability of lognormal dist (benchmark_serving_multi_turn) (#23255 )

2025-09-17 05:53:17 +00:00

overheads

[Misc] Add SPDX-FileCopyrightText (#19100 )

2025-06-03 11:20:17 -07:00

structured_schemas

benchmarks: simplify test jsonschema (#14567 )

2025-03-11 13:39:30 +00:00

backend_request_func.py

[Misc] Add request_id into benchmark_serve.py (#23065 )

2025-08-19 08:32:18 +00:00

benchmark_block_pool.py

fix some typos (#24071 )

2025-09-02 20:44:50 -07:00

benchmark_latency.py

[CI/Build][Doc] Fully deprecate old bench scripts for serving / throughput / latency (#24411 )

2025-09-09 10:02:35 +00:00

benchmark_long_document_qa_throughput.py

[Misc] Modularize CLI Argument Parsing in Benchmark Scripts (#19593 )

2025-06-14 16:54:52 +08:00

benchmark_ngram_proposer.py

fix some typos (#24071 )

2025-09-02 20:44:50 -07:00

benchmark_prefix_caching.py

[Misc] Modularize CLI Argument Parsing in Benchmark Scripts (#19593 )

2025-06-14 16:54:52 +08:00

benchmark_prioritization.py

[Misc] Modularize CLI Argument Parsing in Benchmark Scripts (#19593 )

2025-06-14 16:54:52 +08:00

benchmark_serving_structured_output.py

[Chore] Cleanup guided namespace, move to structured outputs config (#22772 )

2025-09-18 09:20:27 +00:00

benchmark_serving.py

[CI/Build][Doc] Fully deprecate old bench scripts for serving / throughput / latency (#24411 )

2025-09-09 10:02:35 +00:00

benchmark_throughput.py

[CI/Build][Doc] Fully deprecate old bench scripts for serving / throughput / latency (#24411 )

2025-09-09 10:02:35 +00:00

benchmark_utils.py

[Core] [N-gram SD Optimization][1/n] Propose tokens with a single KMP (#22437 )

2025-08-13 14:44:06 -07:00

pyproject.toml

[Doc] Move examples and further reorganize user guide (#18666 )

2025-05-26 07:38:04 -07:00

README.md

[Docs] move benchmarks README to contributing guides (#24820 )

2025-09-16 05:52:57 -07:00

run_structured_output_benchmark.sh

[Benchmarks] Refactor run_structured_output_benchmarks.sh (#17722 )

2025-05-13 01:47:29 -07:00

sonnet.txt

feat(benchmarks): Add Prefix Caching Benchmark to Serving Benchmark (#3277 )

2024-03-27 13:39:26 -07:00

README.md

Benchmarks

This directory used to contain vLLM's benchmark scripts and utilities for performance testing and evaluation.

Serving benchmarks: Scripts for testing online inference performance (latency, throughput)
Throughput benchmarks: Scripts for testing offline batch inference performance
Specialized benchmarks: Tools for testing specific features like structured output, prefix caching, long document QA, request prioritization, and multi-modal inference
Dataset utilities: Framework for loading and sampling from various benchmark datasets (ShareGPT, HuggingFace datasets, synthetic data, etc.)

Usage

For detailed usage instructions, examples, and dataset information, see the Benchmark CLI documentation.

For full CLI reference see:

README.md

Benchmarks

Contents

Usage