xinyun/vllm - vllm - 丝路新云-代码仓

mirror of https://git.datalinker.icu/vllm-project/vllm.git synced 2026-06-24 09:17:13 +08:00

Author	SHA1	Message	Date
Michael Goin	e08a3a3fdb	[CI Failure] Disable FlashInfer RoPE to unblock CI (#25299 ) Signed-off-by: mgoin <mgoin64@gmail.com>	2025-09-20 08:16:56 +00:00
Cyrus Leung	3d9a1d2de5	[V1] Support `LLM.apply_model` (#18465 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>	2025-09-20 07:14:35 +00:00
Roger Wang	be874c0201	[Bugfix] Fix Qwen3-VL-MoE weight loading for EP (#25300 ) Signed-off-by: Roger Wang <hey@rogerw.io>	2025-09-20 00:04:05 -07:00
Chen Zhang	9607d5eb44	[Hybrid Allocator] Support full attention with different hidden size (#25101 ) Signed-off-by: Chen Zhang <zhangch99@outlook.com>	2025-09-19 23:43:59 -07:00
Cyrus Leung	c60e6137f0	[Optimization] Avoid repeated model architecture conversion for pooling models (#25261 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>	2025-09-20 13:30:22 +08:00
Chauncey	f91480b2d4	[Bugfix] fix tool call arguments is empty (#25223 ) Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com> Co-authored-by: xin.li <xin.li@daocloud.io>	2025-09-20 13:29:54 +08:00
Chendi.Xue	6c5f82e5aa	[BUG FIX][NON-CUDA]quick fix to avoid call cudagraph_unsafe in attention (#25298 ) Signed-off-by: Chendi Xue <Chendi.Xue@intel.com>	2025-09-20 04:41:23 +00:00
Nick Hill	b7f186bbb3	[BugFix] Exclude self when checking for port collision (#25286 ) Signed-off-by: Nick Hill <nhill@redhat.com>	2025-09-20 12:28:31 +08:00
JartX	3642909617	[BUGFIX] GPTQ quantization compatibility for Qwen3 Next MOE models (AutoGPTQ and AutoRound-GPTQ) (#25268 ) Signed-off-by: JartX <sagformas@epdcenter.es>	2025-09-20 11:18:13 +08:00
Harry Mellor	c308501cb6	Improve weight loading for encoder models in Transformers backend (#25289 ) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2025-09-20 03:11:03 +00:00
Nick Hill	535d80056b	[Misc] Support more collective_rpc return types (#25294 ) Signed-off-by: Nick Hill <nhill@redhat.com>	2025-09-20 02:02:38 +00:00
Nick Hill	a25ade5d47	[BugFix] Ensure appropriate guards in destructors (#25284 ) Signed-off-by: Nick Hill <nhill@redhat.com> Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>	2025-09-20 09:06:34 +08:00
Boyuan Feng	8945b001db	[torch.compile] CUDAGraph Inductor partition integration (#24281 ) Signed-off-by: Boyuan Feng <boyuan@meta.com> Signed-off-by: Boyuan Feng <fby.1994@gmail.com> Signed-off-by: boyuanfeng <boyuan@meta.com> Co-authored-by: Luka Govedič <ProExpertProg@users.noreply.github.com>	2025-09-20 01:02:15 +00:00
Andrew Sansom	b8a287a0a8	[docs] Prompt Embedding feature support (#25288 ) Signed-off-by: Andrew Sansom <andrew@protopia.ai>	2025-09-19 17:46:23 -07:00
Andrew Sansom	c7e713616a	test: Remove vestigial skip for prompt embeds tests after landing v1 Prompt Embeds support (#25291 ) Signed-off-by: Andrew Sansom <andrew@protopia.ai>	2025-09-19 17:33:40 -07:00
Maximilien de Bayser	a36c675817	Don't skip special tokens with hermes-style tool calling (#25281 ) Signed-off-by: Max de Bayser <mbayser@br.ibm.com>	2025-09-19 17:33:25 -07:00
Lucas Kabela	3da17c2cc2	[Bugfix] Remove VLLM_TEST_DYNAMO_FULLGRAPH_CAPTURE #2969 (#25090 ) Signed-off-by: Lucas Kabela <lucaskabela@meta.com>	2025-09-19 20:27:21 -04:00
Nick Hill	14c1432789	[BugFix] Fix async scheduling CPU tensor race take 2 (#25279 ) Signed-off-by: Nick Hill <nhill@redhat.com>	2025-09-19 16:34:07 -07:00
Lucia Fang	ee7a66dd9a	allow disable flashinfer prefill (#25276 ) Signed-off-by: Lu Fang <fanglu@fb.com>	2025-09-19 22:59:41 +00:00
Zhiyu	431535b522	Enable modelopt gemma3 nvfp4/fp8, make workflow more robust (#22771 ) Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com> Signed-off-by: Michael Goin <mgoin64@gmail.com> Co-authored-by: Michael Goin <mgoin64@gmail.com>	2025-09-19 22:40:33 +00:00
Wentao Ye	711e912946	[Compile] Fix Compile Warning for Ignoring `MIN_BLOCK_PER_SM` (#25193 ) Signed-off-by: yewentao256 <zhyanwentao@126.com>	2025-09-19 16:23:19 -06:00
Alec S	e69e0b8b5f	[Frontend] Responses API messages out, just harmony for now (#24985 ) Signed-off-by: Alec Solder <alecs@fb.com> Co-authored-by: Alec Solder <alecs@fb.com> Co-authored-by: Ye (Charlotte) Qi <yeq@meta.com>	2025-09-19 21:40:16 +00:00
David-Wen	ddc9048394	Fix: Correct FusedMoE layer reference in auto_round quantization (#24818 ) Signed-off-by: David-Wen <18927700430@163.com> Signed-off-by: Michael Goin <mgoin64@gmail.com> Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com> Co-authored-by: Michael Goin <mgoin64@gmail.com> Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>	2025-09-19 20:44:24 +00:00
nvjullin	b1a63d1b3b	[BugFix] Make FlashInferMetadataBuilder non-blocking (#25040 ) Signed-off-by: Julien Lin <jullin@nvidia.com> Co-authored-by: Michael Goin <mgoin64@gmail.com>	2025-09-19 20:36:34 +00:00
Michael Goin	48ecb4438b	[Perf] Use FlashInfer RoPE for RotaryEmbedding.forward_cuda when available (#21126 ) Signed-off-by: mgoin <mgoin64@gmail.com> Signed-off-by: Michael Goin <mgoin64@gmail.com> Co-authored-by: Luka Govedič <ProExpertProg@users.noreply.github.com>	2025-09-19 14:06:49 -06:00
Harry Mellor	e57fc15971	Specify platform in `pip-compile` `pre-commit` hook so it runs on MacOS (#25273 ) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2025-09-19 12:43:33 -07:00
bnellnm	4bdf400218	[Bugfix] Fix chunked a2_scales in modular kernels (#25264 ) Signed-off-by: Bill Nell <bnell@redhat.com>	2025-09-19 19:42:01 +00:00
Varun Sundar Rabindranath	7852b82b93	[Bugfix] GPT OSS Attritbute error on H100 (#25228 ) Signed-off-by: Varun Sundar Rabindranath <vsundarr@redhat.com> Co-authored-by: Varun Sundar Rabindranath <vsundarr@redhat.com>	2025-09-19 13:14:09 -06:00
qizixi	a2a5f79e09	Optimize triton unified attention performance for sliding window attention (#24390 ) Signed-off-by: zixi-qi <qizixi@meta.com>	2025-09-19 13:07:26 -06:00
Or Ozeri	c59a0eca42	[KV offload][4/N] Offloading KV connector (#22595 ) Signed-off-by: Or Ozeri <oro@il.ibm.com>	2025-09-19 19:07:17 +00:00
Lucia Fang	b716ab93a7	[bugfix] fix structured outputs key missing issue from #24929 (#25195 ) Signed-off-by: Lu Fang <fanglu@fb.com>	2025-09-19 18:37:57 +00:00
samzong	138f0d1e75	[Docs] add __init__.py to vllm/model_executor/layers/quantization/compressed_tensors/transform (#24974 ) Signed-off-by: samzong <samzong.lu@gmail.com>	2025-09-19 18:32:27 +00:00
Jialin Ouyang	2506ce5189	[Core][Prefix Hash] Fix prefix hash metrics sliding window maintainance (#24990 ) Signed-off-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>	2025-09-19 12:22:53 -06:00
Chauncey	47fd08aaf9	[CI/Build] fix test function_calling (#25072 ) Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com>	2025-09-19 12:16:32 -06:00
Harry Mellor	12aed7e453	Encoder model support for the Transformers backend (#25174 ) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2025-09-19 19:15:22 +01:00
LJH-LBJ	d90e212a3a	Remove Redundant Assignment in Qwen3_VisionPatchMerger (#25224 ) Signed-off-by: Junhong <liujunhong11@huawei.com> Co-authored-by: Junhong <liujunhong11@huawei.com> Co-authored-by: Roger Wang <hey@rogerw.io>	2025-09-19 12:15:13 -06:00
Jee Jee Li	2821986450	[Core] Modify the initialization parameters of the lora manager (#25249 ) Signed-off-by: Jee Jee Li <pandaleefree@gmail.com>	2025-09-19 18:01:28 +00:00
Cyrus Leung	6c117cff7d	[Frontend] Pass API server count to each process (#23717 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>	2025-09-20 01:15:19 +08:00
Or Ozeri	7ac67ea525	[KV offload][3/N] Add worker-side CPU support (#21448 ) Signed-off-by: Or Ozeri <oro@il.ibm.com>	2025-09-19 09:53:45 -07:00
samzong	ce75e15373	refactor(benchmarks): add type annotations to wait_for_endpoint parameters (#25218 ) Signed-off-by: samzong <samzong.lu@gmail.com>	2025-09-19 16:36:52 +00:00
Harry Mellor	aed16879a9	Move `ModelConfig` from `config/__init__.py` to `config/model.py` (#25252 ) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2025-09-19 16:22:33 +00:00
Harry Mellor	cf278ff3b2	Update CODEOWNERS (#25269 ) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2025-09-19 09:12:55 -07:00
Icey	838d7116ba	[Qwen] Remove cuda hard-code in qwen3 next (#25243 ) Signed-off-by: Icey <1790571317@qq.com>	2025-09-19 12:25:12 +00:00
Cyrus Leung	5089fd749c	[V0 Deprecation] Remove V0 logic from `get_input_embeddings` interface (#25242 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>	2025-09-19 11:10:52 +00:00
Nicolò Lucchesi	a3d087adec	[P/D][Nixl] Introduce `KVTransferMetrics` and aggregation strategy (#22188 ) Signed-off-by: NickLucche <nlucches@redhat.com>	2025-09-19 11:09:14 +00:00
Harry Mellor	058525b997	Move `PoolerConfig` from `config/__init__.py` to `config/pooler.py` (#25181 ) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2025-09-19 11:02:55 +00:00
Roger Wang	1dfea5f4a9	[Bugfix][Perf] Misc fixes for Qwen3 VL (#25238 ) Signed-off-by: Roger Wang <hey@rogerw.io>	2025-09-19 10:46:16 +00:00
Isotr0py	cea91a32f2	[Kernel][Performance] Add Triton kernel for Qwen3-VL interleaved MRoPE (#25055 ) Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn>	2025-09-19 10:27:49 +00:00
Yan Ma	a684c0124c	[bugfix] fix MHA for models like OpenGVLab/InternVL3_5-38B (#25146 ) Signed-off-by: Yan Ma <yan.ma@intel.com> Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn>	2025-09-19 08:45:06 +00:00
Isotr0py	f2718d2948	[Misc] Cleanup test conftest for deprecated encoder-decoder models (#25231 ) Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn>	2025-09-19 07:44:56 +00:00

1 2 3 4 5 ...

9694 Commits