xinyun/vllm - vllm - 丝路新云-代码仓

mirror of https://git.datalinker.icu/vllm-project/vllm.git synced 2026-07-10 04:57:09 +08:00

Author	SHA1	Message	Date
Isotr0py	7c16f3fbcc	[Doc] Add documents for multi-node distributed serving with MP backend (#30509 ) Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn>	2025-12-13 18:02:29 +00:00
Laith Sakka	763963aa73	set assume_32bit_indexing and pass unbacked hints (#30459 ) Signed-off-by: Laith Sakka <lsakka@meta.com>	2025-12-13 15:36:53 +00:00
Cyrus Leung	39cefbdf17	[Refactor] `TokenizerRegistry` only uses lazy imports (#30609 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>	2025-12-13 23:16:22 +08:00
Chen Zhang	ace34e3783	[Bugfix] Qwen3-next with --hf-overrides \{\"num_hidden_layers\":8\} (#30433 ) Signed-off-by: Chen Zhang <zhangch99@outlook.com>	2025-12-13 22:12:45 +08:00
Cyrus Leung	64251f48df	[Chore] Adjust tokenizer import to avoid circular imports (#30601 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>	2025-12-13 04:42:39 -08:00
Nick Hill	1cec5b7ea9	[Scheduer] Simplify stop checking for pooling models (#30591 ) Signed-off-by: Nick Hill <nhill@redhat.com>	2025-12-13 09:45:26 +00:00
Cyrus Leung	b09806e28f	[Bugfix] Dictionary MM embeddings for online chat (#30507 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>	2025-12-13 15:48:56 +08:00
Tsukasa OI	fdc135d768	[Misc][Quantization] Clarify the intent of GGUF `FusedMoE` weight materialization (#30310 ) Signed-off-by: Tsukasa OI <floss_llm@irq.a4lg.com>	2025-12-13 13:55:14 +08:00
Roberto L. Castro	4fa7ce46f3	[Feature] Add SM103 (Blackwell Ultra) Support to vLLM (#30484 ) Signed-off-by: LopezCastroRoberto <robertol.c510@gmail.com> Signed-off-by: Roberto L. Castro <38211239+LopezCastroRoberto@users.noreply.github.com> Co-authored-by: youkaichao <youkaichao@gmail.com>	2025-12-12 19:34:23 -08:00
Matthew Bonanni	86a3261525	[Bugfix] Pass FA version in `MultiHeadAttention` (#30575 ) Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>	2025-12-13 00:02:11 +00:00
rasmith	08f8a5627e	[CI/Build][Kernel][BugFix][AMD] Fix per_token_group_quant_fp8 to use correct fp8 min/max values and update atol/rtol in test_quantfp8_group_functionality (#30292 ) Signed-off-by: Randall Smith <ransmith@amd.com> Co-authored-by: Randall Smith <ransmith@amd.com>	2025-12-12 18:41:56 -05:00
danielafrimi	13618626df	[MoE-FP8-modelopt] Add FlashInfer alignment padding for intermediate dimensions (#29748 ) Signed-off-by: Daniel Afrimi <dafrimi@pool0-00589.cm.cluster> Signed-off-by: dafrimi <dafrimi@nvidia.com> Co-authored-by: Daniel Afrimi <dafrimi@pool0-00589.cm.cluster> Co-authored-by: Tyler Michael Smith <tyler@neuralmagic.com>	2025-12-12 20:42:32 +00:00
danielafrimi	6ec0d8dbe4	[Fix]Load kv-cache dtype from hf_quant_config.json automatically (#29980 ) Signed-off-by: Daniel Afrimi <dafrimi@nvidia.com>	2025-12-12 11:27:47 -08:00
Xin Yang	1f19d8f899	[Perf] Set split_k to 1 for triton_kernels (#30528 ) Signed-off-by: Xin Yang <xyangx@amazon.com>	2025-12-12 14:07:57 -05:00
shivampr	cd7740ac5c	[ROCm] Enable Triton ScaledMM fallback + kernel selection fix (#26668 ) Signed-off-by: Shivam <shivampr.dev@gmail.com> Signed-off-by: Shivam <shivamprasad91@gmail.com>	2025-12-12 13:28:20 -05:00
Wentao Ye	02a5880394	[CI] Fix mypy for vllm/v1/executor (#30517 ) Signed-off-by: yewentao256 <zhyanwentao@126.com>	2025-12-12 18:05:34 +00:00
realliujiaxu	d2c919dcc2	[bugfix] fix bug when top_logprobs=0 with spec decoding (#30059 ) Signed-off-by: realliujiaxu <realliujiaxu@163.com>	2025-12-12 09:03:35 -08:00
Benjamin Bartels	f3237f3f6b	[Frontend] Fixes anthropic streaming message_start usage nesting (#30266 ) Signed-off-by: bbartels <benjamin@bartels.dev>	2025-12-12 16:28:54 +00:00
jvlunteren	9c0ee995a8	[Kernel] Support CUDA Graphs in 3D Triton Attention Kernel (#28306 ) Signed-off-by: Jan van Lunteren <jvl@zurich.ibm.com> Signed-off-by: jvlunteren <161835099+jvlunteren@users.noreply.github.com> Co-authored-by: Thomas Parnell <tom.parnell@gmail.com> Co-authored-by: Thomas Parnell <tpa@zurich.ibm.com>	2025-12-12 16:55:40 +01:00
Michael Goin	09ad3b76b3	[Bug] Fix attention_backend arg string parsing (#30534 ) Signed-off-by: mgoin <mgoin64@gmail.com>	2025-12-12 08:40:50 -07:00
Christina Norman	dc13c99eed	fix(gguf): Disable bfloat16 for GGUF on blackwell device (#30408 ) Signed-off-by: Christina <truffle@gmail.com> Signed-off-by: Isotr0py <2037008807@qq.com> Signed-off-by: Christina Norman <christina@example.com> Co-authored-by: Isotr0py <isotr0py@users.noreply.github.com> Co-authored-by: Isotr0py <2037008807@qq.com> Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>	2025-12-12 23:10:12 +08:00
Vladislav Nosivskoy	3e34adcdfb	[DeepSeek V3.2] Proper drop_thinking logic (#30490 ) Signed-off-by: Vladislav Nosivskoy <vladnosiv@gmail.com>	2025-12-12 15:01:06 +00:00
Lucas Wilkinson	3e41992fec	[Attention] Use sparse prefill kernel for fp8 kv-cache in DeepSeek-v3.2 (#27532 ) Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>	2025-12-12 05:57:47 -08:00
Jaehwang Jung	f90319d5d1	[Bugfix] Schedule failure due to wrong get_image_size_with_most_features (#29692 )	2025-12-12 02:27:20 -08:00
Ben Browning	8f8fda261a	[Bugfix] Multiple fixes for gpt-oss Chat Completion prompting (#28729 ) Signed-off-by: Ben Browning <bbrownin@redhat.com> Co-authored-by: Chauncey <chaunceyjiang@gmail.com>	2025-12-12 12:59:53 +08:00
Zhengxu Chen	fe1787107e	[compile] Parse compile range cache keys as Range during cache loading. (#30516 ) Signed-off-by: zhxchen17 <zhxchen17@fb.com>	2025-12-12 04:30:51 +00:00
Nick Hill	947dfda9c2	[LMCache] Relax lmcache version requirement (#30425 ) Signed-off-by: Nick Hill <nhill@redhat.com>	2025-12-11 18:18:47 -09:00
Michael Goin	9f2fc16a69	[Bugfix][Model] Fix Afmoe rope_parameters issue (#30505 ) Signed-off-by: mgoin <mgoin64@gmail.com> Signed-off-by: Michael Goin <mgoin64@gmail.com> Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2025-12-12 02:53:57 +00:00
Bhanu Prakash Voutharoja	6a6fc41c79	gptq marlin quantization support for fused moe with lora (#30254 ) Signed-off-by: Bhanu068 <voutharoja.bhanu06@gmail.com>	2025-12-12 02:27:22 +00:00
Fadi Arafeh	f355ad5412	[CPU][FIX] Fix build failures on Arm CPUs with torch nightly (#30481 ) Signed-off-by: Fadi Arafeh <fadi.arafeh@arm.com>	2025-12-12 02:09:25 +00:00
Lucas Wilkinson	042da73244	[Core] Refactor `_build_attention_metadata` (#29628 ) Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>	2025-12-11 17:54:12 -08:00
jiahanc	0ab23c2b2b	[fix] fix SM check for Flashinfer TRTLLM MOE (#30314 ) Signed-off-by: jiahanc <173873397+jiahanc@users.noreply.github.com>	2025-12-12 01:00:58 +00:00
Andrew Briand	a00d88973d	[EPLB] Support EPLB w/ NVFP4 (#29804 ) Signed-off-by: Andrew Briand <abriand@nvidia.com> Co-authored-by: Andrew Briand <abriand@nvidia.com>	2025-12-11 22:59:40 +00:00
Wentao Ye	c817b14151	[Perf] Optimize deepgemm experts initialization, 3.9% TTFT improvement (#30494 ) Signed-off-by: yewentao256 <zhyanwentao@126.com> Co-authored-by: li-jinpeng <3332126450@qq.com> Co-authored-by: youkaichao <youkaichao@gmail.com>	2025-12-11 17:28:34 -05:00
Nicolò Lucchesi	0efd9f867c	[Core] Whisper Enable Encoder Batching (#29421 ) Signed-off-by: NickLucche <nlucches@redhat.com>	2025-12-11 21:06:51 +00:00
Xingyu Liu	90d6cf921f	[BugFix][MM]support VLLM_RANDOMIZE_DP_DUMMY_INPUTS (#30472 ) Signed-off-by: Xingyu Liu <charlotteliu12x@gmail.com> Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>	2025-12-11 21:00:15 +00:00
Harry Mellor	cf3eacfe58	Standardise `get_rope` to use `rope_parameters["partial_rotary_factor"]`, not `rotary_dim` (#30389 ) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2025-12-11 20:45:23 +00:00
Zhengxu Chen	92fea56fd1	[compile] Stop one-off setting enable_aot_compile and use context manager instead. (#30503 ) Signed-off-by: zhxchen17 <zhxchen17@fb.com>	2025-12-11 20:28:03 +00:00
Andreas Karatzas	72aaac5b66	[ROCm][Bugfix] Add MLACommonMetadata to allowed attention types for speculative decoding (#30430 ) Signed-off-by: Andreas Karatzas <akaratza@amd.com>	2025-12-11 19:25:01 +00:00
汪志鹏	0e71eaa644	[Feature] AWQ marlin quantization support for fused moe with lora (#30442 ) Signed-off-by: princepride <wangzhipeng628@gmail.com>	2025-12-11 18:03:32 +00:00
Harry Mellor	8781cd6b88	Add Eagle and Eagle3 support to Transformers modeling backend (#30340 ) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2025-12-11 17:02:10 +00:00
Julien Denize	aa3c250c48	[IMPROVEMENT] Change MistralReasoningParser behavior (#30391 ) Signed-off-by: juliendenize <julien.denize@mistral.ai> Signed-off-by: Julien Denize <40604584+juliendenize@users.noreply.github.com> Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com>	2025-12-11 17:53:26 +01:00
Harry Mellor	93db3256a4	Give pooling examples better names (#30488 ) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2025-12-11 16:22:58 +00:00
Harry Mellor	97a042f3bc	Make the `httpx` logger less annoying when Transformers v5 is installed (#30480 ) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2025-12-11 15:44:56 +00:00
Cyrus Leung	3a3b06ee70	[Misc] Improve error message for `is_multimodal` (#30483 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>	2025-12-11 06:39:51 -08:00
Martin Hickey	f4417f8449	[KVConnector] Add KV events to KV Connectors (#28309 ) Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>	2025-12-11 15:30:29 +01:00
Qiu	a11f4a81e0	[Misc][PCP&DCP] relocate PCP feature check (#30050 ) Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com> Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>	2025-12-11 03:36:18 -08:00
Kenichi Maehashi	853611bb18	Fix typo of endpoint name in CLI args docs (#30473 ) Signed-off-by: Kenichi Maehashi <maehashi@preferred.jp>	2025-12-11 11:07:56 +00:00
wang.yuqi	a5f9fb5960	[Deprecation] Deprecation `--convert reward`, use `--convert embed` instead. (#30463 ) Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>	2025-12-11 10:18:25 +00:00
jeremyteboul	4515eb1a0b	[Fix] Update lazing loading of video loader backend (#30444 ) Signed-off-by: Jeremy Teboul <jeremyteboul@fb.com> Co-authored-by: Jeremy Teboul <jeremyteboul@fb.com>	2025-12-11 10:14:57 +00:00

1 2 3 4 5 ...

8540 Commits