xinyun/vllm - vllm - 丝路新云-代码仓

mirror of https://git.datalinker.icu/vllm-project/vllm.git synced 2026-07-20 21:37:15 +08:00

Author	SHA1	Message	Date
Xin Yang	468a8d72ba	[Bugfix] Fix FusedMoEModularKernel for triton backend (#28913 ) Signed-off-by: Xin Yang <xyangx@amazon.com>	2025-11-19 13:05:22 +08:00
Matthew Bonanni	4c23690f43	[Attention] FlashAttention ViT support, make default backend (#28763 ) Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>	2025-11-18 20:06:21 -08:00
Strahinja Stamenkovic	814843e021	Enable bitsandbytes quantization on AMD GPUs that use warp size 32 (#27307 ) Signed-off-by: sstamenk <strahinja.stamenkovic@amd.com>	2025-11-19 03:12:31 +00:00
Li, Jiang	20852c8f4c	[CPU] Refactor CPU WNA16 (#28826 ) Signed-off-by: jiang1.li <jiang1.li@intel.com>	2025-11-19 10:32:00 +08:00
Jialin Ouyang	40b6b38f2c	[Core] Switch Flat logprob control from environment variable to SamplingParams (#28914 ) Signed-off-by: Jialin Ouyang <Jialin.Ouyang@gmail.com> Co-authored-by: 22quinn <33176974+22quinn@users.noreply.github.com>	2025-11-19 02:10:02 +00:00
Jerry Zhang	da94c7c0eb	Move online quantization to `model.load_weights` (#26327 ) Signed-off-by: Jerry Zhang <jerryzh168@gmail.com>	2025-11-18 16:52:41 -08:00
tomeras91	1395461f5f	[Hybrid][torch.compile] Refactor mamba2 forward to avoid obscuring linear projections under custom op (#28587 ) Signed-off-by: Tomer Asida <57313761+tomeras91@users.noreply.github.com>	2025-11-18 16:49:36 -08:00
Varun Sundar Rabindranath	9912b8ccb8	[Build] Add OpenAI triton_kernels (#28788 ) Signed-off-by: Varun Sundar Rabindranath <vsundarr@redhat.com> Co-authored-by: Varun Sundar Rabindranath <vsundarr@redhat.com>	2025-11-18 16:45:20 -08:00
Johnny	49ef847aa8	[NVIDIA] Guard SM100 CUTLASS MoE macro to SM100 builds v2 (#28938 ) Signed-off-by: johnnynunez <johnnynuca14@gmail.com> Signed-off-by: Johnny <johnnynuca14@gmail.com>	2025-11-18 16:44:27 -08:00
Michael Goin	67745d189f	Supress verbose logs from model_hosting_container_standards (#28949 ) Signed-off-by: mgoin <mgoin64@gmail.com>	2025-11-18 12:29:06 -08:00
Kunshang Ji	2a2d5d2780	Replace `torch.cuda.Event` with `torch.Event` for better hardware compatibility (#26985 ) Signed-off-by: Kunshang Ji <kunshang.ji@intel.com>	2025-11-18 11:34:36 -08:00
Chendi.Xue	c3e2978620	[NIXL] fix cpu PD after physical <> logical block_size PR (#28904 ) Signed-off-by: Chendi Xue <chendi.xue@intel.com>	2025-11-18 14:03:23 -05:00
Isotr0py	e4bb2684bc	[Models] Replace all `nn.Conv2d` with vLLM's Conv2dLayer (#28842 ) Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn>	2025-11-18 18:56:04 +00:00
Kevin H. Luu	c64c0b78de	[chore] Move the rest of wikimedia url to S3 (#28921 ) Signed-off-by: Kevin H. Luu <khluu000@gmail.com> Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>	2025-11-18 09:44:18 -08:00
vllmellm	0af3d4f0df	[FEAT] [AITER] [ROCm] integrate aiter sampling ops (#26084 ) Signed-off-by: vllmellm <vllm.ellm@embeddedllm.com>	2025-11-18 17:28:34 +00:00
Nick Hill	da8dadf68b	[Minor] Rename `ec_producer` field to `is_ec_producer` (#28884 ) Signed-off-by: Nick Hill <nhill@redhat.com>	2025-11-18 17:26:07 +00:00
Nicolò Lucchesi	f226a3f0c1	[CI][NIXL] Change default `block_size` for tests (#28927 ) Signed-off-by: NickLucche <nlucches@redhat.com>	2025-11-18 09:22:30 -08:00
Luciano Martins	c2612371ad	[Model] Add Gemma3 GGUF multimodal support (#27772 ) Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com> Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn> Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com> Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn>	2025-11-18 08:56:29 -08:00
Ido Segev	49a986ecd4	[Benchmark] multi_turn: Report warmup-inclusive runtime (#28937 ) Signed-off-by: Ido Segev <idos@pliops.com>	2025-11-18 16:38:22 +00:00
Alex	f6aa122698	[CI Sprint] Quantization CI Cleanup (#24130 ) Signed-off-by: Alex Yun <alexyun04@gmail.com>	2025-11-18 09:21:48 -05:00
Nicolò Lucchesi	184b12fdc6	[Bugfix][NIXL] Fix `block_size_ratio` when logical !=physical blocks (#28925 ) Signed-off-by: NickLucche <nlucches@redhat.com> Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>	2025-11-18 22:07:50 +08:00
Canlin Guo	b9489f51e1	[Model][Perf] Use cos and sin cache in QwenVL (#28798 ) Signed-off-by: gcanlin <canlinguosdu@gmail.com>	2025-11-18 11:51:54 +00:00
Song Zhixin	285eaa4285	[Bugfix] Safeguard against missing backend in AttentionBackendEnum (#28846 ) Signed-off-by: jesse <szxfml@gmail.com> Signed-off-by: Song Zhixin <szxfml@gmail.com> Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>	2025-11-18 10:53:44 +00:00
Nick Hill	439368496d	[BugFix] Fix PP/async scheduling with pooling models (#28899 ) Signed-off-by: Nick Hill <nhill@redhat.com> Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>	2025-11-18 00:20:45 -08:00
Isotr0py	896e41ae04	[CI/Build] Replace wikipedia url with local server ones (#28908 ) Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn>	2025-11-18 08:10:55 +00:00
Kuntai Du	5bb1da5190	[MISC] Remove format.sh (#28906 ) Signed-off-by: Kuntai Du <kuntai@uchicago.edu>	2025-11-18 05:28:31 +00:00
Nick Hill	5bdd155277	[CI] Fix async scheduling + spec decoding test flake (#28902 ) Signed-off-by: Nick Hill <nhill@redhat.com>	2025-11-18 05:26:32 +00:00
Ning Xie	0168f69e50	[Misc] Remove unnecessary parentheses from log statements (#28897 ) Signed-off-by: Andy Xie <andy.xning@gmail.com>	2025-11-17 20:33:46 -08:00
Didier Durand	083cf326dc	[Doc]: fix typos in various files (#28863 ) Signed-off-by: Didier Durand <durand.didier@gmail.com>	2025-11-17 20:32:14 -08:00
Cyrus Leung	bf9e1e8767	[Bugfix] Fix wrong CLI defaults for dynamic `SchedulerConfig` fields (#28872 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>	2025-11-17 20:30:29 -08:00
Wentao Ye	3ddcf46011	[Refactor] Remove Unused Func in Batch Invariant (#28881 ) Signed-off-by: yewentao256 <zhyanwentao@126.com>	2025-11-17 20:29:29 -08:00
xuebwang-amd	d0a73620cc	[ROCm][Quantization] add apply_vllm_mapper in quark config for models like gpt-oss (#28638 ) Signed-off-by: xuebwang-amd <xuebwang@amd.com> Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>	2025-11-18 11:16:45 +08:00
Michael Goin	88ab591f0b	Run macos smoke test workflow on main commit (#28752 ) Signed-off-by: Michael Goin <mgoin64@gmail.com> Signed-off-by: mgoin <mgoin64@gmail.com>	2025-11-18 11:16:03 +08:00
Benjamin Bartels	b6e04390d3	[Bugfix] Fix Kimi-K2 tool parser concatenated tool calls parsing (#28831 ) Signed-off-by: Thomas Mao <yiyeguhu@gmail.com> Signed-off-by: bbartels <benjamin@bartels.dev> Co-authored-by: Thomas Mao <yiyeguhu@gmail.com> Co-authored-by: Chauncey <chaunceyjiang@gmail.com>	2025-11-17 19:13:25 -08:00
Zhuohan Li	552cac95b5	[Misc] Fix wrong comment in scheduler (#28880 ) Signed-off-by: Zhuohan Li <zhuohan123@gmail.com>	2025-11-17 15:32:22 -08:00
Bangsheng Tang	61485844fc	[BugFix] Corner case that could cause out-of-sync with external launcher mode and dp >1 (#28774 )	2025-11-17 15:22:11 -08:00
Pranav	f77bce001a	[Model] Add Afmoe architecture implementation (#28332 ) Signed-off-by: Maziyar Panahi <maziyar.panahi@iscpif.fr> Signed-off-by: Pranav <veldurthipranav@gmail.com> Co-authored-by: Maziyar Panahi <maziyar.panahi@iscpif.fr>	2025-11-17 15:11:20 -08:00
Wentao Ye	a289cc1dde	[Test] Batch Invariant: Rename and organize tests (#27421 ) Signed-off-by: yewentao256 <zhyanwentao@126.com>	2025-11-17 18:09:47 -05:00
Shreyas Kulkarni	95ae50b7d1	[Quantization] [Eagle] Add complete quantization support to the draft model in Eagle (#28435 ) Signed-off-by: Shreyas Kulkarni <shreyas.gp269@gmail.com>	2025-11-17 15:01:34 -08:00
Nick Hill	7765e5ba75	[BugFix] Fix PP performance and PP kv connector output regression (#28768 ) Signed-off-by: Nick Hill <nhill@redhat.com>	2025-11-17 14:08:50 -08:00
Ronald	d8874c61a5	[Core] Async Scheduling X Spec Decoding Compatibility (#24799 ) Signed-off-by: Ronald1995 <ronaldautomobile@163.com> Signed-off-by: Nick Hill <nhill@redhat.com> Signed-off-by: Benjamin Chislett <chislett.ben@gmail.com> Co-authored-by: Nick Hill <nhill@redhat.com> Co-authored-by: Benjamin Chislett <chislett.ben@gmail.com>	2025-11-17 12:16:20 -08:00
Zhewen Li	f8b19c0ffd	[Bugfix] Fix GPT-OSS on AMD after #28603 (#28816 ) Signed-off-by: zhewenli <zhewenli@meta.com>	2025-11-17 13:15:26 -05:00
tiehexue	e42bd8c2e3	Cast return value to int64_t for cache size (#28814 ) Signed-off-by: tiehexue <tiehexue@hotmail.com>	2025-11-17 16:02:32 +00:00
Roger Wang	7f064491f8	[Bugfix][Perf] Revert applying HF processor on text-only inputs for multimodal models (#28858 ) Signed-off-by: Roger Wang <hey@rogerw.io>	2025-11-17 14:49:25 +00:00
Lucas Wilkinson	64e39d667c	[BugFix] Temporary fix for IMA with MTP = 2 and full-cg (#28315 ) Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>	2025-11-17 09:41:22 -05:00
Kunshang Ji	1b82fb0ad3	[XPU] work around for sp, avoid custom op import error (#28822 ) Signed-off-by: Kunshang Ji <kunshang.ji@intel.com>	2025-11-17 13:16:44 +00:00
Jae-Won Chung	d4acf518d0	[Metrics] Fix KV cache usage percent metric multiproc (#28792 ) The `vllm:kv_cache_usage_perc` Gauge metric is missing `multiprocess_mode="mostrecent"` and ends up returning ``` vllm:kv_cache_usage_perc{engine="0",model_name="Qwen/Qwen3-VL-8B-Instruct",pid="277"} 0.0 vllm:kv_cache_usage_perc{engine="0",model_name="Qwen/Qwen3-VL-8B-Instruct",pid="275"} 0.0 vllm:kv_cache_usage_perc{engine="0",model_name="Qwen/Qwen3-VL-8B-Instruct",pid="273"} 0.6530455880475035 ... ``` The deprecated `vllm:gpu_cache_usage_perc` Gauge metric has `multiprocess_mode="mostrecent"`. Signed-off-by: Jae-Won Chung <jwnchung@umich.edu>	2025-11-17 09:54:15 +00:00
wuyaoxuehun	ab01cd14e5	[BugFix] Fix glm4_moe_mtp load weights bug (#28805 ) Signed-off-by: wuyaoxuehun <798143193@qq.com>	2025-11-17 17:13:11 +08:00
Li, Jiang	577bb34fff	[CPU][Bugfix] Fix _to_list in CPU model runner (#28824 ) Signed-off-by: jiang1.li <jiang1.li@intel.com>	2025-11-17 07:47:24 +00:00
Jee Jee Li	3380ed5e11	[Doc] Add llama4 LoRA tag (#28825 ) Signed-off-by: Jee Jee Li <pandaleefree@gmail.com>	2025-11-17 14:08:48 +08:00

... 3 4 5 6 7 ...

11610 Commits