xinyun/vllm - vllm - 丝路新云-代码仓

mirror of https://git.datalinker.icu/vllm-project/vllm.git synced 2026-06-28 15:37:27 +08:00

Author	SHA1	Message	Date
Michael Goin	fbd6523ac0	Refactor dense FP8 tensor/channel/block utils and add CT FP8 block (#21404 )	2025-09-18 08:53:45 -04:00
Shanshan Shen	470484a4f5	[Structured Output][Refactor] Move `apply_grammar_bitmask()` method from `ModelRunner` to structured output utils (#21999 ) Signed-off-by: shen-shanshan <467638484@qq.com>	2025-09-18 20:44:31 +08:00
Roger Wang	21da73343a	[Misc] Clean up flags in `vllm bench serve` (#25138 ) Signed-off-by: Roger Wang <hey@rogerw.io>	2025-09-18 12:43:33 +00:00
Asaf Joseph Gardin	66072b36db	[Bugfix][Mamba] - Fix Conv State Kernel FP32 Support (#24883 ) Signed-off-by: asafg <39553475+Josephasafg@users.noreply.github.com>	2025-09-18 12:21:17 +00:00
Harry Mellor	3ed1ec4af2	Fix `validate-config` pre-commit check (#25157 ) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2025-09-18 12:06:28 +00:00
Harry Mellor	5a33ae9a3f	Fix forward reference warning in documentation (#25150 ) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2025-09-18 11:41:41 +00:00
William Song	c9ff9e6f0c	[Docs] add the parallel sampling usage in LLMEngine and AsyncLLM (#24222 )	2025-09-18 04:37:08 -07:00
Kay Yan	eaffe4486c	[Docs] Fix pooling-params doc references in openai_compatible_server.md (#24939 )	2025-09-18 04:36:47 -07:00
Harry Mellor	8ed039d527	Move `StructuredOutputsConfig` from `config/__init__.py` to `config/structured_outputs.py` (#25153 ) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2025-09-18 11:24:27 +00:00
Jee Jee Li	37970105fe	[Model] Improve Pooling Model (#25149 ) Signed-off-by: Jee Jee Li <pandaleefree@gmail.com>	2025-09-18 11:04:21 +00:00
Chauncey	cc935fdd7e	[Frontend] Support setting logprobs to -1 (#25031 ) Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com>	2025-09-18 10:34:42 +00:00
Aaron Pham	29283e8976	[Chore] Cleanup guided namespace, move to structured outputs config (#22772 ) Signed-off-by: Aaron Pham <contact@aarnphm.xyz> Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com> Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2025-09-18 09:20:27 +00:00
Punitvara	05b044e698	[Doc] Fix cross-reference warnings (#25058 ) Signed-off-by: Punit Vara <punitvara@gmail.com> Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com> Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2025-09-18 02:05:16 -07:00
Tao He	ef7eefe17a	[Qwen] Add fp8 checkpoint support for qwen3-next. (#25079 ) Signed-off-by: Tao He <linzhu.ht@alibaba-inc.com>	2025-09-18 08:16:04 +00:00
rongfu.leng	350c94deb3	[Bugfix] when use s3 model cannot use default load_format (#24435 ) Signed-off-by: rongfu.leng <rongfu.leng@daocloud.io> Co-authored-by: 22quinn <33176974+22quinn@users.noreply.github.com>	2025-09-18 07:47:43 +00:00
Harry Mellor	f4cd80f944	Retrieve `sliding_window` from text config in Gemma3 MM (#25085 ) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2025-09-18 06:29:05 +00:00
Simon Mo	e111d5b0ae	[CLI] Use streaming in CLI chat and completion commands (#23769 ) Signed-off-by: simon-mo <simon.mo@hey.com>	2025-09-17 22:30:26 -07:00
Simon Mo	a904ea78ea	[benchmark] add peak throughput metrics and plot (#23867 ) Signed-off-by: simon-mo <simon.mo@hey.com>	2025-09-17 22:30:02 -07:00
Benjamin Chislett	b7433ca1a4	[Spec Decode] Efficient padded speculation (#24539 ) Signed-off-by: Benjamin Chislett <bchislett@nvidia.com>	2025-09-18 01:07:24 -04:00
YiwenC	9d8a2d86d2	[EPLB] Add EPLB support for hunyuan_v1 (#23078 )	2025-09-18 04:51:35 +00:00
Chaojun Zhang	3bc18127ff	[XPU] Whisper model support on XPU Platform (#25123 ) Signed-off-by: chzhang <chaojun.zhang@intel.com>	2025-09-18 04:30:10 +00:00
Andrew Sansom	bec060fd99	Mark prompt logprobs as incompatible with prompt embeds at API level (#25077 ) Signed-off-by: Andrew Sansom <andrew@protopia.ai>	2025-09-17 21:25:07 -07:00
YiwenC	52bc9d5b3e	[Model] enable data parallel for InternVL vision encoder (#23909 ) Signed-off-by: Yiwen Chen <yiwen66@berkeley.edu> Signed-off-by: YiwenC <54658925+666even666@users.noreply.github.com> Co-authored-by: Roger Wang <hey@rogerw.io>	2025-09-17 21:11:46 -07:00
bnellnm	dc2979c585	[Kernels] Overlap shared experts with combine instead of dispatch (#24254 ) Signed-off-by: Bill Nell <bnell@redhat.com>	2025-09-18 12:10:21 +08:00
toncao	027d37df38	[Bugfix][Qwen3-Next] add prefixes to shared_expert in qwen3-next and mlp in qwen2moe to successfully load ignored params in quantized models (#24960 ) Signed-off-by: toncao <cpatonn@gmail.com> Co-authored-by: toncao <cpatonn@gmail.com> Co-authored-by: Jee Jee Li <pandaleefree@gmail.com>	2025-09-18 12:08:50 +08:00
Lukas Geiger	b98219670f	[Core][MM] Cleanup `MultiModalCache` (#25006 ) Signed-off-by: Lukas Geiger <lukas.geiger94@gmail.com>	2025-09-17 21:08:41 -07:00
Roger Wang	3127274d02	[MM Encoder] Apply DP ViT for Qwen3-VL model series (#24955 ) Signed-off-by: Roger Wang <hey@rogerw.io> Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn> Co-authored-by: Huang Jie <92386084+JJJYmmm@users.noreply.github.com> Co-authored-by: 松灵 <26085463+wulipc@users.noreply.github.com> Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn>	2025-09-17 21:04:21 -07:00
bnellnm	4ac510f484	[Kernels] Enable DeepGEMM by default (#24462 ) Signed-off-by: Bill Nell <bnell@redhat.com>	2025-09-17 20:19:52 -07:00
bnellnm	5963b98b46	[Kernel] Delegate construction of FusedMoEQuantConfig to FusedMoEMethodBase subclasses (#22537 ) Signed-off-by: Bill Nell <bnell@redhat.com>	2025-09-17 17:43:31 -06:00
elvischenv	e67a79db03	[Bugfix] Refactor Flashinfer TRTLLM attention kernel selection logic (#24600 ) Signed-off-by: elvischenv <219235043+elvischenv@users.noreply.github.com> Co-authored-by: Michael Goin <mgoin64@gmail.com>	2025-09-17 15:36:29 -07:00
Douglas Lehr	1a456c7c90	Aiter mha fp8 fix (#24991 ) Signed-off-by: Doug Lehr <douglehr@amd.com> Co-authored-by: Doug Lehr <douglehr@amd.com>	2025-09-17 22:29:14 +00:00
Andrew Xia	bff2e5f1d6	[gpt-oss][2] fix types for streaming (#24556 ) Signed-off-by: Andrew Xia <axia@meta.com>	2025-09-17 22:04:28 +00:00
ahao-anyscale	f20c3b0951	[BUG] Exclude .pth files when pulling remote files (#25092 ) Signed-off-by: ahao-anyscale <ahao@anyscale.com>	2025-09-17 20:42:09 +00:00
Mohammad Miadh Angkad	883131544f	[Bugfix] Update import path for bc_linter_include (#24766 ) Signed-off-by: Mohammad Miadh Angkad <mangkad.bsdsba2027@aim.edu>	2025-09-17 20:33:11 +00:00
afeldman-nm	7ae9887542	[V1] Logits processor docs (#22919 ) Signed-off-by: Andrew Feldman <afeldman@redhat.com> Signed-off-by: afeldman-nm <156691304+afeldman-nm@users.noreply.github.com> Co-authored-by: Joseph Marinier <Joseph.Marinier@gmail.com>	2025-09-17 11:53:12 -07:00
Michael Goin	8b32464ac1	Change log level from info to debug for IOProcessor (#24999 ) Signed-off-by: Michael Goin <mgoin64@gmail.com>	2025-09-17 10:21:28 -07:00
Woosuk Kwon	99cc41ad50	[V0 Deprecation] Remove unused output processor util (#25023 ) Signed-off-by: Woosuk Kwon <woosuk@thinkingmachines.ai>	2025-09-17 09:50:07 -07:00
Simon Mo	4aa8c7b047	cleanup: remove adapter commons (#25045 ) Signed-off-by: Jee Jee Li <pandaleefree@gmail.com> Co-authored-by: Jee Jee Li <pandaleefree@gmail.com>	2025-09-17 16:46:29 +00:00
samzong	4a2d33e371	[Docs] vllm/benchmarks/datasets.py fix docstring param format. (#24970 ) Signed-off-by: samzong <samzong.lu@gmail.com>	2025-09-17 08:11:51 -07:00
Matthew Bonanni	8f3616f422	Remove old cutlass mla (#23961 ) Signed-off-by: Matthew Bonanni <mbonanni001@gmail.com> Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>	2025-09-17 14:31:43 +00:00
samzong	47f670b03b	[Docs] improve code formatting and comments for eliminate griffe build warning. (#25010 ) Signed-off-by: samzong <samzong.lu@gmail.com>	2025-09-17 07:31:20 -07:00
Tao He	dd6a910aac	[Bugfix][Qwen3-Next] fixes the varlen issue in qwen3-next's MTP implementation. (#24957 ) Signed-off-by: Tao He <linzhu.ht@alibaba-inc.com>	2025-09-17 21:59:09 +08:00
Li, Jiang	9fccd04e30	[Bugfix] Fix Stream usage in CPU model runner and OneDNN kernel check (#25046 ) Signed-off-by: jiang1.li <jiang1.li@intel.com>	2025-09-17 05:54:02 -07:00
danielafrimi	252ada5559	Add RADIO Vision Encoder Support to vLLM (#24595 ) Signed-off-by: Daniel Afrimi <danielafrimi8@gmail.com> Co-authored-by: root <root@cw-dfw-h100-001-305-026.cm.cluster>	2025-09-17 05:53:30 -07:00
Shijun Yin	2b85697031	[BugFix] enable DOTALL to match multi-line tool_call parameters in extract_tool_call_required_streaming (#24668 ) Signed-off-by: Shijun Yin <shijun.yin@outlook.com>	2025-09-17 09:21:18 +00:00
Chauncey	544fe76b95	[Frontend] Support returning all prompt logprobs (#24956 ) Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com>	2025-09-17 09:03:52 +00:00
Xinyu Chen	bb58dc8c20	[DP] Create placement groups by ray_device_key (#25026 ) Signed-off-by: Xinyu Chen <xinyu1.chen@intel.com> Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>	2025-09-17 08:57:25 +00:00
Michael Yao	0fb2551c23	[Docs] Fix griffe warning in base_static_graph.py (#25018 ) Signed-off-by: windsonsea <haifeng.yao@daocloud.io>	2025-09-17 08:49:19 +00:00
Zhuohan Li	6c47f6bfa4	[Core] Remove tokenizer group in vLLM (#24078 ) Signed-off-by: Zhuohan Li <zhuohan123@gmail.com>	2025-09-17 08:42:59 +00:00
whx	c15309a730	[Model] Apply SharedFusedMoE to glm4_moe. (#24849 ) Signed-off-by: whx-sjtu <2952154980@qq.com>	2025-09-17 16:02:31 +08:00

1 2 3 4 5 ...

6554 Commits