xinyun/vllm - vllm - 丝路新云-代码仓

mirror of https://git.datalinker.icu/vllm-project/vllm.git synced 2026-06-29 04:17:11 +08:00

Author	SHA1	Message	Date
Ning Xie	6e9cc73f67	[MISC] correct DeviceConfig device field static type analysis (#19699 ) Signed-off-by: Andy Xie <andy.xning@gmail.com>	2025-06-17 17:21:50 -07:00
Ning Xie	c53711bd63	[MISC] correct copy_blocks src_to_dists param type (#19696 ) Signed-off-by: Andy Xie <andy.xning@gmail.com>	2025-06-17 17:21:06 -07:00
Charlie Fu	a44b1c951d	[Feature][ROCm] Add full graph capture support for TritonAttentionBackend (#19158 ) Signed-off-by: charlifu <charlifu@amd.com>	2025-06-17 17:03:06 -04:00
Michael Goin	b447624ee3	[Bugfix] Fix faulty triton importing logic when using Ray for DP (#19734 ) Signed-off-by: mgoin <mgoin64@gmail.com>	2025-06-17 20:59:29 +00:00
Jiayi Yao	cda92307c1	[Misc] Update lmcache connector with the latest connector apis (#19441 ) Signed-off-by: YaoJiayi <120040070@link.cuhk.edu.cn>	2025-06-17 19:57:54 +00:00
Wentao Ye	ffb2cd6b54	[Perf] Optimize `moe_align_block_size` CUDA kernel (#19572 ) Signed-off-by: yewentao256 <zhyanwentao@126.com> Co-authored-by: mgoin <mgoin64@gmail.com>	2025-06-17 11:49:26 -07:00
Isotr0py	ca94d7fa00	[Bugfix] Update multimodel models mapping to fit new checkpoint after Transformers v4.52 (#19151 ) Signed-off-by: Isotr0py <2037008807@qq.com>	2025-06-17 15:58:38 +00:00
CYJiang	5a1c2e15d8	[Mis] remove duplicate engine status checks (#19647 ) Signed-off-by: googs1025 <googs1025@gmail.com>	2025-06-17 08:17:38 -07:00
Nicolò Lucchesi	4c8f64faa7	[V1][Kernel] Flashinfer HND KV cache layout (#19280 ) Signed-off-by: NickLucche <nlucches@redhat.com>	2025-06-17 09:09:22 -04:00
jvlunteren	ccd7c05089	[Kernel] Add Split-KV Support to Unified Triton Attention Kernel (#19152 ) Signed-off-by: Jan van Lunteren <jvl@zurich.ibm.com>	2025-06-17 10:45:07 +00:00
quanliu	5c76b9cdaf	[Core] add remove_seq_from_computed_blocks_tracker to BlockSpaceManager (#19686 ) Signed-off-by: 刘全 <quan.liu2@dbappsecurity.com.cn> Co-authored-by: 刘全 <quan.liu2@dbappsecurity.com.cn>	2025-06-17 04:40:58 +00:00
Driss Guessous	ddfed314f9	Fixes IMA for TP w/ flex-attention (#19712 ) Signed-off-by: drisspg <drisspguessous@gmail.com>	2025-06-17 04:01:50 +00:00
Di Liu	5b3ad5ecf2	[DOC] fix doc typos (#19600 ) Signed-off-by: Di Liu <liu-di@sjtu.edu.cn>	2025-06-17 11:34:53 +08:00
nguyenhoangthuan99	ede5c4ebdf	[Frontend] add chunking audio for > 30s audio (#19597 ) Signed-off-by: nguyenhoangthuan99 <thuanhppro12@gmail.com>	2025-06-17 11:34:00 +08:00
Conroy Cheers	0860087aff	[Fix] Fall back to Gloo when NCCL backend is unavailable (#19641 ) Signed-off-by: conroy-cheers <conroy@corncheese.org>	2025-06-17 08:42:14 +08:00
Dipika Sikka	6bc7b57315	[Quantization] Remove FP4 emulation; Fall-back to marlin for device < 100 (#19563 )	2025-06-16 17:33:51 -04:00
Russell Bryant	90f9c2eb5c	[V1] Change return type on get_multimodal_embeddings() (#19446 ) Signed-off-by: Russell Bryant <rbryant@redhat.com>	2025-06-16 13:32:15 -04:00
qscqesze	387bdf0ab9	[Model] Add support for MiniMaxM1ForCausalLM (shares architecture with MiniMaxText01ForCausalLM) (#19677 ) Signed-off-by: QscQ <qscqesze@gmail.com>	2025-06-16 09:47:14 -07:00
bnellnm	5e5baa91aa	[Kernels] Use empty for modular MoE workspaces (#19667 ) Signed-off-by: Bill Nell <bnell@redhat.com>	2025-06-16 14:58:01 +00:00
Chauncey	836d4ce140	[Bugfix] fix missing 'finish_reason': null in streaming chat (#19662 ) Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com>	2025-06-16 14:10:39 +00:00
Isotr0py	1173804dca	[Bugfix] Fix TP inference for Flex attention backend (#19657 ) Signed-off-by: Isotr0py <2037008807@qq.com>	2025-06-16 11:21:37 +00:00
Shawn Tan	4d5424029b	[Feature]:Allow for Granite MoE Hybrid models with _only_ shared experts. (#19652 ) Signed-off-by: Shawn Tan <shawntan@ibm.com>	2025-06-16 11:14:18 +00:00
Nick Hill	ee35e96ac3	[BugFix] Don't catch BaseException when dumping execute_model errors (#19626 ) Signed-off-by: Nick Hill <nhill@redhat.com>	2025-06-16 11:01:08 +00:00
Szymon Ożóg	dec66d253b	[Kernel] GGUF MMVQ kernel for multiple input vectors (#18754 ) Signed-off-by: SzymonOzog <szymon.ozog@gmail.com>	2025-06-16 17:33:26 +08:00
wang.yuqi	f40f763f12	[CI] Add mteb testing for rerank models (#19344 )	2025-06-16 01:36:43 -07:00
Ning Xie	26bc46ef89	[MISC] typo fix (#19672 ) Signed-off-by: Andy Xie <andy.xning@gmail.com>	2025-06-16 07:18:49 +00:00
Chengji Yao	a77aea59fd	[TPU] support attention head dim smaller than 128 (#19620 ) Signed-off-by: Chengji Yao <chengjiyao@google.com> Co-authored-by: mgoin <mgoin64@gmail.com>	2025-06-16 06:40:53 +00:00
Ye (Charlotte) Qi	b692e9cd07	[Misc] Fix skipped max-model-len validation when deriving max model length from tokenizer config (#19660 ) Signed-off-by: Ye (Charlotte) Qi <yeq@meta.com>	2025-06-16 06:30:29 +00:00
Francesco Bertolotti	367871a469	[Misc][Frontend] passthrough `bad_words` (#19564 ) Signed-off-by: Francesco Bertolotti <francesco.bertolotti@igenius.ai> Co-authored-by: Francesco Bertolotti <francesco.bertolotti@igenius.ai> Co-authored-by: Aaron Pham <Aaronpham0103@gmail.com>	2025-06-16 05:05:13 +00:00
quanliu	92183b41f3	[Bugfix][Core] Prefix caching causes incorrect outputs due to outdated ComputedBlocksTracker (#18957 ) Signed-off-by: 刘全 <quan.liu2@dbappsecurity.com.cn> Co-authored-by: 刘全 <quan.liu2@dbappsecurity.com.cn>	2025-06-15 21:56:37 -07:00
Isotr0py	a5e7242d5f	[Misc] Remove duplicate multiproc method setting for CPU platform (#19649 ) Signed-off-by: Isotr0py <2037008807@qq.com>	2025-06-16 02:26:58 +00:00
Woosuk Kwon	055915e6ce	Enable prefix caching with full cuda graphs (#19617 ) Signed-off-by: Woosuk Kwon <woosuk.kwon@berkeley.edu>	2025-06-15 01:05:05 -07:00
22quinn	0b73736a0d	[Kernel] Raise verbose error and consolidate `num_heads/num_kv_heads` divisibility check (#19339 ) Signed-off-by: 22quinn <33176974+22quinn@users.noreply.github.com>	2025-06-15 13:43:48 +08:00
Lu Fang	ee1531bc38	[Bugfix][2/n] Fix speculative decoding CI - Fix test_ngram_e2e_greedy_correctness (#19644 )	2025-06-14 21:15:41 -07:00
maobaolong	08500011d3	[Fix] Convert kv_transfer_config from dict to KVTransferConfig (#19262 )	2025-06-14 12:32:07 -07:00
Konrad Zawora	861a0a0a39	[Bugfix] Don't attempt to use triton if no driver is active (#19561 )	2025-06-14 12:30:54 -07:00
Isotr0py	2db9044ab6	[Bugfix] Fix auto dtype casting for BatchFeature (#19316 ) Signed-off-by: Isotr0py <2037008807@qq.com> Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn>	2025-06-14 15:13:08 +00:00
Saheli Bhattacharjee	d1e34cc9ac	[V1][Metrics] Deprecate metrics with gpu_ prefix for non GPU specific metrics. (#18354 ) Signed-off-by: Saheli Bhattacharjee <saheli@krai.ai>	2025-06-14 11:07:36 +08:00
Nick Hill	bd517eb9fe	[BugFix] Fix DP Coordinator incorrect debug log message (#19624 ) Signed-off-by: Nick Hill <nhill@redhat.com>	2025-06-14 00:18:03 +00:00
Woosuk Kwon	aafbbd981f	[torch.compile] Use custom ops when use_inductor=False (#19618 )	2025-06-13 15:05:54 -07:00
Luka Govedič	3597b06a4f	[CUDA] Enable full cudagraph for FlashMLA (#18581 ) Signed-off-by: luka <luka@neuralmagic.com>	2025-06-13 18:12:26 +00:00
qscqesze	a24cb91600	[Model] Fix minimax model cache & lm_head precision (#19592 ) Signed-off-by: qingjun <qingjun@minimaxi.com>	2025-06-13 12:08:20 +00:00
Nick Hill	7e8d97dd3f	[BugFix] Honor `enable_caching` in connector-delayed kvcache load case (#19435 ) Signed-off-by: Nick Hill <nhill@redhat.com>	2025-06-13 09:46:32 +00:00
youkaichao	d70bc7c029	[torch.compile] reorganize the cache directory to support compiling multiple models (#19064 ) Signed-off-by: youkaichao <youkaichao@gmail.com>	2025-06-13 15:23:25 +08:00
Boyuan Feng	ce688ad46e	use base version for version comparison (#19587 ) Signed-off-by: Boyuan Feng <boyuan@meta.com>	2025-06-13 15:09:34 +08:00
汪志鹏	cefdb9962d	[Fix] The zip function in Python 3.9 does not have the strict argument (#19549 ) Signed-off-by: 汪志鹏 <wangzhipeng628@gmail.com>	2025-06-13 14:57:48 +08:00
Li, Jiang	6458721108	[CPU] Refine default config for the CPU backend (#19539 ) Signed-off-by: jiang1.li <jiang1.li@intel.com>	2025-06-13 13:27:39 +08:00
Hyogeun Oh (오효근)	bb4a0decef	[Misc] Correct broken docs link (#19553 ) Signed-off-by: Zerohertz <ohg3417@gmail.com>	2025-06-12 22:27:13 -07:00
qizixi	c68698b326	[Bugfix] Fix EAGLE vocab embedding for multimodal target model (#19570 ) Signed-off-by: qizixi <qizixi@meta.com>	2025-06-12 23:09:19 -04:00
Varun Sundar Rabindranath	e3b12667d4	[BugFix] : Fix Batched DeepGemm Experts (#19515 ) Signed-off-by: Varun Sundar Rabindranath <vsundarr@redhat.com> Co-authored-by: Varun Sundar Rabindranath <vsundarr@redhat.com>	2025-06-12 20:43:02 -06:00

1 2 3 4 5 ...

4848 Commits