xinyun/vllm - vllm - 丝路新云-代码仓

mirror of https://git.datalinker.icu/vllm-project/vllm.git synced 2025-12-15 12:05:46 +08:00

Author	SHA1	Message	Date
Nicolò Lucchesi	3cc9af88ff	[TPU][V1] Disable per-request seed/Generator (#16172 ) Signed-off-by: NickLucche <nlucches@redhat.com>	2025-04-10 17:05:44 -04:00
Nicolò Lucchesi	c1b57855ec	[TPU][V1] Use `language_model` interface for getting text backbone in MM (#16410 ) Signed-off-by: NickLucche <nlucches@redhat.com>	2025-04-10 17:32:04 +00:00
Chengji Yao	a454748544	[TPU][V1] Refine tpu_model_runner to mitigate future recompilation issues (#16275 ) Signed-off-by: Chengji Yao <chengjiyao@google.com>	2025-04-09 18:51:51 -06:00
yihong	04149cce27	[BugFix] fix some typos found by typos. (#16314 ) Signed-off-by: yihong0618 <zouzou0208@gmail.com>	2025-04-09 03:43:59 -07:00
Siyuan Liu	87918e40c4	[torch.compile][TPU] Make @support_torch_compile work for XLA backend (#15782 ) Signed-off-by: Siyuan Liu <lsiyuan@google.com> Signed-off-by: mgoin <mgoin64@gmail.com> Co-authored-by: mgoin <mgoin64@gmail.com>	2025-04-08 14:23:53 +08:00
Roger Wang	f2ebb6f541	[V1] Scatter and gather placeholders in the model runner (#16076 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk> Signed-off-by: mgoin <mgoin64@gmail.com> Signed-off-by: Roger Wang <ywang@roblox.com> Co-authored-by: DarkLight1337 <tlleungac@connect.ust.hk> Co-authored-by: mgoin <mgoin64@gmail.com> Co-authored-by: Jennifer Zhao <ai.jenniferzhao@gmail.com>	2025-04-08 10:43:41 +08:00
Roger Wang	af51d80fa1	Revert "[V1] Scatter and gather placeholders in the model runner" (#16075 )	2025-04-04 14:50:57 -07:00
Cyrus Leung	f5722a5052	[V1] Scatter and gather placeholders in the model runner (#15712 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk> Signed-off-by: mgoin <mgoin64@gmail.com> Signed-off-by: Roger Wang <ywang@roblox.com> Co-authored-by: mgoin <mgoin64@gmail.com> Co-authored-by: Roger Wang <ywang@roblox.com>	2025-04-04 21:26:44 +00:00
Chengji Yao	fadc59c0e6	[TPU][V1] Remove ragged attention kernel parameter hard coding (#16041 ) Signed-off-by: Chengji Yao <chengjiyao@google.com>	2025-04-04 07:48:50 -04:00
Nicolò Lucchesi	bd7599d34a	[V1][TPU] Do not compile sampling more than needed (#15883 ) Signed-off-by: NickLucche <nlucches@redhat.com>	2025-04-03 01:36:01 +00:00
Chen Zhang	3a5f0afcd2	[V1] Implement sliding window attention in kv_cache_manager (#14097 ) Signed-off-by: Chen Zhang <zhangch99@outlook.com>	2025-04-01 00:33:17 -07:00
Alexander Matveev	9a2160fa55	[V1] TPU CI - Add basic perf regression test (#15414 ) Signed-off-by: Alexander Matveev <amatveev@redhat.com>	2025-03-31 13:25:20 -04:00
Cyrus Leung	09e974d483	[Bugfix] Check dimensions of multimodal embeddings in V1 (#15816 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>	2025-03-31 09:01:35 -07:00
yarongmu-google	7c1f760024	[Kernel][TPU][ragged-paged-attn] vLLM code change for PR#8896 (#15659 ) Signed-off-by: Yarong Mu <ymu@google.com>	2025-03-28 21:13:15 -07:00
Nicolò Lucchesi	da461f3cbf	[TPU][V1][Bugfix] Fix w8a8 recompiilation with GSM8K (#15714 ) Signed-off-by: NickLucche <nlucches@redhat.com>	2025-03-28 21:13:06 -07:00
Alexander Matveev	c3f687ac22	[V1] TPU - Fix the chunked prompt bug (#15713 ) Signed-off-by: Alexander Matveev <amatveev@redhat.com>	2025-03-28 20:19:04 +00:00
Cyrus Leung	355f66348c	[V1] Remove legacy input registry (#15673 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>	2025-03-27 23:34:34 -07:00
Nicolò Lucchesi	4098b72210	[Bugfix][TPU][V1] Fix recompilation (#15553 ) Signed-off-by: NickLucche <nlucches@redhat.com>	2025-03-27 19:15:06 +00:00
Cyrus Leung	13ac9cab21	[Misc] Avoid direct access of global `mm_registry` in `compute_encoder_budget` (#15621 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>	2025-03-27 17:52:00 +00:00
Chengji Yao	619d3de8bd	[TPU] [V1] fix cases when max_num_reqs is set smaller than MIN_NUM_SEQS (#15583 ) Signed-off-by: Chengji Yao <chengjiyao@google.com>	2025-03-26 22:46:26 -07:00
Alexander Matveev	b2e85e26f4	[V1] TPU - Revert to exponential padding by default (#15565 ) Signed-off-by: Alexander Matveev <amatveev@redhat.com>	2025-03-26 21:35:05 +00:00
Chenyaaang	ac3cd6e83c	[core] add bucket padding to tpu_model_runner (#14995 ) Signed-off-by: Chenyaaang <llccyy1212@gmail.com> Signed-off-by: rshaw@neuralmagic.com <robertgshaw2@gmail.com> Co-authored-by: rshaw@neuralmagic.com <robertgshaw2@gmail.com>	2025-03-25 17:27:22 -04:00
Nicolò Lucchesi	a0dd7dcd49	[TPU][V1] Fix Sampler recompilation (#15309 ) Signed-off-by: NickLucche <nlucches@redhat.com>	2025-03-25 16:43:54 -04:00
Chen Zhang	93a00d7dde	[v1] Refactor KVCacheConfig (#14079 ) Signed-off-by: Chen Zhang <zhangch99@outlook.com>	2025-03-21 04:56:27 -07:00
Siyuan Liu	b15fd2be2a	[Hardware][TPU] Add check for no additional graph compilation during runtime (#14710 ) Signed-off-by: Siyuan Liu <lsiyuan@google.com>	2025-03-21 03:05:28 +00:00
Woosuk Kwon	0c6f5023c3	[V1] Scheduler Refactoring [1/N] - Add Scheduler Interface (#15250 ) Signed-off-by: Woosuk Kwon <woosuk.kwon@berkeley.edu> Co-authored-by: Cody Yu <hao.yu.cody@gmail.com> Co-authored-by: Nick Hill <nhill@redhat.com>	2025-03-20 17:50:43 -07:00
Nicolò Lucchesi	d8c6d7d6b5	[V1][TPU] Support V1 Sampler for ragged attention (#14227 ) Signed-off-by: NickLucche <nlucches@redhat.com>	2025-03-19 21:00:39 -07:00
Nicolò Lucchesi	af35d3a3cc	[TPU][V1][Bugfix] Fix chunked prefill with padding (#15037 ) Signed-off-by: NickLucche <nlucches@redhat.com>	2025-03-18 07:34:45 -07:00
iefgnoix	b4ad56c1bd	[V1][TPU] Apply the ragged paged attention kernel fix and remove the padding. (#14846 ) Signed-off-by: Xiongfei Wei <isaacwxf23@gmail.com>	2025-03-17 01:48:28 -07:00
iefgnoix	863d315c86	[V1][TPU] Pad the block_table.shape[1] so the ragged paged attention can handle correctly (#14597 )	2025-03-11 19:12:26 -04:00
Chengji Yao	212007b168	[Hardware][TPU] Fix the recompiling issue in logits processor after warmup (#14510 ) Signed-off-by: Chengji Yao <chengjiyao@google.com>	2025-03-09 05:44:39 -04:00
iefgnoix	10f7552789	[V1][TPU] Remove unnecessary padding for running on TPU. (#14467 )	2025-03-08 21:56:04 -05:00
Alexander Matveev	cb8bdfade2	[V1] TPU - Add tensor parallel support via Ray (#13618 ) Signed-off-by: Alexander Matveev <amatveev@redhat.com>	2025-03-08 08:19:38 -05:00
Tyler Michael Smith	333681408f	[Bugfix][V1] Handle MLA in kv_cache_interface (#14462 ) Signed-off-by: Tyler Michael Smith <tyler@neuralmagic.com>	2025-03-07 22:18:25 -08:00
Nick Hill	8ed5421aaa	[V1] Eagerly remove finished requests from the batch (#14388 ) Signed-off-by: Nick Hill <nhill@redhat.com>	2025-03-07 10:56:00 -08:00
Chengji Yao	0578e5a462	[Hardware][TPU]Enable ragged paged attention kernel and resolve recompilation issue (#14310 ) Signed-off-by: Chengji Yao <chengjiyao@google.com>	2025-03-06 23:31:05 +00:00
Michael Goin	fbfc3ee37e	[V1][TPU] TPU multimodal model support for ragged attention (#14158 ) Signed-off-by: Michael Goin <mgoin64@gmail.com>	2025-03-04 19:58:48 -05:00
Harry Mellor	cf069aa8aa	Update deprecated Python 3.8 typing (#13971 )	2025-03-02 17:34:51 -08:00
Chen Zhang	e7bd944e08	[v1] Cleanup the BlockTable in InputBatch (#13977 ) Signed-off-by: Chen Zhang <zhangch99@outlook.com>	2025-02-28 19:03:16 +00:00
iefgnoix	c3b6559a10	[V1][TPU] Integrate the new ragged paged attention kernel with vLLM v1 on TPU (#13379 ) Signed-off-by: Xiongfei Wei <isaacwxf23@gmail.com> Signed-off-by: mgoin <mgoin64@gmail.com> Co-authored-by: mgoin <mgoin64@gmail.com>	2025-02-28 11:01:36 -07:00
Harry Mellor	cdc1fa12eb	Remove unused kwargs from model definitions (#13555 )	2025-02-24 17:13:52 -08:00
Nick Hill	30172b4947	[V1] Optimize handling of sampling metadata and req_ids list (#13244 ) Signed-off-by: Nick Hill <nhill@redhat.com>	2025-02-18 12:15:33 -08:00
Woosuk Kwon	cd4a72a28d	[V1][Spec decode] Move drafter to model runner (#13363 ) Signed-off-by: Woosuk Kwon <woosuk.kwon@berkeley.edu>	2025-02-17 15:40:12 -08:00
Lily Liu	80f63a3966	[V1][Spec Decode] Ngram Spec Decode (#12193 ) Signed-off-by: LiuXiaoxuanPKU <lilyliupku@gmail.com>	2025-02-15 18:05:11 -08:00
Alexander Matveev	45f90bcbba	[WIP] TPU V1 Support Refactored (#13049 )	2025-02-14 00:21:53 -08:00

1 2

95 Commits