xinyun/vllm - vllm - 丝路新云-代码仓

mirror of https://git.datalinker.icu/vllm-project/vllm.git synced 2025-12-13 11:15:35 +08:00

Author	SHA1	Message	Date
Shawn Du	f8ece6e17f	[Core][v1] Unify allocating slots in prefill and decode in KV cache manager (#12608 ) As mentioned in RFC https://github.com/vllm-project/vllm/issues/12254, this PR achieves the task: combine allocate_slots and append_slots. There should be no functionality change, except that in decode, also raise exception when num_tokens is zero (like prefill), and change the unit test case accordingly. @comaniac @rickyyx @WoosukKwon @youkaichao @heheda12345 @simon-mo --------- Signed-off-by: Shawn Du <shawnd200@outlook.com>	2025-02-02 16:40:58 +08:00
Mark McLoughlin	f17f1d4608	[V1][Metrics] Add GPU cache usage % gauge (#12561 ) Signed-off-by: Mark McLoughlin <markmc@redhat.com>	2025-01-29 18:31:01 -08:00
Woosuk Kwon	e0cc5f259a	[V1][BugFix] Free encoder cache for aborted requests (#12545 ) Signed-off-by: Woosuk Kwon <woosuk.kwon@berkeley.edu>	2025-01-29 13:47:33 -08:00
Harry Mellor	823ab79633	Update `pre-commit` hooks (#12475 ) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>	2025-01-27 17:23:08 -07:00
Woosuk Kwon	624a1e4711	[V1][Minor] Minor optimizations for update_from_output (#12454 ) Signed-off-by: Woosuk Kwon <woosuk.kwon@berkeley.edu>	2025-01-27 01:09:27 -08:00
Cody Yu	7206ce4ce1	[Core] Support `reset_prefix_cache` (#12284 )	2025-01-22 18:52:27 +00:00
Roger Wang	70755e819e	[V1][Core] Autotune encoder cache budget (#11895 ) Signed-off-by: Roger Wang <ywang@roblox.com>	2025-01-15 11:29:00 -08:00
Chen Zhang	994fc655b7	[V1][Prefix Cache] Move the logic of num_computed_tokens into KVCacheManager (#12003 )	2025-01-15 07:55:30 +00:00
Woosuk Kwon	b7ee940a82	[V1][BugFix] Fix edge case in VLM scheduling (#12065 ) Signed-off-by: Woosuk Kwon <woosuk.kwon@berkeley.edu>	2025-01-14 20:21:28 -08:00
Robert Shaw	9597a095f2	[V1][Core][1/n] Logging and Metrics (#11962 ) Signed-off-by: rshaw@neuralmagic.com <rshaw@neuralmagic.com>	2025-01-12 21:02:02 +00:00
WangErXiao	dc71af0a71	Remove the duplicate imports of MultiModalKwargs and PlaceholderRange… (#11824 )	2025-01-08 04:09:25 +00:00
Woosuk Kwon	73001445fb	[V1] Implement Cascade Attention (#11635 ) Signed-off-by: Woosuk Kwon <woosuk.kwon@berkeley.edu>	2025-01-01 21:56:46 +09:00
Cody Yu	bf8717ebae	[V1] Prefix caching for vision language models (#11187 ) Signed-off-by: Cody Yu <hao.yu.cody@gmail.com>	2024-12-17 16:37:59 -08:00
Roger Wang	59c9b6ebeb	[V1][VLM] Proper memory profiling for image language models (#11210 ) Signed-off-by: Roger Wang <ywang@roblox.com> Co-authored-by: ywang96 <ywang@example.com>	2024-12-16 22:10:57 -08:00
Mark McLoughlin	6d917d0eeb	Enable mypy checking on V1 code (#11105 ) Signed-off-by: Mark McLoughlin <markmc@redhat.com>	2024-12-14 09:54:04 -08:00
Cody Yu	9855aea21b	[Bugfix][V1] Re-compute an entire block when fully cache hit (#11186 ) Signed-off-by: Cody Yu <hao.yu.cody@gmail.com>	2024-12-13 17:08:23 -08:00
Tyler Michael Smith	28b3a1c7e5	[V1] Multiprocessing Tensor Parallel Support for v1 (#9856 ) Signed-off-by: Tyler Michael Smith <tyler@neuralmagic.com>	2024-12-10 06:28:14 +00:00
Roger Wang	a11f326528	[V1] Initial support of multimodal models for V1 re-arch (#10699 ) Signed-off-by: Roger Wang <ywang@roblox.com>	2024-12-08 12:50:51 +00:00
Woosuk Kwon	a79b122400	[V1] Do not allocate beyond the max_model_len (#10730 ) Signed-off-by: Woosuk Kwon <woosuk.kwon@berkeley.edu>	2024-11-28 00:13:15 -08:00
Woosuk Kwon	bbd3e86926	[V1] Support VLMs with fine-grained scheduling (#9871 ) Signed-off-by: Woosuk Kwon <woosuk.kwon@berkeley.edu> Co-authored-by: Roger Wang <ywang@roblox.com>	2024-11-13 04:53:13 +00:00
Robert Shaw	6ace6fba2c	[V1] `AsyncLLM` Implementation (#9826 ) Signed-off-by: Nick Hill <nickhill@us.ibm.com> Signed-off-by: rshaw@neuralmagic.com <rshaw@neuralmagic.com> Signed-off-by: Nick Hill <nhill@redhat.com> Co-authored-by: Nick Hill <nickhill@us.ibm.com> Co-authored-by: Varun Sundar Rabindranath <varun@neuralmagic.com> Co-authored-by: Nick Hill <nhill@redhat.com> Co-authored-by: Tyler Michael Smith <tyler@neuralmagic.com>	2024-11-11 23:05:38 +00:00
Cody Yu	201fc07730	[V1] Prefix caching (take 2) (#9972 ) Signed-off-by: Cody Yu <hao.yu.cody@gmail.com>	2024-11-07 17:34:44 -08:00
Woosuk Kwon	42b4f46b71	[V1] Add all_token_ids attribute to Request (#10135 ) Signed-off-by: Woosuk Kwon <woosuk.kwon@berkeley.edu>	2024-11-07 17:08:24 -08:00
Woosuk Kwon	6c5af09b39	[V1] Implement vLLM V1 [1/N] (#9289 )	2024-10-22 01:24:07 -07:00

24 Commits