xinyun/vllm - vllm - 丝路新云-代码仓

mirror of https://git.datalinker.icu/vllm-project/vllm.git synced 2026-06-07 15:02:18 +08:00

Author	SHA1	Message	Date
yansh97	94d545a1a1	[Doc] Fix typo in the help message of '--guided-decoding-backend' (#11440 )	2024-12-23 20:20:44 +00:00
Ricky Xu	584f0ae40d	[V1] Make AsyncLLMEngine v1-v0 opaque (#11383 ) Signed-off-by: Ricky Xu <xuchen727@hotmail.com>	2024-12-21 15:14:08 +08:00
omer-dayan	995f56236b	[Core] Loading model from S3 using RunAI Model Streamer as optional loader (#10192 ) Signed-off-by: OmerD <omer@run.ai>	2024-12-20 16:46:24 +00:00
Yanyi Liu	5aef49806d	[Feature] Add load generation config from model (#11164 ) Signed-off-by: liuyanyi <wolfsonliu@163.com> Signed-off-by: Yanyi Liu <wolfsonliu@163.com> Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk> Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>	2024-12-19 10:50:38 +00:00
Alexander Matveev	fdea8ec167	[V1] VLM - enable processor cache by default (#11305 ) Signed-off-by: Alexander Matveev <alexm@neuralmagic.com>	2024-12-18 18:54:46 -05:00
Konrad Zawora	866fa4550d	[Bugfix] Restore support for larger block sizes (#11259 ) Signed-off-by: Konrad Zawora <kzawora@habana.ai>	2024-12-17 16:39:07 -08:00
Cody Yu	bf8717ebae	[V1] Prefix caching for vision language models (#11187 ) Signed-off-by: Cody Yu <hao.yu.cody@gmail.com>	2024-12-17 16:37:59 -08:00
Joe Runde	2d1b9baa8f	[Bugfix] Fix request cancellation without polling (#11190 )	2024-12-17 12:26:32 -08:00
wangxiyuan	e88db68cf5	[Platform] platform agnostic for EngineArgs initialization (#11225 ) Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>	2024-12-16 22:11:06 -08:00
youkaichao	551603feff	[core] overhaul memory profiling and fix backward compatibility (#10511 ) Signed-off-by: youkaichao <youkaichao@gmail.com>	2024-12-16 13:32:25 -08:00
chenqianfzh	69ba344de8	[Bugfix] Fix block size validation (#10938 )	2024-12-15 16:38:40 -08:00
Brad Hilton	9c3dadd1c9	[Frontend] Add `logits_processors` as an extra completion argument (#11150 ) Signed-off-by: Brad Hilton <brad.hilton.nw@gmail.com>	2024-12-14 16:46:42 +00:00
Cyrus Leung	eeec9e3390	[Frontend] Separate pooling APIs in offline inference (#11129 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>	2024-12-13 10:40:07 +00:00
Gregory Shtrasberg	00c1bde5d8	[ROCm][AMD] Disable auto enabling chunked prefill on ROCm (#11146 ) Signed-off-by: Gregory Shtrasberg <Gregory.Shtrasberg@amd.com>	2024-12-13 05:31:26 +00:00
Jeremy Arnold	9f3974a319	Fix logging of the vLLM Config (#11143 )	2024-12-12 12:05:57 -08:00
Alexander Matveev	4e11683368	[V1] VLM preprocessor hashing (#11020 ) Signed-off-by: Roger Wang <ywang@roblox.com> Signed-off-by: Alexander Matveev <alexm@neuralmagic.com> Co-authored-by: Michael Goin <michael@neuralmagic.com> Co-authored-by: Roger Wang <ywang@roblox.com>	2024-12-12 00:55:30 +00:00
Cyrus Leung	cad5c0a6ed	[Doc] Update docs to refer to pooling models (#11093 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>	2024-12-11 13:36:27 +00:00
Cyrus Leung	8f10d5e393	[Misc] Split up pooling tasks (#10820 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>	2024-12-11 01:28:00 -08:00
Woosuk Kwon	134810b3d9	[V1][Bugfix] Always set enable_chunked_prefill = True for V1 (#11061 ) Signed-off-by: Woosuk Kwon <woosuk.kwon@berkeley.edu>	2024-12-10 14:41:23 -08:00
Joe Runde	9b9cef3145	[Bugfix] Backport request id validation to v0 (#11036 ) Signed-off-by: Joe Runde <Joseph.Runde@ibm.com>	2024-12-10 16:38:23 +00:00
youkaichao	ebf778061d	monitor metrics of tokens per step using cudagraph batchsizes (#11031 ) Signed-off-by: youkaichao <youkaichao@gmail.com>	2024-12-09 22:35:36 -08:00
Cyrus Leung	391d7b2763	[Bugfix] Fix usage of `deprecated` decorator (#11025 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>	2024-12-10 13:45:47 +08:00
youkaichao	46004e83a2	[misc] clean up and unify logging (#10999 ) Signed-off-by: youkaichao <youkaichao@gmail.com>	2024-12-08 17:28:27 -08:00
Roger Wang	a11f326528	[V1] Initial support of multimodal models for V1 re-arch (#10699 ) Signed-off-by: Roger Wang <ywang@roblox.com>	2024-12-08 12:50:51 +00:00
youkaichao	fd57d2b534	[torch.compile] allow candidate compile sizes (#10984 ) Signed-off-by: youkaichao <youkaichao@gmail.com>	2024-12-08 11:05:21 +00:00
Russell Bryant	69d357ba12	[Core] Cleanup startup logging a bit (#10961 ) Signed-off-by: Russell Bryant <rbryant@redhat.com>	2024-12-07 02:30:23 +00:00
youkaichao	b031a455a9	[torch.compile] add logging for compilation time (#10941 ) Signed-off-by: youkaichao <youkaichao@gmail.com> Co-authored-by: Woosuk Kwon <woosuk.kwon@berkeley.edu>	2024-12-06 10:07:15 +00:00
Cyrus Leung	aa39a8e175	[Doc] Create a new "Usage" section (#10827 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>	2024-12-05 11:19:35 +08:00
Xin Yang	01d079fd8e	[LoRA] Change lora_tokenizers capacity (#10796 ) Signed-off-by: Xin Yang <xyang19@gmail.com>	2024-12-04 17:40:16 +00:00
tomeras91	7c32b6861e	[Frontend] correctly record prefill and decode time metrics (#10853 ) Signed-off-by: Tomer Asida <tomera@ai21.com>	2024-12-03 19:13:31 +00:00
Aaron Pham	9323a3153b	[Core][Performance] Add XGrammar support for guided decoding and set it as default (#10785 ) Signed-off-by: Aaron Pham <contact@aarnphm.xyz> Signed-off-by: mgoin <michael@neuralmagic.com> Co-authored-by: mgoin <michael@neuralmagic.com>	2024-12-03 15:17:00 +08:00
Cyrus Leung	3257d449fa	[Misc] Remove deprecated names (#10817 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>	2024-12-03 06:52:57 +00:00
cduk	b7954776fd	[core] Avoid metrics log noise when idle - include speculative decodi… (#10809 )	2024-12-02 01:49:48 +00:00
Kuntai Du	0590ec3fd9	[Core] Implement disagg prefill by StatelessProcessGroup (#10502 ) This PR provides initial support for single-node disaggregated prefill in 1P1D scenario. Signed-off-by: KuntaiDu <kuntai@uchicago.edu> Co-authored-by: ApostaC <yihua98@uchicago.edu> Co-authored-by: YaoJiayi <120040070@link.cuhk.edu.cn>	2024-12-01 19:01:00 -06:00
Cyrus Leung	d2f058e76c	[Misc] Rename embedding classes to pooling (#10801 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>	2024-12-01 14:36:51 +08:00
Ricky Xu	d9b4b3f069	[Bug][CLI] Allow users to disable prefix caching explicitly (#10724 ) Signed-off-by: rickyx <rickyx@anyscale.com>	2024-11-27 23:59:28 -08:00
Mor Zusman	197b4484a3	[Bugfix][Mamba] Fix Multistep on Mamba-like models (#10705 ) Signed-off-by: mzusman <mor.zusmann@gmail.com>	2024-11-27 19:02:27 +00:00
Michael Goin	9a99273b48	[Bugfix] Fix using `-O[0,3]` with LLM entrypoint (#10677 ) Signed-off-by: mgoin <michael@neuralmagic.com>	2024-11-26 10:44:01 -08:00
Ricky Xu	519e8e4182	[v1] EngineArgs for better config handling for v1 (#10382 ) Signed-off-by: rickyx <rickyx@anyscale.com>	2024-11-25 21:09:43 -08:00
Wallas Henrique	c27df94e1f	[Bugfix] Fix chunked prefill with model dtype float32 on Turing Devices (#9850 ) Signed-off-by: Wallas Santos <wallashss@ibm.com> Co-authored-by: Michael Goin <michael@neuralmagic.com>	2024-11-25 12:23:32 -05:00
youkaichao	25d806e953	[misc] add torch.compile compatibility check (#10618 ) Signed-off-by: youkaichao <youkaichao@gmail.com>	2024-11-24 23:40:08 -08:00
Russell Bryant	28598f3939	[Core] remove temporary local variables in LLMEngine.__init__ (#10577 ) Signed-off-by: Russell Bryant <rbryant@redhat.com>	2024-11-22 16:22:53 -08:00
youkaichao	a111d0151f	[platforms] absorb worker cls difference into platforms folder (#10555 ) Signed-off-by: youkaichao <youkaichao@gmail.com> Co-authored-by: Nick Hill <nhill@redhat.com>	2024-11-21 21:00:32 -08:00
youkaichao	7560ae5caf	[8/N] enable cli flag without a space (#10529 ) Signed-off-by: youkaichao <youkaichao@gmail.com>	2024-11-21 12:30:42 -08:00
Cyrus Leung	b4be5a8adb	[Bugfix] Enforce no chunked prefill for embedding models (#10470 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>	2024-11-20 05:12:51 +00:00
Russell Bryant	efa9084628	[Core] Avoid metrics log noise when idle (#8868 ) Signed-off-by: Russell Bryant <rbryant@redhat.com>	2024-11-19 21:05:25 +00:00
youkaichao	803f37eaaa	[6/N] torch.compile rollout to users (#10437 ) Signed-off-by: youkaichao <youkaichao@gmail.com>	2024-11-19 10:09:03 -08:00
Russell Bryant	5390d6664f	[Doc] Add the start of an arch overview page (#10368 )	2024-11-19 09:52:11 +00:00
Travis Johnson	272e31c0bd	[Bugfix] Guard for negative counter metrics to prevent crash (#10430 ) Signed-off-by: Travis Johnson <tsjohnso@us.ibm.com>	2024-11-19 04:57:10 +00:00
Cyrus Leung	32e46e000f	[Frontend] Automatic detection of chat content format from AST (#9919 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>	2024-11-16 13:35:40 +08:00

... 3 4 5 6 7 ...

692 Commits