xinyun/vllm - vllm - 丝路新云-代码仓

mirror of https://git.datalinker.icu/vllm-project/vllm.git synced 2025-12-10 00:06:06 +08:00

Author	SHA1	Message	Date
Alexandre Payot	937e7b7d7c	Build docker image with shared objects from "build" step (#2237 )	2024-01-04 09:35:18 -08:00
ljss	aee8ef661a	Miner fix of type hint (#2340 )	2024-01-03 21:27:56 -08:00
Woosuk Kwon	2e0b6e7757	Bump up to v0.2.7 (#2337 ) v0.2.7	2024-01-03 17:35:56 -08:00
Woosuk Kwon	941767127c	Revert the changes in test_cache (#2335 )	2024-01-03 17:32:05 -08:00
Ronen Schaffer	74d8d77626	Remove unused const TIMEOUT_TO_PREVENT_DEADLOCK (#2321 )	2024-01-03 15:49:07 -08:00
Zhuohan Li	fd4ea8ef5c	Use NCCL instead of ray for control-plane communication to remove serialization overhead (#2221 )	2024-01-03 11:30:22 -08:00
Ronen Schaffer	1066cbd152	Remove deprecated parameter: concurrency_count (#2315 )	2024-01-03 09:56:21 -08:00
Woosuk Kwon	6ef00b03a2	Enable CUDA graph for GPTQ & SqueezeLLM (#2318 )	2024-01-03 09:52:29 -08:00
Roy	9140561059	[Minor] Fix typo and remove unused code (#2305 )	2024-01-02 19:23:15 -08:00
Jee Li	77af974b40	[FIX] Support non-zero CUDA devices in custom kernels (#1959 )	2024-01-02 19:09:59 -08:00
Jong-hun Shin	4934d49274	Support GPT-NeoX Models without attention biases (#2301 )	2023-12-30 11:42:04 -05:00
Zhuohan Li	358c328d69	[BUGFIX] Fix communication test (#2285 )	2023-12-27 17:18:11 -05:00
Zhuohan Li	4aaafdd289	[BUGFIX] Fix the path of test prompts (#2273 )	2023-12-26 10:37:21 -08:00
Zhuohan Li	66b108d142	[BUGFIX] Fix API server test (#2270 )	2023-12-26 10:37:06 -08:00
Zhuohan Li	e0ff920001	[BUGFIX] Do not return ignored sentences twice in async llm engine (#2258 )	2023-12-26 13:41:09 +08:00
blueceiling	face83c7ec	[Docs] Add "About" Heading to README.md (#2260 )	2023-12-25 16:37:07 -08:00
Shivam Thakkar	1db83e31a2	[Docs] Update installation instructions to include CUDA 11.8 xFormers (#2246 )	2023-12-22 23:20:02 -08:00
Woosuk Kwon	a1b9cb2a34	[BugFix] Fix recovery logic for sequence group (#2186 )	2023-12-20 21:52:37 -08:00
Woosuk Kwon	3a4fd5ca59	Disable Ray usage stats collection (#2206 )	2023-12-20 21:52:08 -08:00
Ronen Schaffer	c17daa9f89	[Docs] Fix broken links (#2222 )	2023-12-20 12:43:42 -08:00
Antoni Baum	bd29cf3d3a	Remove Sampler copy stream (#2209 )	2023-12-20 00:04:33 -08:00
Hanzhi Zhou	31bff69151	Make _prepare_sample non-blocking and use pinned memory for input buffers (#2207 )	2023-12-19 16:52:46 -08:00
Woosuk Kwon	ba4f826738	[BugFix] Fix weight loading for Mixtral with TP (#2208 )	2023-12-19 16:16:11 -08:00
avideci	de60a3fb93	Added DeciLM-7b and DeciLM-7b-instruct (#2062 )	2023-12-19 02:29:33 -08:00
Woosuk Kwon	21d5daa4ac	Add warning on CUDA graph memory usage (#2182 )	2023-12-18 18:16:17 -08:00
Suhong Moon	290e015c6c	Update Help Text for --gpu-memory-utilization Argument (#2183 )	2023-12-18 11:33:24 -08:00
kliuae	1b7c791d60	[ROCm] Fixes for GPTQ on ROCm (#2180 )	2023-12-18 10:41:04 -08:00
JohnSaxon	bbe4466fd9	[Minor] Fix typo (#2166 ) Co-authored-by: John-Saxon <zhang.xiangxuan@oushu.com>	2023-12-17 23:28:49 -08:00
Harry Mellor	08133c4d1a	Add SSL arguments to API servers (#2109 )	2023-12-18 10:56:23 +08:00
Woosuk Kwon	76a7983b23	[BugFix] Fix RoPE kernel on long sequences(#2164 )	2023-12-17 17:09:10 -08:00
Woosuk Kwon	8041b7305e	[BugFix] Raise error when max_model_len is larger than KV cache (#2163 )	2023-12-17 17:08:23 -08:00
Suhong Moon	3ec8c25cd0	[Docs] Update documentation for gpu-memory-utilization option (#2162 )	2023-12-17 10:51:57 -08:00
Woosuk Kwon	671af2b1c0	Bump up to v0.2.6 (#2157 ) v0.2.6	2023-12-17 10:34:56 -08:00
Woosuk Kwon	6f41f0e377	Disable CUDA graph for SqueezeLLM (#2161 )	2023-12-17 10:24:25 -08:00
Woosuk Kwon	2c9b638065	[Minor] Fix a typo in .pt weight support (#2160 )	2023-12-17 10:12:44 -08:00
Antoni Baum	a7347d9a6d	Make sampler less blocking (#1889 )	2023-12-17 23:03:49 +08:00
Woosuk Kwon	f8c688d746	[Minor] Add Phi 2 to supported models (#2159 )	2023-12-17 02:54:57 -08:00
Woosuk Kwon	c9fadda543	[Minor] Fix xformers version (#2158 )	2023-12-17 02:28:02 -08:00
Woosuk Kwon	30fb0956df	[Minor] Add more detailed explanation on `quantization` argument (#2145 )	2023-12-17 01:56:16 -08:00
Woosuk Kwon	3a765bd5e1	Temporarily enforce eager mode for GPTQ models (#2154 )	2023-12-17 01:51:12 -08:00
Woosuk Kwon	26c52a5ea6	[Docs] Add CUDA graph support to docs (#2148 )	2023-12-17 01:49:20 -08:00
Woosuk Kwon	c3372e87be	Remove dependency on CuPy (#2152 )	2023-12-17 01:49:07 -08:00
Woosuk Kwon	b0a1d667b0	Pin PyTorch & xformers versions (#2155 )	2023-12-17 01:46:54 -08:00
Woosuk Kwon	e1d5402238	Fix all-reduce memory usage (#2151 )	2023-12-17 01:44:45 -08:00
Woosuk Kwon	3d1cfbfc74	[Minor] Delete Llama tokenizer warnings (#2146 )	2023-12-16 22:05:18 -08:00
Woosuk Kwon	37ca558103	Optimize model execution with CUDA graph (#1926 ) Co-authored-by: Chen Shen <scv119@gmail.com> Co-authored-by: Antoni Baum <antoni.baum@protonmail.com>	2023-12-16 21:12:08 -08:00
Roy	eed74a558f	Simplify weight loading logic (#2133 )	2023-12-16 12:41:23 -08:00
Woosuk Kwon	2acd76f346	[ROCm] Temporarily remove GPTQ ROCm support (#2138 )	2023-12-15 17:13:58 -08:00
Woosuk Kwon	b81a6a6bb3	[Docs] Add supported quantization methods to docs (#2135 )	2023-12-15 13:29:22 -08:00
CHU Tianxiang	0fbfc4b81b	Add GPTQ support (#916 )	2023-12-15 03:04:22 -08:00

1 2 3 4 5 ...

624 Commits