xinyun/vllm - vllm - 丝路新云-代码仓

mirror of https://git.datalinker.icu/vllm-project/vllm.git synced 2025-12-18 10:06:11 +08:00

Author	SHA1	Message	Date
Woosuk Kwon	37ca558103	Optimize model execution with CUDA graph (#1926 ) Co-authored-by: Chen Shen <scv119@gmail.com> Co-authored-by: Antoni Baum <antoni.baum@protonmail.com>	2023-12-16 21:12:08 -08:00
Roy	eed74a558f	Simplify weight loading logic (#2133 )	2023-12-16 12:41:23 -08:00
CHU Tianxiang	0fbfc4b81b	Add GPTQ support (#916 )	2023-12-15 03:04:22 -08:00
Antoni Baum	21d93c140d	Optimize Mixtral with expert parallelism (#2090 )	2023-12-13 23:55:07 -08:00
Woosuk Kwon	518369d78c	Implement lazy model loader (#2044 )	2023-12-12 22:21:45 -08:00
Woosuk Kwon	cb3f30c600	Upgrade transformers version to 4.36.0 (#2046 )	2023-12-11 18:39:14 -08:00
Woosuk Kwon	31d2ab4aff	Remove python 3.10 requirement (#2040 )	2023-12-11 12:26:42 -08:00
Woosuk Kwon	6120e5aaea	Fix import error msg for megablocks (#2038 )	2023-12-11 11:40:56 -08:00
Woosuk Kwon	81ce2a4b26	[Minor] Fix type annotation in Mixtral (#2036 )	2023-12-11 11:32:39 -08:00
Woosuk Kwon	4ff0203987	Minor fixes for Mixtral (#2015 )	2023-12-11 09:16:15 -08:00
Pierre Stock	b5f882cc98	Mixtral 8x7B support (#2011 ) Co-authored-by: Pierre Stock <p@mistral.ai> Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>	2023-12-11 01:09:15 -08:00