xinyun/vllm - vllm - 丝路新云-代码仓

mirror of https://git.datalinker.icu/vllm-project/vllm.git synced 2026-05-22 10:17:52 +08:00

Author	SHA1	Message	Date
Ram	2eaa81b236	Update README.md to add megablocks requirement for mixtral (#2033 )	2023-12-11 11:37:34 -08:00
Woosuk Kwon	81ce2a4b26	[Minor] Fix type annotation in Mixtral (#2036 )	2023-12-11 11:32:39 -08:00
Woosuk Kwon	5dd80d3777	Fix latency benchmark script (#2035 )	2023-12-11 11:19:08 -08:00
Woosuk Kwon	beeee69bc9	Revert adding Megablocks (#2030 )	2023-12-11 10:49:00 -08:00
Ram	9bf28d0b69	Update requirements.txt for mixtral (#2029 )	2023-12-11 10:39:29 -08:00
Ikko Eltociear Ashimine	c0ce15dfb2	Update run_on_sky.rst (#2025 ) sharable -> shareable	2023-12-11 10:32:58 -08:00
Woosuk Kwon	b9bcdc7158	Change the load format to pt for Mixtral (#2028 )	2023-12-11 10:32:17 -08:00
Woosuk Kwon	4ff0203987	Minor fixes for Mixtral (#2015 )	2023-12-11 09:16:15 -08:00
Pierre Stock	b5f882cc98	Mixtral 8x7B support (#2011 ) Co-authored-by: Pierre Stock <p@mistral.ai> Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>	2023-12-11 01:09:15 -08:00
Simon Mo	2e8fc0d4c3	Fix completion API echo and logprob combo (#1992 )	2023-12-10 13:20:30 -08:00
wbn	dacaf5a400	Replace head_mapping params with num_kv_heads to attention kernel. (#1997 ) Co-authored-by: wangguoya <wangguoya@baidu.com> Co-authored-by: Yang Zhao <zhaoyangstar@foxmail.com>	2023-12-10 10:12:53 -08:00
Woosuk Kwon	24cde76a15	[Minor] Add comment on skipping rope caches (#2004 )	2023-12-10 10:04:12 -08:00
Jin Shang	1aa1361510	Fix OpenAI server completion_tokens referenced before assignment (#1996 )	2023-12-09 21:01:21 -08:00
Woosuk Kwon	fe470ae5ad	[Minor] Fix code style for baichuan (#2003 )	2023-12-09 19:24:29 -08:00
Jun Gao	3a8c2381f7	Fix for KeyError on Loading LLaMA (#1978 )	2023-12-09 15:59:57 -08:00
Simon Mo	c85b80c2b6	[Docker] Add cuda arch list as build option (#1950 )	2023-12-08 09:53:47 -08:00
firebook	2b981012a6	Fix Baichuan2-7B-Chat (#1987 )	2023-12-08 09:38:36 -08:00
TJian	6ccc0bfffb	Merge EmbeddedLLM/vllm-rocm into vLLM main (#1836 ) Co-authored-by: Philipp Moritz <pcmoritz@gmail.com> Co-authored-by: Amir Balwel <amoooori04@gmail.com> Co-authored-by: root <kuanfu.liu@akirakan.com> Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com> Co-authored-by: kuanfu <kuanfu.liu@embeddedllm.com> Co-authored-by: miloice <17350011+kliuae@users.noreply.github.com>	2023-12-07 23:16:52 -08:00
Daya Khudia	c8e7eb1eb3	fix typo in getenv call (#1972 )	2023-12-07 16:04:41 -08:00
AguirreNicolas	24f60a54f4	[Docker] Adding number of nvcc_threads during build as envar (#1893 )	2023-12-07 11:00:32 -08:00
gottlike	42c02f5892	Fix quickstart.rst typo jinja (#1964 )	2023-12-07 08:34:44 -08:00
Jie Li	ebede26ebf	Make InternLM follow `rope_scaling` in `config.json` (#1956 ) Co-authored-by: lijie8 <lijie8@sensetime.com>	2023-12-07 08:32:08 -08:00
Peter Götz	d940ce497e	Fix typo in adding_model.rst (#1947 ) adpated -> adapted	2023-12-06 10:04:26 -08:00
Antoni Baum	05ff90b692	Save pytorch profiler output for latency benchmark (#1871 ) * Save profiler output * Apply feedback from code review	2023-12-05 20:55:55 -08:00
dancingpipi	1d9b737e05	Support ChatGLMForConditionalGeneration (#1932 ) Co-authored-by: shujunhua1 <shujunhua1@jd.com>	2023-12-05 10:52:48 -08:00
Roy	60dc62dc9e	add custom server params (#1868 )	2023-12-03 12:59:18 -08:00
Woosuk Kwon	0f90effc66	Bump up to v0.2.3 (#1903 ) v0.2.3	2023-12-03 12:27:47 -08:00
Woosuk Kwon	464dd985e3	Fix num_gpus when TP > 1 (#1852 )	2023-12-03 12:24:30 -08:00
Massimiliano Pronesti	c07a442854	chore(examples-docs): upgrade to OpenAI V1 (#1785 )	2023-12-03 01:11:22 -08:00
Woosuk Kwon	cd3aa153a4	Fix broken worker test (#1900 )	2023-12-02 22:17:33 -08:00
Woosuk Kwon	9b294976a2	Add PyTorch-native implementation of custom layers (#1898 )	2023-12-02 21:18:40 -08:00
Simon Mo	5313c2cb8b	Add Production Metrics in Prometheus format (#1890 )	2023-12-02 16:37:44 -08:00
Woosuk Kwon	5f09cbdb63	Fix broken sampler tests (#1896 ) Co-authored-by: Antoni Baum <antoni.baum@protonmail.com>	2023-12-02 16:06:17 -08:00
Simon Mo	4cefa9b49b	[Docs] Update the AWQ documentation to highlight performance issue (#1883 )	2023-12-02 15:52:47 -08:00
Jerry	f86bd6190a	Fix the typo in SamplingParams' docstring (#1886 )	2023-12-01 02:06:36 -08:00
Woosuk Kwon	e5452ddfd6	Normalize head weights for Baichuan 2 (#1876 )	2023-11-30 20:03:58 -08:00
Woosuk Kwon	d06980dfa7	Fix Baichuan tokenizer error (#1874 )	2023-11-30 18:35:50 -08:00
Adam Brusselback	66785cc05c	Support chat template and `echo` for chat API (#1756 )	2023-11-30 16:43:13 -08:00
Massimiliano Pronesti	05a38612b0	docs: add instruction for langchain (#1162 )	2023-11-30 10:57:44 -08:00
Roy	d27f4bae39	Fix rope cache key error (#1867 )	2023-11-30 08:29:28 -08:00
aisensiy	8d8c2f6ffe	Support max-model-len argument for throughput benchmark (#1858 )	2023-11-30 08:10:24 -08:00
Woosuk Kwon	51d3cb951d	Remove max_num_seqs in latency benchmark script (#1855 )	2023-11-30 00:00:32 -08:00
Woosuk Kwon	e74b1736a1	Add profile option to latency benchmark script (#1839 )	2023-11-29 23:42:52 -08:00
Allen	f07c1ceaa5	[FIX] Fix docker build error (#1831 ) (#1832 ) Co-authored-by: Antoni Baum <antoni.baum@protonmail.com>	2023-11-29 23:06:50 -08:00
Jee Li	63b2206ad0	Avoid multiple instantiations of the RoPE class (#1828 )	2023-11-29 23:06:27 -08:00
Woosuk Kwon	27feead2f8	Refactor Worker & InputMetadata (#1843 )	2023-11-29 22:16:37 -08:00
Michael McCulloch	c782195662	Disable Logs Requests should Disable Logging of requests. (#1779 ) Co-authored-by: Michael McCulloch <mjm.gitlab@fastmail.com>	2023-11-29 21:50:02 -08:00
Simon Mo	0f621c2c7d	[Docs] Add information about using shared memory in docker (#1845 )	2023-11-29 18:33:56 -08:00
Woosuk Kwon	a9e4574261	Refactor Attention (#1840 )	2023-11-29 15:37:31 -08:00
FlorianJoncour	0229c386c5	Better integration with Ray Serve (#1821 ) Co-authored-by: FlorianJoncour <florian@zetta-sys.com>	2023-11-29 13:25:43 -08:00

1 2 3 4 5 ...

553 Commits