Ram
|
2eaa81b236
|
Update README.md to add megablocks requirement for mixtral (#2033)
|
2023-12-11 11:37:34 -08:00 |
|
Woosuk Kwon
|
81ce2a4b26
|
[Minor] Fix type annotation in Mixtral (#2036)
|
2023-12-11 11:32:39 -08:00 |
|
Woosuk Kwon
|
5dd80d3777
|
Fix latency benchmark script (#2035)
|
2023-12-11 11:19:08 -08:00 |
|
Woosuk Kwon
|
beeee69bc9
|
Revert adding Megablocks (#2030)
|
2023-12-11 10:49:00 -08:00 |
|
Ram
|
9bf28d0b69
|
Update requirements.txt for mixtral (#2029)
|
2023-12-11 10:39:29 -08:00 |
|
Ikko Eltociear Ashimine
|
c0ce15dfb2
|
Update run_on_sky.rst (#2025)
sharable -> shareable
|
2023-12-11 10:32:58 -08:00 |
|
Woosuk Kwon
|
b9bcdc7158
|
Change the load format to pt for Mixtral (#2028)
|
2023-12-11 10:32:17 -08:00 |
|
Woosuk Kwon
|
4ff0203987
|
Minor fixes for Mixtral (#2015)
|
2023-12-11 09:16:15 -08:00 |
|
Pierre Stock
|
b5f882cc98
|
Mixtral 8x7B support (#2011)
Co-authored-by: Pierre Stock <p@mistral.ai>
Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>
|
2023-12-11 01:09:15 -08:00 |
|
Simon Mo
|
2e8fc0d4c3
|
Fix completion API echo and logprob combo (#1992)
|
2023-12-10 13:20:30 -08:00 |
|
wbn
|
dacaf5a400
|
Replace head_mapping params with num_kv_heads to attention kernel. (#1997)
Co-authored-by: wangguoya <wangguoya@baidu.com>
Co-authored-by: Yang Zhao <zhaoyangstar@foxmail.com>
|
2023-12-10 10:12:53 -08:00 |
|
Woosuk Kwon
|
24cde76a15
|
[Minor] Add comment on skipping rope caches (#2004)
|
2023-12-10 10:04:12 -08:00 |
|
Jin Shang
|
1aa1361510
|
Fix OpenAI server completion_tokens referenced before assignment (#1996)
|
2023-12-09 21:01:21 -08:00 |
|
Woosuk Kwon
|
fe470ae5ad
|
[Minor] Fix code style for baichuan (#2003)
|
2023-12-09 19:24:29 -08:00 |
|
Jun Gao
|
3a8c2381f7
|
Fix for KeyError on Loading LLaMA (#1978)
|
2023-12-09 15:59:57 -08:00 |
|
Simon Mo
|
c85b80c2b6
|
[Docker] Add cuda arch list as build option (#1950)
|
2023-12-08 09:53:47 -08:00 |
|
firebook
|
2b981012a6
|
Fix Baichuan2-7B-Chat (#1987)
|
2023-12-08 09:38:36 -08:00 |
|
TJian
|
6ccc0bfffb
|
Merge EmbeddedLLM/vllm-rocm into vLLM main (#1836)
Co-authored-by: Philipp Moritz <pcmoritz@gmail.com>
Co-authored-by: Amir Balwel <amoooori04@gmail.com>
Co-authored-by: root <kuanfu.liu@akirakan.com>
Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com>
Co-authored-by: kuanfu <kuanfu.liu@embeddedllm.com>
Co-authored-by: miloice <17350011+kliuae@users.noreply.github.com>
|
2023-12-07 23:16:52 -08:00 |
|
Daya Khudia
|
c8e7eb1eb3
|
fix typo in getenv call (#1972)
|
2023-12-07 16:04:41 -08:00 |
|
AguirreNicolas
|
24f60a54f4
|
[Docker] Adding number of nvcc_threads during build as envar (#1893)
|
2023-12-07 11:00:32 -08:00 |
|
gottlike
|
42c02f5892
|
Fix quickstart.rst typo jinja (#1964)
|
2023-12-07 08:34:44 -08:00 |
|
Jie Li
|
ebede26ebf
|
Make InternLM follow rope_scaling in config.json (#1956)
Co-authored-by: lijie8 <lijie8@sensetime.com>
|
2023-12-07 08:32:08 -08:00 |
|
Peter Götz
|
d940ce497e
|
Fix typo in adding_model.rst (#1947)
adpated -> adapted
|
2023-12-06 10:04:26 -08:00 |
|
Antoni Baum
|
05ff90b692
|
Save pytorch profiler output for latency benchmark (#1871)
* Save profiler output
* Apply feedback from code review
|
2023-12-05 20:55:55 -08:00 |
|
dancingpipi
|
1d9b737e05
|
Support ChatGLMForConditionalGeneration (#1932)
Co-authored-by: shujunhua1 <shujunhua1@jd.com>
|
2023-12-05 10:52:48 -08:00 |
|
Roy
|
60dc62dc9e
|
add custom server params (#1868)
|
2023-12-03 12:59:18 -08:00 |
|
Woosuk Kwon
|
0f90effc66
|
Bump up to v0.2.3 (#1903)
v0.2.3
|
2023-12-03 12:27:47 -08:00 |
|
Woosuk Kwon
|
464dd985e3
|
Fix num_gpus when TP > 1 (#1852)
|
2023-12-03 12:24:30 -08:00 |
|
Massimiliano Pronesti
|
c07a442854
|
chore(examples-docs): upgrade to OpenAI V1 (#1785)
|
2023-12-03 01:11:22 -08:00 |
|
Woosuk Kwon
|
cd3aa153a4
|
Fix broken worker test (#1900)
|
2023-12-02 22:17:33 -08:00 |
|
Woosuk Kwon
|
9b294976a2
|
Add PyTorch-native implementation of custom layers (#1898)
|
2023-12-02 21:18:40 -08:00 |
|
Simon Mo
|
5313c2cb8b
|
Add Production Metrics in Prometheus format (#1890)
|
2023-12-02 16:37:44 -08:00 |
|
Woosuk Kwon
|
5f09cbdb63
|
Fix broken sampler tests (#1896)
Co-authored-by: Antoni Baum <antoni.baum@protonmail.com>
|
2023-12-02 16:06:17 -08:00 |
|
Simon Mo
|
4cefa9b49b
|
[Docs] Update the AWQ documentation to highlight performance issue (#1883)
|
2023-12-02 15:52:47 -08:00 |
|
Jerry
|
f86bd6190a
|
Fix the typo in SamplingParams' docstring (#1886)
|
2023-12-01 02:06:36 -08:00 |
|
Woosuk Kwon
|
e5452ddfd6
|
Normalize head weights for Baichuan 2 (#1876)
|
2023-11-30 20:03:58 -08:00 |
|
Woosuk Kwon
|
d06980dfa7
|
Fix Baichuan tokenizer error (#1874)
|
2023-11-30 18:35:50 -08:00 |
|
Adam Brusselback
|
66785cc05c
|
Support chat template and echo for chat API (#1756)
|
2023-11-30 16:43:13 -08:00 |
|
Massimiliano Pronesti
|
05a38612b0
|
docs: add instruction for langchain (#1162)
|
2023-11-30 10:57:44 -08:00 |
|
Roy
|
d27f4bae39
|
Fix rope cache key error (#1867)
|
2023-11-30 08:29:28 -08:00 |
|
aisensiy
|
8d8c2f6ffe
|
Support max-model-len argument for throughput benchmark (#1858)
|
2023-11-30 08:10:24 -08:00 |
|
Woosuk Kwon
|
51d3cb951d
|
Remove max_num_seqs in latency benchmark script (#1855)
|
2023-11-30 00:00:32 -08:00 |
|
Woosuk Kwon
|
e74b1736a1
|
Add profile option to latency benchmark script (#1839)
|
2023-11-29 23:42:52 -08:00 |
|
Allen
|
f07c1ceaa5
|
[FIX] Fix docker build error (#1831) (#1832)
Co-authored-by: Antoni Baum <antoni.baum@protonmail.com>
|
2023-11-29 23:06:50 -08:00 |
|
Jee Li
|
63b2206ad0
|
Avoid multiple instantiations of the RoPE class (#1828)
|
2023-11-29 23:06:27 -08:00 |
|
Woosuk Kwon
|
27feead2f8
|
Refactor Worker & InputMetadata (#1843)
|
2023-11-29 22:16:37 -08:00 |
|
Michael McCulloch
|
c782195662
|
Disable Logs Requests should Disable Logging of requests. (#1779)
Co-authored-by: Michael McCulloch <mjm.gitlab@fastmail.com>
|
2023-11-29 21:50:02 -08:00 |
|
Simon Mo
|
0f621c2c7d
|
[Docs] Add information about using shared memory in docker (#1845)
|
2023-11-29 18:33:56 -08:00 |
|
Woosuk Kwon
|
a9e4574261
|
Refactor Attention (#1840)
|
2023-11-29 15:37:31 -08:00 |
|
FlorianJoncour
|
0229c386c5
|
Better integration with Ray Serve (#1821)
Co-authored-by: FlorianJoncour <florian@zetta-sys.com>
|
2023-11-29 13:25:43 -08:00 |
|