Philipp Moritz
|
317b29de0f
|
Remove Yi model definition, please use LlamaForCausalLM instead (#2854)
Co-authored-by: Roy <jasonailu87@gmail.com>
|
2024-02-13 14:22:22 -08:00 |
|
Woosuk Kwon
|
a463c333dd
|
Use CuPy for CUDA graphs (#2811)
|
2024-02-13 11:32:06 -08:00 |
|
Philipp Moritz
|
ea356004d4
|
Revert "Refactor llama family models (#2637)" (#2851)
This reverts commit 5c976a7e1a1bec875bf6474824b7dff39e38de18.
|
2024-02-13 09:24:59 -08:00 |
|
Roy
|
5c976a7e1a
|
Refactor llama family models (#2637)
|
2024-02-13 00:09:23 -08:00 |
|
Simon Mo
|
f964493274
|
[CI] Ensure documentation build is checked in CI (#2842)
|
2024-02-12 22:53:07 -08:00 |
|
Roger Wang
|
a4211a4dc3
|
Serving Benchmark Refactoring (#2433)
|
2024-02-12 22:53:00 -08:00 |
|
Rex
|
563836496a
|
Refactor 2 awq gemm kernels into m16nXk32 (#2723)
Co-authored-by: Chunan Zeng <chunanzeng@Chunans-Air.attlocal.net>
|
2024-02-12 11:02:17 -08:00 |
|
Philipp Moritz
|
4ca2c358b1
|
Add documentation section about LoRA (#2834)
|
2024-02-12 17:24:45 +01:00 |
|
Hongxia Yang
|
0580aab02f
|
[ROCm] support Radeon™ 7900 series (gfx1100) without using flash-attention (#2768)
|
2024-02-10 23:14:37 -08:00 |
|
Woosuk Kwon
|
3711811b1d
|
Disable custom all reduce by default (#2808)
|
2024-02-08 09:58:03 -08:00 |
|
SangBin Cho
|
65b89d16ee
|
[Ray] Integration compiled DAG off by default (#2471)
|
2024-02-08 09:57:25 -08:00 |
|
Philipp Moritz
|
931746bc6d
|
Add documentation on how to do incremental builds (#2796)
|
2024-02-07 14:42:02 -08:00 |
|
Hongxia Yang
|
c81dddb45c
|
[ROCm] Fix build problem resulted from previous commit related to FP8 kv-cache support (#2790)
|
2024-02-06 22:36:59 -08:00 |
|
Lily Liu
|
fe6d09ae61
|
[Minor] More fix of test_cache.py CI test failure (#2750)
|
2024-02-06 11:38:38 -08:00 |
|
liuyhwangyh
|
ed70c70ea3
|
modelscope: fix issue when model parameter is not a model id but path of the model. (#2489)
|
2024-02-06 09:57:15 -08:00 |
|
Woosuk Kwon
|
f0d4e14557
|
Add fused top-K softmax kernel for MoE (#2769)
|
2024-02-05 17:38:02 -08:00 |
|
Douglas Lehr
|
2ccee3def6
|
[ROCm] Fixup arch checks for ROCM (#2627)
|
2024-02-05 14:59:09 -08:00 |
|
Lukas
|
b92adec8e8
|
Set local logging level via env variable (#2774)
|
2024-02-05 14:26:50 -08:00 |
|
Hongxia Yang
|
56f738ae9b
|
[ROCm] Fix some kernels failed unit tests (#2498)
|
2024-02-05 14:25:36 -08:00 |
|
Woosuk Kwon
|
72d3a30c63
|
[Minor] Fix benchmark_latency script (#2765)
|
2024-02-05 12:45:37 -08:00 |
|
whyiug
|
c9b45adeeb
|
Require triton >= 2.1.0 (#2746)
Co-authored-by: yangrui1 <yangrui@lanjingren.com>
|
2024-02-04 23:07:36 -08:00 |
|
Rex
|
5a6c81b051
|
Remove eos tokens from output by default (#2611)
|
2024-02-04 14:32:42 -08:00 |
|
dancingpipi
|
51cd22ce56
|
set&get llm internal tokenizer instead of the TokenizerGroup (#2741)
Co-authored-by: shujunhua1 <shujunhua1@jd.com>
|
2024-02-04 14:25:36 -08:00 |
|
Massimiliano Pronesti
|
5ed704ec8c
|
docs: fix langchain (#2736)
|
2024-02-03 18:17:55 -08:00 |
|
Cheng Su
|
4abf6336ec
|
Add one example to run batch inference distributed on Ray (#2696)
|
2024-02-02 15:41:42 -08:00 |
|
zspo
|
0e163fce18
|
Fix default length_penalty to 1.0 (#2667)
|
2024-02-01 15:59:39 -08:00 |
|
Kunshang Ji
|
96b6f475dd
|
Remove hardcoded device="cuda" to support more devices (#2503)
Co-authored-by: Jiang Li <jiang1.li@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2024-02-01 15:46:39 -08:00 |
|
Pernekhan Utemuratov
|
c410f5d020
|
Use revision when downloading the quantization config file (#2697)
Co-authored-by: Pernekhan Utemuratov <pernekhan@deepinfra.com>
|
2024-02-01 15:41:58 -08:00 |
|
Simon Mo
|
bb8c697ee0
|
Update README for meetup slides (#2718)
|
2024-02-01 14:56:53 -08:00 |
|
Simon Mo
|
b9e96b17de
|
fix python 3.8 syntax (#2716)
|
2024-02-01 14:00:58 -08:00 |
|
zhaoyang-star
|
923797fea4
|
Fix compile error when using rocm (#2648)
|
2024-02-01 09:35:09 -08:00 |
|
Fengzhe Zhou
|
cd9e60c76c
|
Add Internlm2 (#2666)
|
2024-02-01 09:27:40 -08:00 |
|
Robert Shaw
|
93b38bea5d
|
Refactor Prometheus and Add Request Level Metrics (#2316)
|
2024-01-31 14:58:07 -08:00 |
|
Philipp Moritz
|
d0d93b92b1
|
Add unit test for Mixtral MoE layer (#2677)
|
2024-01-31 14:34:17 -08:00 |
|
Philipp Moritz
|
89efcf1ce5
|
[Minor] Fix test_cache.py CI test failure (#2684)
|
2024-01-31 10:12:11 -08:00 |
|
zspo
|
c664b0e683
|
fix some bugs (#2689)
|
2024-01-31 10:09:23 -08:00 |
|
Tao He
|
d69ff0cbbb
|
Fixes assertion failure in prefix caching: the lora index mapping should respect prefix_len (#2688)
Signed-off-by: Tao He <sighingnow@gmail.com>
|
2024-01-31 18:00:13 +01:00 |
|
Zhuohan Li
|
1af090b57d
|
Bump up version to v0.3.0 (#2656)
v0.3.0
|
2024-01-31 00:07:07 -08:00 |
|
Woosuk Kwon
|
3dad944485
|
Add quantized mixtral support (#2673)
|
2024-01-30 16:34:10 -08:00 |
|
Woosuk Kwon
|
105a40f53a
|
[Minor] Fix false warning when TP=1 (#2674)
|
2024-01-30 14:39:40 -08:00 |
|
Philipp Moritz
|
bbe9bd9684
|
[Minor] Fix a small typo (#2672)
|
2024-01-30 13:40:37 -08:00 |
|
Vladimir
|
4f65af0e25
|
Add swap_blocks unit tests (#2616)
|
2024-01-30 09:30:50 -08:00 |
|
Wen Sun
|
d79ced3292
|
Fix 'Actor methods cannot be called directly' when using --engine-use-ray (#2664)
* fix: engine-useray complain
* fix: typo
|
2024-01-30 17:17:05 +01:00 |
|
Philipp Moritz
|
ab40644669
|
Fused MOE for Mixtral (#2542)
Co-authored-by: chen shen <scv119@gmail.com>
|
2024-01-29 22:43:37 -08:00 |
|
wangding zeng
|
5d60def02c
|
DeepseekMoE support with Fused MoE kernel (#2453)
Co-authored-by: roy <jasonailu87@gmail.com>
|
2024-01-29 21:19:48 -08:00 |
|
Rasmus Larsen
|
ea8489fce2
|
ROCm: Allow setting compilation target (#2581)
|
2024-01-29 10:52:31 -08:00 |
|
Hanzhi Zhou
|
1b20639a43
|
No repeated IPC open (#2642)
|
2024-01-29 10:46:29 -08:00 |
|
zhaoyang-star
|
b72af8f1ed
|
Fix error when tp > 1 (#2644)
Co-authored-by: zhaoyang-star <zhao.yang16@zte.com.cn>
|
2024-01-28 22:47:39 -08:00 |
|
zhaoyang-star
|
9090bf02e7
|
Support FP8-E5M2 KV Cache (#2279)
Co-authored-by: zhaoyang <zhao.yang16@zte.com.cn>
Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>
|
2024-01-28 16:43:54 -08:00 |
|
Simon Mo
|
7d648418b8
|
Update Ray version requirements (#2636)
|
2024-01-28 14:27:22 -08:00 |
|