xinyun/vllm - vllm - 丝路新云-代码仓

mirror of https://git.datalinker.icu/vllm-project/vllm.git synced 2026-05-11 00:38:00 +08:00

Author	SHA1	Message	Date
Antoni Baum	ccdc490dda	[Core] Change LoRA embedding sharding to support loading methods (#5038 )	2024-06-06 19:07:57 -07:00
Antoni Baum	a31cab7556	[Core] Avoid copying prompt/output tokens if no penalties are used (#5289 )	2024-06-06 18:12:00 -07:00
Philipp Moritz	abe855d637	[Kernel] Retune Mixtral 8x22b configs for FP8 on H100 (#5294 )	2024-06-06 09:29:29 -07:00
Breno Faria	7b0a0dfb22	[Frontend][Core] Update Outlines Integration from `FSM` to `Guide` (#4109 ) Co-authored-by: Simon Mo <simon.mo@hey.com> Co-authored-by: Breno Faria <breno.faria@intrafind.com>	2024-06-05 16:49:12 -07:00
Woosuk Kwon	6a7c7711a2	[Misc] Skip for logits_scale == 1.0 (#5291 )	2024-06-05 15:19:02 -07:00
Philipp Moritz	51a08e7d8f	[Kernel] Re-tune Mixtral MoE configurations for FP8 on H100 (#5238 )	2024-06-05 10:59:14 -07:00
Cody Yu	5563a4dea8	[Model] Correct Mixtral FP8 checkpoint loading (#5231 )	2024-06-05 10:58:50 -07:00
Woosuk Kwon	41ca62cf03	[Misc] Add CustomOp interface for device portability (#5255 )	2024-06-05 09:18:19 -07:00
Woosuk Kwon	3a434b07ed	[Kernel] Enhance MoE benchmarking & tuning script (#4921 )	2024-06-03 20:06:59 -07:00
Toshiki Kataoka	06b2550cbb	[Bugfix] Support `prompt_logprobs==0` (#5217 )	2024-06-03 17:59:30 -07:00
Breno Faria	f775a07e30	[FRONTEND] OpenAI `tools` support named functions (#5032 )	2024-06-03 18:25:29 -05:00
Tyler Michael Smith	cbb2f59cc8	[Kernel] Pass a device pointer into the quantize kernel for the scales (#5159 )	2024-06-03 09:52:30 -07:00
Cyrus Leung	7a64d24aad	[Core] Support image processor (#4197 )	2024-06-02 22:56:41 -07:00
Divakar Verma	a66cf40b20	[Kernel][ROCm][AMD] enable fused topk_softmax kernel for moe layer (#4927 ) This PR enables the fused topk_softmax kernel used in moe layer for HIP	2024-06-02 14:13:26 -07:00
chenqianfzh	b9c0605a8e	[Feature][Kernel] Support bitsandbytes quantization and QLoRA (#4776 )	2024-06-01 14:51:10 -06:00
Ye Cao	c354072828	[Minor] Fix the path typo in loader.py: save_sharded_states.py -> save_sharded_state.py (#5151 ) Signed-off-by: Ye Cao <caoye.cao@alibaba-inc.com>	2024-06-01 17:11:22 +00:00
Tyler Michael Smith	260d119e86	[Kernel] Refactor CUTLASS kernels to always take scales that reside on the GPU (#5137 )	2024-06-01 06:45:32 +00:00
Cody Yu	e9899fb7a4	[Model] Enable FP8 QKV in MoE and refine kernel tuning script (#5039 )	2024-05-31 14:29:19 -07:00
Robert Shaw	b35be5403f	[Bugfix] Avoid Warnings in SparseML Activation Quantization (#5120 )	2024-05-30 17:04:37 -07:00
Alexander Matveev	5bf185a1c4	[Bugfix] gptq_marlin: Ensure g_idx_sort_indices is not a Parameter (#5108 )	2024-05-30 00:30:18 +00:00
Divakar Verma	dd8de11f0a	[Kernel][ROCm][AMD] Add fused_moe Triton configs for MI300X (#4951 ) This PR adds Triton kernel configs for the MoE kernel for MI300X	2024-05-28 16:03:23 +00:00
Isotr0py	890aa93d27	[Model] Add support for falcon-11B (#5069 )	2024-05-27 16:41:43 -07:00
sasha0552	fbdb7b3ee2	[Core] Allow AQLM on Pascal (#5058 )	2024-05-27 15:26:14 -07:00
Zhuohan Li	1102bef219	[Bugfix / Core] Prefix Caching Guards (merged with main) (#4846 ) Co-authored-by: rsnm2 <rshaw@neuralmagic.com> Co-authored-by: Robert Shaw <114415538+robertgshaw2-neuralmagic@users.noreply.github.com>	2024-05-27 15:18:17 -07:00
Eric Xihui Lin	8e192ff967	[Kernel][Backend][Model] Blocksparse flash attention kernel and Phi-3-Small model (#4799 ) Co-authored-by: beagleski <yunanzhang@microsoft.com> Co-authored-by: bapatra <bapatra@microsoft.com> Co-authored-by: Barun Patra <codedecde@users.noreply.github.com> Co-authored-by: Michael Goin <michael@neuralmagic.com>	2024-05-24 22:00:52 -07:00
Robert Shaw	919770957f	[Bugfix] Fix Mistral v0.3 Weight Loading (#5005 ) Co-authored-by: Cody Yu <hao.yu.cody@gmail.com>	2024-05-24 12:28:27 +00:00
Elisei Smirnov	e3470f8753	[Core]: Option To Use Prompt Token Ids Inside Logits Processor (#4985 ) Co-authored-by: Elisei Smirnov <el.smirnov@innopolis.university>	2024-05-23 22:04:24 +00:00
Dipika Sikka	a1242324c9	[Kernel] Initial Activation Quantization Support (#4525 ) Co-authored-by: Varun Sundar Rabindranath <varunsundar08@gmail.com> Co-authored-by: Varun Sundar Rabindranath <varun@neuralmagic.com>	2024-05-23 21:29:18 +00:00
Alexander Matveev	6066253296	Marlin 24 prefill performance improvement (about 25% better on average) (#4983 )	2024-05-23 02:39:27 -04:00
Philipp Moritz	a36de682d4	[Minor] Fix small typo in llama.py: QKVParallelLinear -> QuantizationConfig (#4991 )	2024-05-22 22:26:56 +00:00
raywanb	97b030005c	[Model] LoRA gptbigcode implementation (#3949 )	2024-05-22 13:58:59 -07:00
Cody Yu	a3a73ab069	[Misc] Load FP8 kv-cache scaling factors from checkpoints (#4893 ) The 2nd PR for #4532. This PR supports loading FP8 kv-cache scaling factors from a FP8 checkpoint (with .kv_scale parameter).	2024-05-22 13:28:20 -07:00
Isotr0py	f12c3b5b3d	[Model] Add Phi-2 LoRA support (#4886 )	2024-05-21 14:24:17 +09:00
HUANG Fei	d130b573a0	[Model] add rope_scaling support for qwen2 (#4930 )	2024-05-21 05:22:22 +00:00
Aurick Qiao	1937e29848	[Core] Sharded State Loader download from HF (#4889 )	2024-05-20 11:46:12 -07:00
Mor Zusman	f0eecee610	[Bugfix] Fix dummy weight for fp8 (#4916 ) Allow dummy load format for fp8, torch.uniform_ doesn't support FP8 at the moment Co-authored-by: Mor Zusman <morz@ai21.com>	2024-05-20 18:44:25 +00:00
Cyrus Leung	6287537a0c	[Model] LLaVA model refactor (#4910 )	2024-05-20 08:11:25 +00:00
Alexander Matveev	27ce85476e	[Kernel] Add marlin_24 unit tests (#4901 )	2024-05-19 11:37:34 -04:00
Cyrus Leung	f68470e803	[Bugfix][Model] Add base class for vision-language models (#4809 )	2024-05-19 00:13:33 -07:00
SangBin Cho	2e9a2227ec	[Lora] Support long context lora (#4787 ) Currently we need to call rotary embedding kernel for each LoRA, which makes it hard to serve multiple long context length LoRA. Add batched rotary embedding kernel and pipe it through. It replaces the rotary embedding layer to the one that is aware of multiple cos-sin-cache per scaling factors. Follow up of https://github.com/vllm-project/vllm/pull/3095/files	2024-05-18 16:05:23 +09:00
eigenLiu	48d5985a08	Sync huggingface modifications of qwen Moe model (#4774 )	2024-05-17 09:43:19 -07:00
Jinzhen Lin	33e0823de5	[Bugfix] fix rope error when load models with different dtypes (#4835 )	2024-05-17 18:43:34 +09:00
Alexander Matveev	6979ade384	Add GPTQ Marlin 2:4 sparse structured support (#4790 ) Co-authored-by: Robert Shaw <rshaw@neuralmagic.com>	2024-05-16 12:56:15 -04:00
Jinzhen Lin	99caa49106	[Kernel] add bfloat16 support for gptq marlin kernel (#4788 )	2024-05-16 09:55:29 -04:00
alexm-nm	5c342570d7	Add marlin unit tests and marlin benchmark script (#4815 )	2024-05-16 09:36:49 -04:00
Aurick Qiao	30e754390c	[Core] Implement sharded state loader (#4690 ) Co-authored-by: Woosuk Kwon <woosuk.kwon@berkeley.edu>	2024-05-15 22:11:54 -07:00
SangBin Cho	65bf2ac165	[Core][2/N] Model runner refactoring part 2. Combine prepare prefill / decode to a single API (#4681 ) This PR combines prepare_prompt and prepare_decode into a single API. This PR also coelsce the attn metadata for prefill/decode to a single class and allow to slice them when running attn backend. It also refactors subquery_start_loc which was not refactored in the previous PR	2024-05-15 14:00:10 +09:00
Philipp Moritz	33d3914b1e	[Bugfix] Fix dynamic FP8 quantization for Mixtral (#4793 )	2024-05-13 19:00:27 -04:00
Sanger Steel	8bc68e198c	[Frontend] [Core] perf: Automatically detect vLLM-tensorized model, update `tensorizer` to version 2.9.0 (#4208 )	2024-05-13 14:57:07 -07:00
Woosuk Kwon	0fca3cdcf2	[Misc] Enhance attention selector (#4751 )	2024-05-13 10:47:25 -07:00

... 3 4 5 6 7 ...

606 Commits