xinyun/vllm - vllm - 丝路新云-代码仓

mirror of https://git.datalinker.icu/vllm-project/vllm.git synced 2026-04-21 07:07:02 +08:00

Author	SHA1	Message	Date
Russell Bryant	d3d6bb13fb	Set weights_only=True when using torch.load() (#12366 ) Signed-off-by: Russell Bryant <rbryant@redhat.com>	2025-01-24 02:17:30 +00:00
Jee Jee Li	07934cc237	[Misc][LoRA] Improve the readability of LoRA error messages (#12102 ) Signed-off-by: Jee Jee Li <pandaleefree@gmail.com>	2025-01-17 19:32:28 +08:00
youkaichao	bf53e0c70b	Support torchrun and SPMD-style offline inference (#12071 ) Signed-off-by: youkaichao <youkaichao@gmail.com>	2025-01-16 19:58:53 +08:00
Varun Sundar Rabindranath	ebd8c669ef	[Bugfix] Fix _get_lora_device for HQQ marlin (#12090 ) Signed-off-by: Varun Sundar Rabindranath <varun@neuralmagic.com> Co-authored-by: Varun Sundar Rabindranath <varun@neuralmagic.com>	2025-01-15 19:59:42 +00:00
Shanshan Shen	a7d59688fb	[Platform] Move get_punica_wrapper() function to Platform (#11516 ) Signed-off-by: Shanshan Shen <467638484@qq.com>	2025-01-13 13:12:10 +00:00
Akshat Tripathi	8bddb73512	[Hardware][CPU] Multi-LoRA implementation for the CPU backend (#11100 ) Signed-off-by: Akshat Tripathi <akshat@krai.ai> Signed-off-by: Oleg Mosalov <oleg@krai.ai> Signed-off-by: Jee Jee Li <pandaleefree@gmail.com> Co-authored-by: Oleg Mosalov <oleg@krai.ai> Co-authored-by: Jee Jee Li <pandaleefree@gmail.com> Co-authored-by: Isotr0py <2037008807@qq.com>	2025-01-12 13:01:52 +00:00
Joe Runde	ac2f3f7fee	[Bugfix] Validate lora adapters to avoid crashing server (#11727 ) Signed-off-by: Joe Runde <Joseph.Runde@ibm.com> Co-authored-by: Jee Jee Li <pandaleefree@gmail.com>	2025-01-10 15:56:36 +08:00
Cyrus Leung	d848800e88	[Misc] Move `print_*_once` from utils to logger (#11298 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk> Signed-off-by: Maxime Fournioux <55544262+mfournioux@users.noreply.github.com> Co-authored-by: Maxime Fournioux <55544262+mfournioux@users.noreply.github.com>	2025-01-09 12:48:12 +08:00
Jee Jee Li	b278557935	[Kernel][LoRA]Punica prefill kernels fusion (#11234 ) Signed-off-by: Jee Jee Li <pandaleefree@gmail.com> Signed-off-by: Abatom <abzhonghua@gmail.com> Co-authored-by: Zhonghua Deng <abatom@163.com>	2025-01-07 04:01:39 +00:00
Lucas Tucker	9c749713f6	[mypy] Forward pass function type hints in lora (#11740 ) Signed-off-by: lucast2021 <lucast2021@headroyce.org> Co-authored-by: lucast2021 <lucast2021@headroyce.org>	2025-01-06 07:59:36 +00:00
ZincCat	61fed92c7e	[Bugfix] Fix ColumnParallelLinearWithLoRA slice (#11708 ) Signed-off-by: ZincCat <zincchloride@outlook.com>	2025-01-03 21:02:34 +00:00
John Giorgi	82c49d3260	[Misc][LoRA] Support Rank Stabilized LoRA (RSLoRA) (#6909 ) Signed-off-by: Jee Jee Li <pandaleefree@gmail.com> Co-authored-by: Jee Jee Li <pandaleefree@gmail.com>	2024-12-30 22:15:58 -08:00
Jee Jee Li	aa25985bd1	[Misc][LoRA] Fix LoRA weight mapper (#11495 ) Signed-off-by: Jee Jee Li <pandaleefree@gmail.com>	2024-12-26 15:52:48 +08:00
Jee Jee Li	b1b1038fbd	[Bugfix] Fix Qwen2-VL LoRA weight loading (#11430 ) Signed-off-by: Jee Jee Li <pandaleefree@gmail.com>	2024-12-24 09:56:10 +00:00
Jason T. Greene	f1d1bf6288	[Bugfix] Fix fully sharded LoRAs with Mixtral (#11390 ) Signed-off-by: Jason Greene <jason.greene@redhat.com>	2024-12-22 23:25:10 +08:00
Jee Jee Li	3cb5769883	[Misc] Minor improvements to the readability of PunicaWrapperBase (#11200 ) Signed-off-by: Jee Jee Li <pandaleefree@gmail.com>	2024-12-14 16:38:27 +00:00
Sanju C Sudhakaran	8195824206	[Hardware][Intel-Gaudi] Enable LoRA support for Intel Gaudi (HPU) (#10565 ) Signed-off-by: Sanju C Sudhakaran <scsudhakaran@habana.ai>	2024-12-12 08:09:28 +00:00
Jee Jee Li	d05f88679b	[Misc][LoRA] Add PEFTHelper for LoRA (#11003 ) Signed-off-by: Jee Jee Li <pandaleefree@gmail.com>	2024-12-10 11:12:01 +00:00
Jee Jee Li	ca871491ed	[Misc][LoRA] Abstract PunicaWrapper (#10955 ) Signed-off-by: Jee Jee Li <pandaleefree@gmail.com>	2024-12-09 12:54:44 -08:00
Isotr0py	b26b4cd03c	[Misc][LoRA] Refactor and clean MergedQKVParallelLinearWithLora implementation (#10958 ) Signed-off-by: Isotr0py <2037008807@qq.com>	2024-12-07 18:33:49 +08:00
Jee Jee Li	571da8fc43	[Misc][LoRA] Clean up the function interface of Punica (#10917 ) Signed-off-by: Jee Jee Li <pandaleefree@gmail.com>	2024-12-05 13:22:28 +00:00
Jee Jee Li	a4cf256159	[Bugfix] Fix QKVParallelLinearWithShardedLora bias bug (#10844 ) Signed-off-by: Jee Jee Li <pandaleefree@gmail.com>	2024-12-03 12:10:29 +08:00
Jee Jee Li	b45f0d7946	[Misc][LoRA] Move the implementation of lora bias to punica.py (#10829 ) Signed-off-by: Jee Jee Li <pandaleefree@gmail.com>	2024-12-02 17:53:36 +00:00
Jee Jee Li	1700c543a5	[Bugfix] Fix LoRA weight sharding (#10450 ) Signed-off-by: Jee Jee Li <pandaleefree@gmail.com> Co-authored-by: Cyrus Leung <cyrus.tl.leung@gmail.com>	2024-11-23 17:23:17 -08:00
Jee Jee Li	2385b60d83	[Kernel] Register punica ops directly (#10522 ) Signed-off-by: Jee Jee Li <pandaleefree@gmail.com>	2024-11-21 09:18:11 -08:00
Angus Wang	c2170a5b39	[Kernel] Explicitly specify other value in tl.load calls (#9014 ) Signed-off-by: Angus Wang <wangjadehao@gmail.com>	2024-11-18 11:39:40 -08:00
Jee Jee Li	1d65ec7eeb	[Bugfix] Fix fully sharded LoRA bug (#10352 ) Signed-off-by: Jee Jee Li <pandaleefree@gmail.com>	2024-11-15 10:34:58 +00:00
Umesh	8a06428c70	[LoRA] Adds support for bias in LoRA (#5733 ) Signed-off-by: Umesh Deshpande <udeshpa@us.ibm.com> Co-authored-by: Umesh Deshpande <udeshpa@us.ibm.com>	2024-11-12 11:08:40 -08:00
Jee Jee Li	7f5edb5900	[Misc][LoRA] Replace hardcoded cuda device with configurable argument (#10223 ) Signed-off-by: Jee Jee Li <pandaleefree@gmail.com>	2024-11-12 11:10:15 +08:00
Jee Jee Li	36e4acd02a	[LoRA][Kernel] Remove the unused libentry module (#10214 ) Signed-off-by: Jee Jee Li <pandaleefree@gmail.com>	2024-11-11 09:43:23 +00:00
Aaron Pham	21063c11c7	[CI/Build] drop support for Python 3.8 EOL (#8464 ) Signed-off-by: Aaron Pham <contact@aarnphm.xyz>	2024-11-06 07:11:55 +00:00
Jee Jee Li	7a4df5f200	[Model][LoRA]LoRA support added for Qwen (#9622 ) Signed-off-by: Jee Jee Li <pandaleefree@gmail.com>	2024-10-29 04:14:07 +00:00
Jee Jee Li	250e26a63e	[Bugfix]Fix MiniCPM's LoRA bug (#9286 )	2024-10-12 09:36:47 -07:00
Jee Jee Li	36ea79079b	[Misc][LoRA] Support loading LoRA weights for target_modules in reg format (#9275 )	2024-10-11 12:31:21 +00:00
Ahmad Fahadh Ilyas	21906a6f50	[Bugfix] Fix lora loading for Compressed Tensors in #9120 (#9179 )	2024-10-09 12:10:44 +00:00
Cyrus Leung	0e36fd4909	[Misc] Move registry to its own file (#9064 )	2024-10-04 10:01:37 +00:00
Jee Jee Li	3d49776bbb	[Model][LoRA]LoRA support added for MiniCPMV2.5 (#7199 )	2024-09-29 06:59:45 +00:00
Jee Jee Li	9b0e3ec970	[Kernel][LoRA] Add assertion for punica sgmv kernels (#7585 )	2024-09-23 18:57:42 +00:00
Jiaxin Shan	260d40b5ea	[Core] Support Lora lineage and base model metadata management (#6315 )	2024-09-20 06:20:56 +00:00
Dipika Sikka	23f322297f	[Misc] Remove `SqueezeLLM` (#8220 )	2024-09-06 16:29:03 -06:00
Jiaxin Shan	db3bf7c991	[Core] Support load and unload LoRA in api server (#6566 ) Co-authored-by: Jee Jee Li <pandaleefree@gmail.com>	2024-09-05 18:10:33 -07:00
bnellnm	3cdfe1f38b	[Bugfix] Make torch registration of punica ops optional (#7970 )	2024-08-28 16:11:49 -06:00
Kunshang Ji	6e4658c7aa	[Intel GPU] fix xpu not support punica kernel (which use torch.library.custom_op) (#7685 )	2024-08-20 12:01:09 -07:00
SangBin Cho	ff7ec82c4d	[Core] Optimize SPMD architecture with delta + serialization optimization (#7109 )	2024-08-18 17:57:20 -07:00
bnellnm	9f69856356	[Kernel] register punica functions as torch ops (#7591 )	2024-08-16 13:59:38 -07:00
William Lin	57b7be0e1c	[Speculative decoding] [Multi-Step] decouple should_modify_greedy_probs_inplace (#6971 )	2024-08-09 05:42:45 +00:00
Murali Andoorveedu	6dffa4b0a6	[Bugfix] Fix LoRA with PP (#7292 )	2024-08-08 00:02:27 -07:00
Jee Jee Li	9118217f58	[LoRA] Relax LoRA condition (#7146 )	2024-08-06 01:57:25 +00:00
Jacob Schein	89b8db6bb2	[Bugfix] Specify device when loading LoRA and embedding tensors (#7129 ) Co-authored-by: Jacob Schein <jacobschein@Jacobs-MacBook-Pro-2.local>	2024-08-05 16:35:47 -07:00
Jee Jee Li	f80ab3521c	Clean up remaining Punica C information (#7027 )	2024-08-04 15:37:08 -07:00

1 2

100 Commits