Lucas Wilkinson
|
a8d604ca2a
|
[Misc] Disambiguate quantized types via a new ScalarType (#6396)
|
2024-08-02 13:51:58 -07:00 |
|
Tyler Michael Smith
|
61a97c32f6
|
[Kernel] Fix marlin divide-by-zero warnings (#6904)
|
2024-07-30 01:26:07 +00:00 |
|
Alexander Matveev
|
75acdaa4b6
|
[Kernel] Increase precision of GPTQ/AWQ Marlin kernel (#6795)
|
2024-07-27 17:52:33 -04:00 |
|
Alexander Matveev
|
396d92d5e0
|
[Kernel][Core] Add AWQ support to the Marlin kernel (#6612)
|
2024-07-21 19:41:42 -04:00 |
|
bnellnm
|
5467ac3196
|
[Kernel][Misc] Use TORCH_LIBRARY instead of PYBIND11_MODULE for custom ops (#5047)
|
2024-06-09 16:23:30 -04:00 |
|
Michael Goin
|
5f6d10c14c
|
[CI/Build] Enforce style for C++ and CUDA code with clang-format (#4722)
|
2024-05-22 07:18:41 +00:00 |
|
Alexander Matveev
|
da5a0b539d
|
Remove marlin warning (#4918)
|
2024-05-20 14:55:34 +00:00 |
|
Jinzhen Lin
|
99caa49106
|
[Kernel] add bfloat16 support for gptq marlin kernel (#4788)
|
2024-05-16 09:55:29 -04:00 |
|
alexm-nm
|
e288df0632
|
[Bugfix] Fine-tune gptq_marlin configs to be more similar to marlin (#4626)
|
2024-05-08 17:14:31 -07:00 |
|
alexm-nm
|
7038e8b803
|
[Kernel] Support running GPTQ 8-bit models in Marlin (#4533)
|
2024-05-02 12:56:22 -04:00 |
|
Robert Shaw
|
73c8d677e5
|
[Kernel] Marlin Expansion: Support AutoGPTQ Models with Marlin (#3922)
Co-authored-by: alexm <alexm@neuralmagic.com>
Co-authored-by: mgoin <michael@neuralmagic.com>
|
2024-04-29 09:35:34 -07:00 |
|