184 Commits

Author SHA1 Message Date
comfyanonymous
cbb1bde521 Fix some issues with inference slowing down. 2024-08-10 16:21:25 -04:00
comfyanonymous
1280396253 Fix regression. 2024-08-09 03:36:40 -04:00
comfyanonymous
114af84d12 Try to improve inference speed on some machines. 2024-08-08 17:29:27 -04:00
comfyanonymous
4f196c19a3 Fix. 2024-08-08 15:16:51 -04:00
comfyanonymous
ede5622146 Partial model shift support. 2024-08-08 14:45:06 -04:00
comfyanonymous
5a25225024 Make supported_dtypes a priority list. 2024-08-07 15:00:06 -04:00
comfyanonymous
001ff4d6a6 Fix OOMs happening in some cases.
A cloned model patcher sometimes reported a model was loaded on a device
when it wasn't.
2024-08-06 13:36:04 -04:00
comfyanonymous
5228fbf0ad Unload models and load them back in lowvram mode no free vram. 2024-08-06 03:22:39 -04:00
comfyanonymous
f7183e2c73 Improve performance on some lowend GPUs. 2024-08-05 16:24:04 -04:00
comfyanonymous
10ad73d520 Fix crash. 2024-08-03 16:55:38 -04:00
comfyanonymous
e7287d2fc2 Tweak lowvram memory formula. 2024-08-03 16:44:50 -04:00
comfyanonymous
d4ca060482 Lower lowvram memory to 1/3 of free memory. 2024-08-03 15:14:07 -04:00
comfyanonymous
f27f0f8a48 Fix some issues. 2024-08-03 15:06:40 -04:00
comfyanonymous
e83d97d649 Cap lowvram to half of free memory. 2024-08-03 14:50:20 -04:00
comfyanonymous
753d9640f9 Automatically use fp8 for diffusion model weights if:
Checkpoint contains weights in fp8.

There isn't enough memory to load the diffusion model in GPU vram.
2024-08-03 13:45:19 -04:00
comfyanonymous
f176649e08 Make ComfyUI split batches a higher priority than weight offload. 2024-08-01 16:39:59 -04:00
comfyanonymous
f5b2fc6e6b Fix bfloat16 potentially not being enabled on mps. 2024-08-01 16:18:44 -04:00
comfyanonymous
faad5f1cdb Make lowvram more aggressive on low memory machines. 2024-08-01 12:11:57 -04:00
comfyanonymous
d6d48306d8 Load flux t5 in fp8 if weights are in fp8. 2024-08-01 11:05:56 -04:00
comfyanonymous
339a710a7a Mac supports bf16 just make sure you are using the latest pytorch. 2024-08-01 09:42:17 -04:00
comfyanonymous
78c273b0ba Make lowvram less aggressive when there are large amounts of free memory. 2024-08-01 03:58:58 -04:00
comfyanonymous
160c550977 Use fp16 as the default vae dtype for the audio VAE. 2024-06-16 13:12:54 -04:00
comfyanonymous
12e44a5d50 Add a --force-channels-last to inference models in channel last mode. 2024-06-15 01:08:12 -04:00
Simon Lui
20052d83b5 Exempt IPEX from non_blocking previews fixing segmentation faults. (#3708) 2024-06-13 18:51:14 -04:00
comfyanonymous
917810a2c6 Load the SD3 T5xxl model in the same dtype stored in the checkpoint. 2024-06-11 17:03:26 -04:00
comfyanonymous
b6a9851535 Add function to get the list of currently loaded models. 2024-06-05 23:25:16 -04:00
comfyanonymous
20f463902f pytorch xpu should be flash or mem efficient attention? 2024-06-04 17:44:14 -04:00
comfyanonymous
5f826c9921 Add an annoying print to a function I want to remove. 2024-06-01 12:47:31 -04:00
comfyanonymous
9cea5b91a0 Disable non_blocking when --deterministic or directml. 2024-05-30 11:07:38 -04:00
comfyanonymous
93094050bb Remove some unused imports. 2024-05-27 19:08:27 -04:00
comfyanonymous
94a508adc2 Fix OSX latent2rgb previews. 2024-05-22 13:56:28 -04:00
comfyanonymous
bdbaa4491c Work around black image bug on Mac 14.5 by forcing attention upcasting. 2024-05-21 16:56:33 -04:00
comfyanonymous
bcc5cd02c7 Log the pytorch version. 2024-05-20 06:22:29 -04:00
comfyanonymous
367f84be8d Don't automatically switch to lowvram mode on GPUs with low memory. 2024-05-17 00:31:32 -04:00
Simon Lui
4cd1bc2cce Fix Intel GPU memory allocation accuracy and documentation update. (#3459)
* Change calculation of memory total to be more accurate, allocated is actually smaller than reserved.

* Update README.md install documentation for Intel GPUs.
2024-05-12 06:36:30 -04:00
comfyanonymous
43079b39cd Fix lowvram issue with saving checkpoints.
The previous fix didn't cover the case where the model was loaded in
lowvram mode right before.
2024-05-12 06:13:45 -04:00
comfyanonymous
397b840846 No longer necessary. 2024-05-12 05:34:43 -04:00
comfyanonymous
ae31a1b885 Fix issue with lowvram mode breaking model saving. 2024-05-11 21:55:20 -04:00
Simon Lui
87e7b35146 Change torch.xpu to ipex.optimize, xpu device initialization and remove workaround for text node issue from older IPEX. (#3388) 2024-05-02 03:26:50 -04:00
comfyanonymous
982b8e2276 Fix some memory related issues. 2024-04-14 12:08:58 -04:00
comfyanonymous
d93e3f1991 Fix issue with controlnet models getting loaded multiple times. 2024-04-06 18:38:39 -04:00
comfyanonymous
cd0be2a31e Fix some performance issues with weight loading and unloading.
Lower peak memory usage when changing model.

Fix case where model weights would be unloaded and reloaded.
2024-03-28 18:04:42 -04:00
comfyanonymous
af7900cd79 Optimize memory unload strategy for more optimized performance. 2024-03-24 02:36:30 -04:00
comfyanonymous
2db9fdbacc Fix regression with model merging. 2024-03-20 13:56:12 -04:00
comfyanonymous
1681937288 Don't unload model weights for non weight patches. 2024-03-20 02:27:58 -04:00
comfyanonymous
7e04cdafc0 Lower memory usage for loras in lowvram mode at the cost of perf. 2024-03-13 20:07:27 -04:00
comfyanonymous
46e5a518c3 Change log levels.
Logging level now defaults to info. --verbose sets it to debug.
2024-03-11 13:54:56 -04:00
comfyanonymous
35d0fd9a5d Replace prints with logging and add --verbose argument. 2024-03-10 12:14:23 -04:00
comfyanonymous
8726a79de9 Add some tesla pascal GPUs to the fp16 working but slower list. 2024-03-02 17:16:31 -05:00
comfyanonymous
f4e74ae786 Enable fp16 by default on mps. 2024-02-19 12:00:48 -05:00