Simon Lui
|
87e7b35146
|
Change torch.xpu to ipex.optimize, xpu device initialization and remove workaround for text node issue from older IPEX. (#3388)
|
2024-05-02 03:26:50 -04:00 |
|
comfyanonymous
|
982b8e2276
|
Fix some memory related issues.
|
2024-04-14 12:08:58 -04:00 |
|
comfyanonymous
|
d93e3f1991
|
Fix issue with controlnet models getting loaded multiple times.
|
2024-04-06 18:38:39 -04:00 |
|
comfyanonymous
|
cd0be2a31e
|
Fix some performance issues with weight loading and unloading.
Lower peak memory usage when changing model.
Fix case where model weights would be unloaded and reloaded.
|
2024-03-28 18:04:42 -04:00 |
|
comfyanonymous
|
af7900cd79
|
Optimize memory unload strategy for more optimized performance.
|
2024-03-24 02:36:30 -04:00 |
|
comfyanonymous
|
2db9fdbacc
|
Fix regression with model merging.
|
2024-03-20 13:56:12 -04:00 |
|
comfyanonymous
|
1681937288
|
Don't unload model weights for non weight patches.
|
2024-03-20 02:27:58 -04:00 |
|
comfyanonymous
|
7e04cdafc0
|
Lower memory usage for loras in lowvram mode at the cost of perf.
|
2024-03-13 20:07:27 -04:00 |
|
comfyanonymous
|
46e5a518c3
|
Change log levels.
Logging level now defaults to info. --verbose sets it to debug.
|
2024-03-11 13:54:56 -04:00 |
|
comfyanonymous
|
35d0fd9a5d
|
Replace prints with logging and add --verbose argument.
|
2024-03-10 12:14:23 -04:00 |
|
comfyanonymous
|
8726a79de9
|
Add some tesla pascal GPUs to the fp16 working but slower list.
|
2024-03-02 17:16:31 -05:00 |
|
comfyanonymous
|
f4e74ae786
|
Enable fp16 by default on mps.
|
2024-02-19 12:00:48 -05:00 |
|
comfyanonymous
|
cabdbc2529
|
Manual cast for bf16 on older GPUs.
|
2024-02-17 09:01:17 -05:00 |
|
comfyanonymous
|
b3c820556b
|
Make --force-fp32 disable loading models in bf16.
|
2024-02-16 23:01:54 -05:00 |
|
comfyanonymous
|
deebc95781
|
Stable Cascade Stage C.
|
2024-02-16 10:55:08 -05:00 |
|
comfyanonymous
|
a3ada6af3b
|
Small refactor of is_device_* functions.
|
2024-02-15 21:10:10 -05:00 |
|
comfyanonymous
|
1bbaf69344
|
Don't use is_bf16_supported to check for fp16 support.
|
2024-02-04 20:53:35 -05:00 |
|
comfyanonymous
|
5efb046482
|
Speed up SDXL on 16xx series with fp16 weights and manual cast.
|
2024-02-04 13:23:43 -05:00 |
|
comfyanonymous
|
32be22279f
|
Always use fp16 for the text encoders.
|
2024-02-02 10:02:49 -05:00 |
|
comfyanonymous
|
96232a0a45
|
Only auto enable bf16 VAE on nvidia GPUs that actually support it.
|
2024-01-15 03:10:22 -05:00 |
|
comfyanonymous
|
25ec26ff9c
|
Add argument to run the VAE on the CPU.
|
2023-12-30 05:49:07 -05:00 |
|
comfyanonymous
|
48d2b242ee
|
Load weights that can't be lowvramed to target device.
|
2023-12-28 21:41:10 -05:00 |
|
comfyanonymous
|
0baa3cc687
|
--disable-smart-memory now unloads everything like it did originally.
|
2023-12-23 04:25:06 -05:00 |
|
comfyanonymous
|
59e4fbf255
|
Greatly improve lowvram sampling speed by getting rid of accelerate.
Let me know if this breaks anything.
|
2023-12-22 14:38:45 -05:00 |
|
comfyanonymous
|
65ada00818
|
Add --deterministic option to make pytorch use deterministic algorithms.
|
2023-12-17 16:59:21 -05:00 |
|
comfyanonymous
|
45acae2d06
|
Add an option --fp16-unet to force using fp16 for the unet.
|
2023-12-11 18:36:29 -05:00 |
|
comfyanonymous
|
f64ac62540
|
Use faster manual cast for fp8 in unet.
|
2023-12-11 18:24:44 -05:00 |
|
comfyanonymous
|
58b39151a4
|
Switch text encoder to manual cast.
Use fp16 text encoder weights for CPU inference to lower memory usage.
|
2023-12-10 23:00:54 -05:00 |
|
comfyanonymous
|
260b25aef8
|
Disable non blocking on mps.
|
2023-12-10 01:30:35 -05:00 |
|
comfyanonymous
|
63349484b8
|
Make --gpu-only put intermediate values in GPU memory instead of cpu.
|
2023-12-08 02:35:45 -05:00 |
|
comfyanonymous
|
1f1ef695bb
|
Slightly faster lora applying.
|
2023-12-06 05:13:14 -05:00 |
|
comfyanonymous
|
dfa7737afb
|
Use .itemsize to get dtype size for fp8.
|
2023-12-04 11:52:06 -05:00 |
|
comfyanonymous
|
080f5f4e84
|
UNET weights can now be stored in fp8.
--fp8_e4m3fn-unet and --fp8_e5m2-unet are the two different formats
supported by pytorch.
|
2023-12-04 11:10:00 -05:00 |
|
comfyanonymous
|
c013b8e94c
|
Add some command line arguments to store text encoder weights in fp8.
Pytorch supports two variants of fp8:
--fp8_e4m3fn-text-enc (the one that seems to give better results)
--fp8_e5m2-text-enc
|
2023-11-17 02:56:59 -05:00 |
|
comfyanonymous
|
9f546f0cb3
|
Disable xformers when it can't load properly.
|
2023-11-13 12:31:10 -05:00 |
|
comfyanonymous
|
21ea9c3263
|
Allow different models to estimate memory usage differently.
|
2023-11-12 04:03:52 -05:00 |
|
comfyanonymous
|
7f861d49fd
|
Empty the cache when torch cache is more than 25% free mem.
|
2023-10-22 13:58:12 -04:00 |
|
comfyanonymous
|
daabf7fd3a
|
Add some Quadro cards to the list of cards with broken fp16.
|
2023-10-16 16:48:46 -04:00 |
|
comfyanonymous
|
db653f4908
|
Add a --bf16-unet to test running the unet in bf16.
|
2023-10-13 14:51:10 -04:00 |
|
comfyanonymous
|
728139a5b9
|
Refactor code so model can be a dtype other than fp32 or fp16.
|
2023-10-13 14:41:17 -04:00 |
|
comfyanonymous
|
494ddf7717
|
pytorch_attention_enabled can now return True when xformers is enabled.
|
2023-10-11 21:30:57 -04:00 |
|
comfyanonymous
|
18e4504de7
|
Pull some small changes from the other repo.
|
2023-10-11 20:38:48 -04:00 |
|
Simon Lui
|
47164eb065
|
Allow Intel GPUs to LoRA cast on GPU since it supports BF16 natively.
|
2023-09-22 21:11:27 -07:00 |
|
comfyanonymous
|
795f5b3163
|
Only do the cast on the device if the device supports it.
|
2023-09-20 17:52:41 -04:00 |
|
comfyanonymous
|
cdbbeb584d
|
Enable pytorch attention by default on xpu.
|
2023-09-17 04:09:19 -04:00 |
|
comfyanonymous
|
cac135d12f
|
Don't run text encoders on xpu because there are issues.
|
2023-09-14 12:16:07 -04:00 |
|
comfyanonymous
|
ef0c0892f6
|
Add a force argument to soft_empty_cache to force a cache empty.
|
2023-09-04 00:58:18 -04:00 |
|
Simon Lui
|
1148c2dec7
|
Some fixes to generalize CUDA specific functionality to Intel or other GPUs.
|
2023-09-02 18:22:10 -07:00 |
|
comfyanonymous
|
ae3f7060d8
|
Enable bf16-vae by default on ampere and up.
|
2023-08-27 23:06:19 -04:00 |
|
comfyanonymous
|
90bfcef833
|
Fix lowvram model merging.
|
2023-08-26 11:52:07 -04:00 |
|