947 Commits

Author SHA1 Message Date
comfyanonymous
d4ca060482 Lower lowvram memory to 1/3 of free memory. 2024-08-03 15:14:07 -04:00
comfyanonymous
f27f0f8a48 Fix some issues. 2024-08-03 15:06:40 -04:00
comfyanonymous
e83d97d649 Cap lowvram to half of free memory. 2024-08-03 14:50:20 -04:00
comfyanonymous
753d9640f9 Automatically use fp8 for diffusion model weights if:
Checkpoint contains weights in fp8.

There isn't enough memory to load the diffusion model in GPU vram.
2024-08-03 13:45:19 -04:00
comfyanonymous
2181bb7a57 Load T5 in fp8 if it's in fp8 in the Flux checkpoint. 2024-08-03 12:39:33 -04:00
comfyanonymous
2878ffd75a More aggressive batch splitting. 2024-08-03 11:53:30 -04:00
comfyanonymous
b0eaa09c6a Better per model memory usage estimations. 2024-08-02 18:09:24 -04:00
comfyanonymous
e2284e5609 Tweak regular SD memory formula. 2024-08-02 17:34:30 -04:00
comfyanonymous
ebc7233402 Better Flux vram estimation. 2024-08-02 17:02:35 -04:00
Alexander Brown
90267ffb3a Fix clip_g/clip_l mixup (#4168) 2024-08-01 21:40:56 -04:00
comfyanonymous
a06d0ff763 Hack to make all resolutions work on Flux models. 2024-08-01 21:39:18 -04:00
comfyanonymous
49c0c333d3 Tweak the memory usage formulas for Flux and SD. 2024-08-01 17:53:45 -04:00
comfyanonymous
f176649e08 Make ComfyUI split batches a higher priority than weight offload. 2024-08-01 16:39:59 -04:00
comfyanonymous
516d71d227 Fast preview support for Flux. 2024-08-01 16:28:11 -04:00
comfyanonymous
f5b2fc6e6b Fix bfloat16 potentially not being enabled on mps. 2024-08-01 16:18:44 -04:00
comfyanonymous
bd6f023ff6 Try to fix mac issue. 2024-08-01 13:41:27 -04:00
comfyanonymous
5606d85b07 Add a way to load the diffusion model in fp8 with UNETLoader node. 2024-08-01 13:30:51 -04:00
comfyanonymous
30a15dd5e1 Better Mac support on flux model. 2024-08-01 13:10:50 -04:00
comfyanonymous
faad5f1cdb Make lowvram more aggressive on low memory machines. 2024-08-01 12:11:57 -04:00
comfyanonymous
5d5b9dcf89 Fix .sft file loading (they are safetensors files). 2024-08-01 11:32:58 -04:00
comfyanonymous
d6d48306d8 Load flux t5 in fp8 if weights are in fp8. 2024-08-01 11:05:56 -04:00
comfyanonymous
96dbbd23d8 Fix old python versions no longer working. 2024-08-01 09:57:20 -04:00
comfyanonymous
fc6a27f36c Basic Flux Schnell and Flux Dev model implementation. 2024-08-01 09:49:29 -04:00
comfyanonymous
339a710a7a Mac supports bf16 just make sure you are using the latest pytorch. 2024-08-01 09:42:17 -04:00
comfyanonymous
78c273b0ba Make lowvram less aggressive when there are large amounts of free memory. 2024-08-01 03:58:58 -04:00
comfyanonymous
a4b37c2939 Fix to get fp8 working on T5 base. 2024-07-31 02:00:19 -04:00
comfyanonymous
c5b754fbb6 Fix hunyuan dit text encoder weights always being in fp32. 2024-07-31 01:34:57 -04:00
comfyanonymous
4e9abfb0d5 Lower CLIP memory usage by a bit. 2024-07-31 01:32:35 -04:00
comfyanonymous
7f0000bc5a Lower T5 memory usage by a few hundred MB. 2024-07-31 00:52:34 -04:00
comfyanonymous
671fba8c0d Fix potential issue with non clip text embeddings. 2024-07-30 14:41:13 -04:00
comfyanonymous
36b0b2b215 Use common function for casting weights to input. 2024-07-30 10:49:14 -04:00
comfyanonymous
d0cc951625 Remove unnecessary code. 2024-07-30 05:01:34 -04:00
comfyanonymous
7f112b9d5d Improve artifacts on hydit, auraflow and SD3 on specific resolutions.
This breaks seeds for resolutions that are not a multiple of 16 in pixel
resolution by using circular padding instead of reflection padding but
should lower the amount of artifacts when doing img2img at those
resolutions.
2024-07-29 20:48:50 -04:00
comfyanonymous
79ea2d3d12 Refactor: Move sd2_clip.py to text_encoders folder. 2024-07-28 01:19:20 -04:00
comfyanonymous
f4a8f8db50 Don't treat Bert model like CLIP.
Bert can accept up to 512 tokens so any prompt with more than 77 should
just be passed to it as is instead of splitting it up like CLIP.
2024-07-26 13:08:12 -04:00
comfyanonymous
a787d2dfc0 Let hunyuan dit work with all prompt lengths. 2024-07-26 12:11:32 -04:00
comfyanonymous
1824ded587 Hunyuan dit can now accept longer prompts. 2024-07-26 11:52:58 -04:00
comfyanonymous
0a7029fabf Own BertModel implementation that works with lowvram. 2024-07-26 04:47:17 -04:00
comfyanonymous
080a8326a3 Hunyuan DiT lora support. 2024-07-25 22:42:54 -04:00
comfyanonymous
06fb93c0d1 Basic hunyuan dit implementation. (#4102)
* Let tokenizers return weights to be stored in the saved checkpoint.

* Basic hunyuan dit implementation.

* Fix some resolutions not working.

* Support hydit checkpoint save.

* Init with right dtype.

* Switch to optimized attention in pooler.

* Fix black images on hunyuan dit.
2024-07-25 18:21:08 -04:00
comfyanonymous
15a95668ae Let tokenizers return weights to be stored in the saved checkpoint. 2024-07-25 10:52:09 -04:00
comfyanonymous
33e584b3bc Make it possible to load tokenizer data from checkpoints. 2024-07-24 16:43:53 -04:00
comfyanonymous
ba8920ca60 Remove duplicate code. 2024-07-24 01:12:59 -04:00
comfyanonymous
fe94df0894 Support MT5. 2024-07-23 15:35:28 -04:00
comfyanonymous
f5ebfaba43 Allow SPieceTokenizer to load model from a byte string. 2024-07-23 14:17:42 -04:00
comfyanonymous
b63c469865 More generic unet prefix detection code. 2024-07-23 14:13:32 -04:00
comfyanonymous
a7f1068957 Rename LLAMATokenizer to SPieceTokenizer. 2024-07-22 12:21:45 -04:00
comfyanonymous
e96e0fdf8b "auto" type is only relevant to the SetUnionControlNetType node. 2024-07-22 11:30:38 -04:00
Chenlei Hu
c79f809b56 Add error message on union controlnet (#4081) 2024-07-22 11:27:32 -04:00
comfyanonymous
78746a7499 Only append zero to noise schedule if last sigma isn't zero. 2024-07-20 12:37:30 -04:00