xinyun/ComfyUI-CogVideoXWrapper

mirror of https://git.datalinker.icu/kijai/ComfyUI-CogVideoXWrapper.git synced 2026-06-21 01:37:02 +08:00

Go to file

Jukka Seppänen 49767f1cda

Update readme.md

2024-08-07 02:15:27 +03:00

cleanup, fix vid2vid

2024-08-07 02:11:37 +03:00

__init__.py

initial

2024-08-06 01:56:25 +03:00

.gitattributes

Initial commit

2024-08-06 01:54:04 +03:00

.gitignore

example

2024-08-06 02:02:04 +03:00

nodes.py

cleanup, fix vid2vid

2024-08-07 02:11:37 +03:00

pipeline_cogvideox.py

cleanup, fix vid2vid

2024-08-07 02:11:37 +03:00

readme.md

Update readme.md

2024-08-07 02:15:27 +03:00

requirements.txt

initial

2024-08-06 01:56:25 +03:00

readme.md

WORK IN PROGRESS

Currently requires diffusers with PR: https://github.com/huggingface/diffusers/pull/9082

This is specified in requirements.txt

Uses same T5 model than SD3 and Flux, fp8 works fine too. Memory requirements depend mostly on the video length. VAE decoding seems to be the only big that takes a lot of VRAM when everything is offloaded, peaks at around 13-14GB momentarily at that stage. Sampling itself takes only maybe 5-6GB.

Hacked in img2img to attempt vid2vid workflow, works interestingly with some inputs, highly experimental.

https://github.com/user-attachments/assets/e6951ef4-ea7a-4752-94f6-cf24f2503d83

https://github.com/user-attachments/assets/9e41f37b-2bb3-411c-81fa-e91b80da2559

Original repo: https://github.com/THUDM/CogVideo