mirror of
https://git.datalinker.icu/comfyanonymous/ComfyUI
synced 2026-08-14 14:30:11 +08:00
As commented. The former logic computed independent pieces of QKV in parallel which help more inference intermediates in VRAM spiking VRAM usage. Fully roping Q and garbage collecting the intermediates before touching K reduces the peak inference VRAM usage.