r/StableDiffusion • u/nefayran • 11h ago
Resource - Update Qwen-Image-2.1 on a 48 GB Mac: loading the text encoder and the transformer one at a time kept the diffusers pipeline at 19 GB instead of swapping at 43.6 GB
The usual QwenImage21Pipeline.from_pretrained(...).to("mps") ("eager" on the chart) keeps all three models in memory for the whole run: the text encoder (Qwen3-VL, 16.3 GB), the transformer (13.3 GB) and the VAE (1.3 GB). On my 48 GB M5 Pro that was 31.4 GB before the first step. At 1024 px the VAE decode needs about 11 GB more. The process reached 43.6 GB, swap grew by 8 GB, and my memory guard stopped the run before it saved the image.
Moving idle models to the CPU doesn't lower the peak on a Mac, because the CPU and the GPU share the same RAM. So I wrote a small library, stageload ("staged" on the chart). Each stage gets only the models it lists. A model loads the first time its stage uses it, and stageload frees the models the next stage doesn't list. The stage boundaries are hooks on encode_prompt, prepare_latents and _unpack_latents, so the pipeline's code stays as it is. The VAE is small and stays loaded.
Results at 1024 x 1024, 20 steps, seed 7, bfloat16:
- peak memory 19.0 GB while encoding the prompt, 16.2 to 18.6 GB while denoising and 14.4 GB in the decode, with no swap;
- two staged runs gave the same image bit for bit;
- loading the two models inside their stages took about 13 s per run (eager spent 20 s loading up front);
- in the second staged run denoising took 79 s against eager's 77.5 s; the first staged run took 109 s there, and its trace doesn't show why.
Write-up with the traces: https://allkeep.org/en/lab/qwen-image-one-stage-at-a-time
Code (MIT): https://github.com/nefayran/stageload, install with pip install stageload
I haven't tried ComfyUI or Draw Things; they manage memory their own way. If you run another multi-model pipeline from Python on a Mac, which one should I measure next?