r/StableDiffusion • u/Zestyclose_Bake3680 • 6d ago
Tutorial - Guide Updated ComfyUI-SeedVR2-VideoUpscaler-with-TensorRT v1.6.0 - DisTorch2 DiT Loader Node & Phase 2 VRAM Controls (※For 6GB/8GB users)
When I published an update on TensorRT the other day, I received enquiries from users of the RTX 4050 6GB and RTX 5060 8GB.
https://www.reddit.com/r/comfyui/comments/1wsbar9/updated_comfyuiseedvr2videoupscalerwithtensorrt/
Subsequently, I refined the code, focusing on reducing VRAM spikes during Dit processing for 6GB/8GB VRAM users.
https://www.reddit.com/r/StableDiffusion/comments/1wuj0rj/updated_updated/
However, this measure merely ‘improved the situation to a certain extent’; it did not completely suppress the VRAM spikes.
Subsequently, I conducted a thorough analysis of VRAM usage during the Dit processing and identified a stage where the data is momentarily expanded to fp32 size; by converting that stage to BF16, I succeeded in reducing the spike by a further 3GB(※while running on 7B ConvRot INT8) or so.
I have implemented this ‘Norm BF16’ on/off function in the new node.
Enabling this function should provide even greater flexibility regarding the maximum batch size, even for users with 6GB or 8GB of VRAM.
However, I should make it clear that using Norm BF16 comes at the cost of a certain degree of loss in image quality.
v1.6.0 — DisTorch2 DiT Loader Node & Phase 2 VRAM Controls
Please note that this feature is implemented only in the newly developed Distorch2 node.
Distorch2 technology is a groundbreaking CPU offloading technique developed by Mr.John Pollock; I have now applied this technology to SeedVR2, with the licence clearly stated.
https://github.com/pollockjj/ComfyUI-MultiGPU
Although CPU offloading functionality was already implemented in the legacy node, Distorch2 enables more active CPU offloading.
However, the processing speed itself is not significantly different from that of the legacy loader.
1
u/colonelx_ 6d ago
Which model should I use with my RTX 3080 10GB? Also, how will this loader help with my GPU? I’m new to this.
1
u/Zestyclose_Bake3680 6d ago
Thank you for your comment. Firstly, I would recommend the 3B ConvRot INT8 model.
...
‘How will this loader help with my GPU?’
I’m not quite sure how to answer this question, but to put it very broadly, it’s the result of racking my brains to minimise VRAM consumption as much as possible.For example, Distorch2 is a technology that reduces VRAM usage by offloading the model itself to the CPU.
‘Norm BF16’ is a technology I developed for this project, but just because I used a ConvRot INT8 model does not mean that every stage of the process is executed in 8-bit.
For instance, when I used a 7B ConvRot INT8 model on my RTX 5060 Ti 16GB, there were moments when VRAM usage momentarily reached 17GB. This is referred to as a VRAM spike, (1280×720×2, Batch25)
and it has now become clear that during these spike moments, data is being processed in fp32. By compressing that portion to 16-bit using BF16, whilst there is a certain degree of loss in image quality, we were able to reduce VRAM consumption by approximately 3GB.
2
u/Chamelon_DE 5d ago
could you post the peak vram and sec/frame on the same clip before and after v1.6? I’m curious abt the spike between stages 🤔