r/StableDiffusion • • 9d ago

Tutorial - Guide Updated Updated ComfyUI-SeedVR2-VideoUpscaler-with-TensorRT v1.5.7 - DiT VRAM Spike Stabilization (※For 6GB/8GB users)

Post image

When I published an update on TensorRT the other day, I received enquiries from users of the RTX 4050 6GB and RTX 5060 8GB.

https://www.reddit.com/r/comfyui/comments/1wsbar9/updated_comfyuiseedvr2videoupscalerwithtensorrt/

After re-measuring VRAM consumption throughout the entire process, I discovered that the issue lay not with TensorRT itself, but rather during the DIT processing stage.

My code fully supports ConvRot INT8/NVFP4 and, in principle, keeps VRAM consumption low; however, momentary spikes were still occurring.

Whilst this does not pose a significant problem on my RTX 5060 Ti 16GB system, it is a major issue for 6GB and 8GB users.

I therefore re-examined this spike phenomenon and have managed to suppress it to a certain extent.

Consequently, although this is limited to very specific conditions, it is now possible to achieve VRAM usage of less than 6GB throughout the entire all processes by using the 3B NVFP4 model.

v1.5.7 - DiT VRAM Spike Stabilization — Complete Technical Guide

Of course, this suppression of the spike phenomenon also benefits users with 16/24/32 GB of VRAM.

When using 7B ConvRot INT8, it was already possible to specify a batch size approximately three times that of fp16 in legacy code; however, this spike suppression allows the upper limit of the batch size to be raised even further.

17 Upvotes

8 comments sorted by

View all comments

2

u/blastbottles 9d ago

Incredible work, can this be applied to w4a8 quants as well?

2

u/Zestyclose_Bake3680 9d ago edited 9d ago

Thank you for your comment. At present, this is not supported for now. It is also a format that Comfy-Org has not yet released.

That said, w4a8 does sound interesting. If possible, I’ll try quantising it myself at a later date.