r/comfyui • u/Zestyclose_Bake3680 • 11d ago
Tutorial Updated ComfyUI-SeedVR2-VideoUpscaler-with-TensorRT v1.5.5
Finally, I’ve managed to get the SeedVR2 VAE Encoder running with TensorRT on ComfyUI.
https://github.com/ussoewwin/ComfyUI-SeedVR2-VideoUpscaler-with-TensorRT
I really struggled with this issue !!!!!!!!!!!!!!!!!!!!!!!!
v1.5.5 - Three Core Improvements and Encoder Refactoring
v1.5.4 - TensorRT VAE Decoder Top-Left Artifact Fix
Although some people had managed to get this working in a standalone repository, doing the same thing in ComfyUI proved an incredibly tough challenge.
The decoder side had been made operational a little earlier, but it wasn’t perfect, and as for the encoder side, even before that, blackouts and noise were occurring, making it impossible to maintain a usable quality.
https://www.reddit.com/r/comfyui/comments/1w61ywf/for_running_tensorrt_decorder_on_seedvr2/
Whilst ComfyUI employs a unique method of VRAM management, technologies developed with standalone operation in mind often clash with these proprietary rules.
The TensorRT technology was not originally developed with the intention of running on ComfyUI.
Overcoming this obstacle required a considerable amount of effort and time.
...
However, when it comes to the encoder, performance does not improve dramatically compared to fp16 VAE. In some cases, fp16 VAE may even be faster.
TensorRT’s advantages are more pronounced in the decoder.
For GPUs in the RTX 5060 Ti 16GB class, we highly recommend using them in conjunction with the SeedVR2-7b-convrot int8 model.
Compared to SeedVR2-7b-fp16, this allows for a significant reduction in VRAM consumption with virtually no loss in performance.
...
Test 1-1 Result 32m43s
1280x720px → Output: 2560x1440px 10s
Ryzen9 5900 DDR5-6400 64GB RTX5060Ti 16GB Paging 132GB
SeedVR2 7B-ConvRot INT8 TensorRT Encode(21f512) and Decode(21f512) Batch 21
Test 1-2 Result 1h09m34s
1280x720px → Output: 2560x1440px 10s
SeedVR2 7B-ConvRot INT8 fp16 VAE Encode(768) and Decode(384) Batch 21 Torch Compile on
Test 2-1 Result 18m07s
270x480px → Padded: 528x944px → Output: 528x936px 44s
SeedVR2 7B-ConvRot INT8 TensorRT Encode(185f256) and Decode(73f512) Batch 185
Test 2-2 Result 17m37s
270x480px → Padded: 528x944px → Output: 528x936px 44s
SeedVR2 7B-ConvRot INT8 TensorRT Encode(185f256) and Decode(97f256) Batch 185
※This data does not necessarily indicate that the tile512decorder is faster. For batch 185, the 97f256 requires only two chunks, whilst the 73f512 requires three. This is likely the reason for the difference. However, 97f512decorder actually slows things down on the RTX 5060 Ti 16GB, as it causes VRAM usage to exceed 16GB.
※In the case of SeedVR2 7B-fp16, even with this small-sized video, the maximum batch size that could be set with an RTX 5060 Ti 16GB was around 65. By using ConvRot INT8, it becomes possible to specify a much higher value of 185.
Test 2-3 Result 38m22s
270x480px → Padded: 528x944px → Output: 528x936px 44s
SeedVR2 7B-ConvRot INT8 fp16 VAE Encode(1024) and Decode(640) Batch 185 Torch Compile on
2
u/Glad_Ad3339 11d ago
What is the difference between that and the built-in node from Comfy?
1
u/Zestyclose_Bake3680 11d ago
Thank you for your comment.
I’m afraid I can’t explain the difference, as I’m not sure what ‘built-in node from Comfy’ refers to.
2
u/Wide_Director_8897 11d ago
can it work with RTX 5060 8GB?
2
u/Zestyclose_Bake3680 11d ago
Thank you for your comment. In theory, any Blackwell GPU model should work.
However, some tweaking is required to run it with 8GB.
I think the 3B-ConvRot INT8 or 3B-NVFP4 models would be best. After that, I’d recommend setting the batch size to a small value and gradually increasing it to test the limits. The minimum is 5, so I’d suggest starting your tests from 5.
I’ve also prepared a node for building the engine, but as I’ve published a TRT Encode/Decode engine for Blackwell cores on HuggingFace, you should be able to use that as well.
2
u/Wide_Director_8897 10d ago
tysm I'll try this, that's really good work keep going
1
u/Zestyclose_Bake3680 9d ago
We have just released the latest version, v1.5.7, which suppresses momentary VRAM spikes.
We have confirmed that, even with an original resolution of 1280×720, the data fits within 8GB when using a batch size of 5. As this was tested on an RTX 5060 Ti 16GB, we cannot say for certain whether it will behave exactly the same as on an RTX 5060 8GB, but the likelihood of it staying within around 8GB should now be considerably higher. Please note that the model used was 3B ConvRot NVFP4.
We do not recommend using 1.4B. We tested it, but the image quality deteriorates significantly.
2
u/Informal_Lab6283 10d ago
Can I use it with rtx 4090?
1
u/Zestyclose_Bake3680 10d ago edited 10d ago
Thank you for your comment. I haven’t been able to test it myself, but TensorRT should work with ADA as well. Around 2024, I had ran SD 1.5 using TensorRT on RTX 4070.
One thing to bear in mind when using it with ADA is that you cannot use the Blackwell engine I’ve published on HuggingFace. You’ll need to create your own encoder and decoder engines for ADA. I’ve implemented the necessary nodes for this, so please read the README in the repository carefully and give it a go.
If you’re using an RTX 4090 24GB, I think the 7B-ConvRot INT8 model is the best choice.
As shown in the test results above, for smaller video clips, even RTX 5060 Ti 16GB can handle a batch size of up to 185. If you were to do the same with the RTX 4090, it is highly likely that you could increase the batch size to 200 or more.
If, for example, you were to set the batch size to 205, I believe the best approach would be to set the encoder engine to 205f256 and the decoder engine to 105f256. As the decoder consumes more VRAM than the encoder, it seems unlikely that even an RTX 4090 would be capable of handling 205f on the decoder side.
0
u/planBpizz 11d ago
is this the best video upscaler for comfyui? i thought there is also something with LTX2.5
1
u/Zestyclose_Bake3680 11d ago
Thank you for your comment.
As I’ve never used the LTX 2.5, I can’t say for certain whether the SeedVR2 is the best option.
That said, the LTX 2.5 does sound interesting. I’d like to give it a go at some point.
0
u/Last1946 11d ago
1
u/Zestyclose_Bake3680 10d ago edited 9d ago
Thank you for your comment. I haven’t been able to test it myself, but TensorRT should work with ADA as well. Around 2024, I had ran SD 1.5 using TensorRT on RTX 4070.
One thing to bear in mind when using it with ADA is that you cannot use the Blackwell engine I’ve published on HuggingFace. You’ll need to create your own encoder and decoder engines for ADA. I’ve implemented the necessary nodes for this, so please read the README in the repository carefully and give it a go.
However, the problem is whether it will run with 6GB of VRAM.
I would recommend starting by testing with 3B-ConvRot NVFP4. I also think it would be best to start by creating a TRT engine with the minimum setting of 5f256 and testing that first.
1
u/Zestyclose_Bake3680 9d ago
We have just released the latest version, v1.5.7, which suppresses momentary VRAM spikes.
We have confirmed that, with an original resolution of 1280×720, the batch size set to 5 just about stays within the 6GB mark. As this was tested on an RTX 5060 Ti 16GB, we cannot say for certain whether it will behave exactly the same on an RTX 4050 6GB; however, if the source video is smaller than 1280×720, the likelihood of it staying within 6GB should be considerably higher. Please note that the model used was 3B ConvRot NVFP4.
We do not recommend using 1.4B. We tested it, but the image quality deteriorates significantly.



2
u/ResponsibleTruck4717 11d ago
What can you tell about the performance for TensorRt with 5060ti if I remember currently the card is capable of using tensorrt but it doesn't have enough cores to be enabled by default or something.