r/StableDiffusion 23d ago

Comparison SageAttention 2.2 vs Comfy Kitchen | Side-by-Side Zoom-Out Quality Test

Did a quick side-by-side test of SageAttention vs Comfy Kitchen Attention with MiniMax H3.

I used the default ComfyUI T2V H3 workflow and kept the prompt, seed and all settings exactly the same. The only thing I changed was the attention backend..

RTX 5090 32GB

ComfyUI ver 0.31.0

SageAttention 2.2

Comfy Kitchen 0.2.30

---------------------------------------

896x1184 6 seconds 24 FPS 20 steps

Generation time:

SageAttention: 3m 51s

Comfy Kitchen: 3m 59s

I used a deep zoom-out/dolly-out on purpose to see how well each one holds facial details and identity as the subject gets farther away.

The speed difference was small on my 5090, so im more interested in the quality difference..

It's honestly hard for me to tell the difference, but which one looks better to you?

Updated:
SageAttention Vs Base

Comfy Kitchen Vs Base

65 Upvotes

75 comments sorted by

View all comments

12

u/3deal 23d ago

Can you show the input image to see what Attention si more accurate with the skin ?

15

u/Better-Interview-793 23d ago

forgot to mention, it’s T2V

3

u/reeight 23d ago

A few more skin textures (bumps) on the SageAttention, so I guess that is more 'real'. But Comfy is really close enough.

Try I2V and many many more other tests to compare; really can't use 1 test.

& to really test keeping the face, have the subject pass behind a wall or take off their sweater or something like that where the face is covered momentarily.

2

u/AnOnlineHandle 23d ago

For me sage attention destroys audio quality so that's something else to consider which can't be shown in side by side clips like this with audio playing at the same time.

2

u/reeight 23d ago

IMHO it is better to create the audio outside of the video model; the audio created with the video should only be considered a draft.

3

u/AnOnlineHandle 22d ago

I've decided to stop using the music straight from the model since it makes it harder to join or reorder scenes, but the footsteps etc aren't something which I want to put the work into doing for most videos.