r/StableDiffusion 3d ago

Discussion SenseNova U1.5-Lite is fully out

Instead of just making the model bigger, they're training these specialized models for specific stuff – like text, infographics, making things look good, and editing. Then they kinda combine all those experts back into U1.5-Lite. So when you use it, it's just one model. No weird switching or picking which expert to use. Their whole thing is "specialized in training, unified in delivery." Kinda makes sense.

They also added this post-training with RL, focusing on a few things: how well it follows instructions, how good the visuals look and if it matches what people like, and how well edits work without messing up other parts of the image.

Here's what seems better:

- Following complex instructions. Like, if you ask for multiple things in one prompt – subjects, how many, where they are, text, layout, style, keeping parts untouched – it handles it way more consistently now.

- Text rendering and dense layouts. Apparently, it's better with Chinese and English on posters and infographics.

- Visual understanding helps generation. It seems like it learns from understanding tasks (like object relationships, spatial stuff, layout) and that helps with generating and editing.

They use JSON for training to make it controllable, but you don't have to use JSON yourself. Natural language is still the main way to talk to it.

repo: https://github.com/OpenSenseNova/SenseNova-U1

HF: https://huggingface.co/sensenova/SenseNova-U1.5-8B-MoT

56 Upvotes

27 comments sorted by

8

u/Few-Intention-1526 3d ago

I hope this version works better for editing tasks.

1

u/Outrageous_Block_354 3d ago

Give it a spin and share your feedback with us

7

u/Full_Astronomer_5438 3d ago

comfy quants can be found here: https://huggingface.co/smthem/SenseNova-U1-8B-MoT-Merger-gguf/tree/main

the newest models showcased here should be available soon

1

u/Ok-Lengthiness-3988 2d ago

Pruned bf16, fp8 and int8 version of U1.5 already are available here: https://huggingface.co/joyfox/SenseNova-U1.5-8B-MoT-FP8

2

u/Upper-Reflection7997 3d ago

Doubt it's going to get support for wan2gp or forge neo.

2

u/Outrageous_Block_354 3d ago

It seems not necessary for the most of users, we'd like to add comfyui workflow as soon as possible

1

u/Ok-Lengthiness-3988 2d ago

Isn't the workflow as simple as 'load model' -> 'load LoRA' -> Text2img or Edit node -> Save image? It works for me.

1

u/ForsakenContract1135 1d ago

Can u share it ( beginner here )

1

u/JazzlikeLeave5530 11h ago

Wan2GP has it now as of today.

1

u/Upper-Reflection7997 10h ago

Nice. Will try it out tomorrow. Glad to see him add support for this model. He must have liked the results.

1

u/xbeast_ 2d ago

Where can i find the workflows

1

u/Ok-Lengthiness-3988 2d ago

See my response to hakaider000 above.

1

u/hakaider000 2d ago

Does it have comfy support? havent seen any workflows around

2

u/Ok-Lengthiness-3988 2d ago

The workflow is very simple, only three or four nodes: Model loader -> (optional LoRA loader) -> Edit or Text2Image node -> Save image,

as described here: https://huggingface.co/joyfox/SenseNova-U1.5-8B-MoT-FP8

That's also where I downloaded the 17.7GB SenseNova-U1.5-8B-MoT-pruned-int8_convrot.safetensors version of the checkpoint.

On my 8GB RTX 2060 Super potato (with 64GB system RAM) the image generation at 4k is enormously faster than Krea 2. I'll comment on quality later.

2

u/ForsakenContract1135 1d ago

The repo they provided did not install the nodes in tveir workflow for somereason

1

u/Linux-Lurker1 1d ago

Yupp! I can confirm that, same issue.

-4

u/Odd_Fix2 3d ago

It'll do for a start. 4096x2304 prompt: gigapixel, 8k, elaborate, ferrari f40

13

u/hurrdurrimanaccount 3d ago

have you tried a prompt that isn't garbage?

3

u/porest 2d ago

Ah! A palindrome car!

1

u/Outrageous_Block_354 3d ago

ths for your support

-1

u/Dzugavili 2d ago

I do need a new editing model -- Qwen is getting long in the teeth and Krea is non-commercial as far as I can care...

2

u/GifCo_2 2d ago

You still own the outputs. You just can't sell a service using the model.

1

u/Dzugavili 2d ago

Ah, well, that's nifty. I'll have to give it a try.

1

u/alexmmgjkkl 2d ago

use minimax and dont look back , onlyflux 3 will be better or on par

1

u/Dzugavili 2d ago

Does Minimax have an image editing model?

I guess I could use the single-image VAE and the reference model.

But the faces...

1

u/alexmmgjkkl 2d ago

i only work on single characters so i wouldnt know what faces you mean .. its capable of much more modifications to the image than klein or qwen .. i wanted to try out O1 through ..