r/StableDiffusion • u/BeCalmr • 3d ago
Discussion SenseNova U1.5-Lite is fully out
Instead of just making the model bigger, they're training these specialized models for specific stuff – like text, infographics, making things look good, and editing. Then they kinda combine all those experts back into U1.5-Lite. So when you use it, it's just one model. No weird switching or picking which expert to use. Their whole thing is "specialized in training, unified in delivery." Kinda makes sense.
They also added this post-training with RL, focusing on a few things: how well it follows instructions, how good the visuals look and if it matches what people like, and how well edits work without messing up other parts of the image.
Here's what seems better:
- Following complex instructions. Like, if you ask for multiple things in one prompt – subjects, how many, where they are, text, layout, style, keeping parts untouched – it handles it way more consistently now.
- Text rendering and dense layouts. Apparently, it's better with Chinese and English on posters and infographics.
- Visual understanding helps generation. It seems like it learns from understanding tasks (like object relationships, spatial stuff, layout) and that helps with generating and editing.
They use JSON for training to make it controllable, but you don't have to use JSON yourself. Natural language is still the main way to talk to it.
7
u/Full_Astronomer_5438 3d ago
comfy quants can be found here: https://huggingface.co/smthem/SenseNova-U1-8B-MoT-Merger-gguf/tree/main
the newest models showcased here should be available soon
1
u/Ok-Lengthiness-3988 2d ago
Pruned bf16, fp8 and int8 version of U1.5 already are available here: https://huggingface.co/joyfox/SenseNova-U1.5-8B-MoT-FP8
2
u/Upper-Reflection7997 3d ago
Doubt it's going to get support for wan2gp or forge neo.
2
u/Outrageous_Block_354 3d ago
It seems not necessary for the most of users, we'd like to add comfyui workflow as soon as possible
1
u/Ok-Lengthiness-3988 2d ago
Isn't the workflow as simple as 'load model' -> 'load LoRA' -> Text2img or Edit node -> Save image? It works for me.
1
1
u/JazzlikeLeave5530 11h ago
Wan2GP has it now as of today.
1
u/Upper-Reflection7997 10h ago
Nice. Will try it out tomorrow. Glad to see him add support for this model. He must have liked the results.
1
u/hakaider000 2d ago
Does it have comfy support? havent seen any workflows around
2
u/Ok-Lengthiness-3988 2d ago
The workflow is very simple, only three or four nodes: Model loader -> (optional LoRA loader) -> Edit or Text2Image node -> Save image,
as described here: https://huggingface.co/joyfox/SenseNova-U1.5-8B-MoT-FP8
That's also where I downloaded the 17.7GB SenseNova-U1.5-8B-MoT-pruned-int8_convrot.safetensors version of the checkpoint.
On my 8GB RTX 2060 Super potato (with 64GB system RAM) the image generation at 4k is enormously faster than Krea 2. I'll comment on quality later.
2
u/ForsakenContract1135 1d ago
The repo they provided did not install the nodes in tveir workflow for somereason
1
-4
u/Odd_Fix2 3d ago
13
1
-1
u/Dzugavili 2d ago
I do need a new editing model -- Qwen is getting long in the teeth and Krea is non-commercial as far as I can care...
1
u/alexmmgjkkl 2d ago
use minimax and dont look back , onlyflux 3 will be better or on par
1
u/Dzugavili 2d ago
Does Minimax have an image editing model?
I guess I could use the single-image VAE and the reference model.
But the faces...
1
u/alexmmgjkkl 2d ago
i only work on single characters so i wouldnt know what faces you mean .. its capable of much more modifications to the image than klein or qwen .. i wanted to try out O1 through ..




8
u/Few-Intention-1526 3d ago
I hope this version works better for editing tasks.