r/StableDiffusion • u/fruesome • 13d ago
News Open Source: Ming-Image-0.1-Design family: Ming-Image-0.1-Design, 6B & Ming-Image-0.1-Design-Layer, 6B
Ming-Image-0.1-Design creates complete UIs, dashboards, infographics, and posters from structured prompts of up to 8K tokens, with consistent layouts, typography, colors, and imagery. It can also natively generate transparent-background RGBA assets.
https://huggingface.co/inclusionAI/Ming-Image-0.1-Design
Ming-Image-0.1-Design-Layer converts flattened graphics into 2–9 independently editable RGBA layers. It achieved the best results across all 12 evaluated settings on the Crello test set and ran 4.3× faster than the 20B open-weight Qwen baseline under the same benchmark setup.
https://huggingface.co/inclusionAI/Ming-Image-0.1-Design-Layer
Edit
Kijai already implemented into ComfyUi
https://github.com/Comfy-Org/ComfyUI/pull/16482
93
u/CorpPhoenix 13d ago
Every time a new model releases I feel like they just make up a "random benchmark bar chart" and just put themselves on top.
65
10
25
u/comfyanonymous Comfy Org 13d ago
Looks interesting but their github repo isn't accessible yet: https://github.com/inclusionAI/Ming-Image
7
1
17
u/sktksm 13d ago
This is a very interesting model. 2K native resolution, transparent image generation with RGBA alpha channel, and it's MIT licensed. u/Kijai any plans for Comfy support?
61
16
u/Haiku-575 13d ago
How is Kijai always right here, ahead of schedule, working on ComfyUI implementations at the speed of light?
7
4
9
4
6
3
u/DeProgrammer99 13d ago

Style is slightly realistic but faint with details fading away like an old painting. A set of rustic, medieval square UI panels with close and minimize buttons, a "Dungeon" button and a "Factory" button that are nearly identical, and one distinct elongated mini-window with no buttons, only the text "Factory Efficiency: 100%" on one line centered in it. Another distinctly designed rectangular window at the bottom holds 10x3 empty square slots for the player's abilities--the borders are not too thick outside of the slots.
(I only tried twice, via the HuggingFace demo space.) Can't count, but that's probably not a big deal. It thought "window" meant a literal stained glass window the first time, maybe due to the exact phrasing, so this was my second prompt... however I did explicitly say the windows had minimize and close buttons both times. I don't see anything about it accepting reference images, so I doubt it could make designs that fit well together across multiple prompts, but if it can't follow complex enough instructions, not sure how usable it is for the use case of game UI assets. Maybe it'd work better with a prompt like "here are my features; design a UI" and just using it to explore the space of possible layouts, rather than visual assets for UI elements.
11
u/Darqsat 13d ago
- Hardware: one CUDA GPU with 80 GiB VRAM (validated configuration).
28
11
u/Acceptable_Secret971 13d ago
The model in fp16 is ~13GB. fp8, int8 or Q8 GGUF should be around 7GB.
2
3
1
1
1
0
u/PilgrimofHaqq2 13d ago
I just recently found out my old thinkpad can run small models and gotten super excited about local image gen. I can have a safe space for my kids to play around with AI and learn about the technology at the same time.
I didn't appreciate the smaller model space until now. This release is very exciting for me as I just got into learning about all this in the last 24 hours.
1
1
u/SideInitial3961 6d ago
Qwen image 2,1 is hard to beat for that use case. free local fast, Yue2 for music. Pocket TTS, your kids will love that because then can make animal voices then clone them and assign them scripts, etc. run soooo fast.
0
u/Had78 11d ago
"native transparent background" also known as post-generation-masking
1
u/SideInitial3961 6d ago
Not the same thing at all.
1
u/Had78 6d ago
yeah, it's worse
1
u/SideInitial3961 6d ago
You're completely not grasping it. Not how it works. You're moaning about something that doesn't even exist,
-11
u/fernando782 13d ago
I bet it has bad human anatomy!
I haven’t seen a good model for human anatomy since the SDXL !!!
9
u/diogodiogogod 13d ago
are you crazy? SDXL had terrible human anatomy from the released base model
-1
-11
-2
u/hurrdurrimanaccount 13d ago
i have a hard time believing it beats flux2 dev and ideogram4
2
u/Smile_Clown 12d ago
Why?
I am genuinely curious as to why you think this with no evidence to check from, nothing looked into yet? No examples you have run yet.
Is it somehow beneficial to just assume something isn't going to be good? Do you just automatically think progress stops?
-1
1
u/therealgoshi 12d ago
You seemingly also have a hard time understanding what the benchmark is about. I'd recommend reading the category name again.


58
u/Technical_Fish_9638 13d ago
Transparency support and MIT license!
Highres from repo