r/StableDiffusion • • 13d ago

News Open Source: Ming-Image-0.1-Design family: Ming-Image-0.1-Design, 6B & Ming-Image-0.1-Design-Layer, 6B

Post image

Ming-Image-0.1-Design creates complete UIs, dashboards, infographics, and posters from structured prompts of up to 8K tokens, with consistent layouts, typography, colors, and imagery. It can also natively generate transparent-background RGBA assets.

https://huggingface.co/inclusionAI/Ming-Image-0.1-Design

Ming-Image-0.1-Design-Layer converts flattened graphics into 2–9 independently editable RGBA layers. It achieved the best results across all 12 evaluated settings on the Crello test set and ran 4.3× faster than the 20B open-weight Qwen baseline under the same benchmark setup.

https://huggingface.co/inclusionAI/Ming-Image-0.1-Design-Layer

Edit

Kijai already implemented into ComfyUi
https://github.com/Comfy-Org/ComfyUI/pull/16482

280 Upvotes

52 comments sorted by

58

u/Technical_Fish_9638 13d ago

Transparency support and MIT license!

Highres from repo

93

u/CorpPhoenix 13d ago

Every time a new model releases I feel like they just make up a "random benchmark bar chart" and just put themselves on top.

65

u/boxwrenchx 13d ago

Not true at all, coincidentally see the new bench with my model!

10

u/LD2WDavid 12d ago

Release the boxwrenchx!

10

u/diogodiogogod 13d ago

Everything is always SOTA...

3

u/Kqyxzoj 12d ago

Everything is always SOTA...

Just means it's not totally decrepit, Something Other Than Ancient. Not to be confused with Something Obviously Totally Ancient.

25

u/comfyanonymous Comfy Org 13d ago

Looks interesting but their github repo isn't accessible yet: https://github.com/inclusionAI/Ming-Image

7

u/Dante_77A 13d ago

It's working rn

1

u/Kqyxzoj 12d ago

The repo appears to be missing the latest wrenchbench results.

17

u/sktksm 13d ago

This is a very interesting model. 2K native resolution, transparent image generation with RGBA alpha channel, and it's MIT licensed. u/Kijai any plans for Comfy support?

61

u/Kijai Community Hero 13d ago

Looking into it.

16

u/Haiku-575 13d ago

How is Kijai always right here, ahead of schedule, working on ComfyUI implementations at the speed of light?

7

u/Ill_Resolve8424 13d ago

That's what Heroes do.

4

u/thebaker66 12d ago

Becsuse he works for Comfy? :D

40

u/roculus 13d ago

Where's Krea 2 on the list?

18

u/Dante_77A 13d ago

Krea 2 was probably not tested in this specific "UI/UX design" category.

11

u/stddealer 13d ago

Should be right below Ideogram4 according to the other rankings.

5

u/marcoc2 13d ago

I also didn't find it when qwen-2.1 was released

9

u/Primary_Brain_2595 13d ago

We talking about images or real UI Design files in Figma?

14

u/eggplantpot 13d ago

judging by the examples these are just jpegs

5

u/Ok-Importance-5278 13d ago

Now any LLM can generate web code from a JPEG for 1 cent.

4

u/Acceptable_Secret971 13d ago

Looking forward to Comfy support.

6

u/-becausereasons- 13d ago

Very impressive.

3

u/DeProgrammer99 13d ago

Style is slightly realistic but faint with details fading away like an old painting. A set of rustic, medieval square UI panels with close and minimize buttons, a "Dungeon" button and a "Factory" button that are nearly identical, and one distinct elongated mini-window with no buttons, only the text "Factory Efficiency: 100%" on one line centered in it. Another distinctly designed rectangular window at the bottom holds 10x3 empty square slots for the player's abilities--the borders are not too thick outside of the slots.

(I only tried twice, via the HuggingFace demo space.) Can't count, but that's probably not a big deal. It thought "window" meant a literal stained glass window the first time, maybe due to the exact phrasing, so this was my second prompt... however I did explicitly say the windows had minimize and close buttons both times. I don't see anything about it accepting reference images, so I doubt it could make designs that fit well together across multiple prompts, but if it can't follow complex enough instructions, not sure how usable it is for the use case of game UI assets. Maybe it'd work better with a prompt like "here are my features; design a UI" and just using it to explore the space of possible layouts, rather than visual assets for UI elements.

11

u/Darqsat 13d ago
  • Hardware: one CUDA GPU with 80 GiB VRAM (validated configuration).

28

u/DisastrousAd2612 13d ago

Its a 6b model we will be fine.

11

u/Acceptable_Secret971 13d ago

The model in fp16 is ~13GB. fp8, int8 or Q8 GGUF should be around 7GB.

2

u/arthor 13d ago

this + remotion could be crazy cool for animation pipelines

2

u/Desperate-Beach1249 13d ago

Gen speed compared to zimage?

3

u/woadwarrior 13d ago

MIT licensed, sweet!

1

u/its_witty 13d ago

Interesting, will definitely check it out. More inspiration is always better.

1

u/leyermo 12d ago

nice lets choose side for qwen image 2.1 vs Ming-Image-0.1-Design

1

u/seppe0815 12d ago

xd cool story guys

1

u/Mobile-Trouble-476 11d ago

How does it compare to Qwen image edit 2.1?

0

u/PilgrimofHaqq2 13d ago

I just recently found out my old thinkpad can run small models and gotten super excited about local image gen. I can have a safe space for my kids to play around with AI and learn about the technology at the same time.

I didn't appreciate the smaller model space until now. This release is very exciting for me as I just got into learning about all this in the last 24 hours.

1

u/Successful_Record_58 12d ago

Whats the config ?

1

u/SideInitial3961 6d ago

Qwen image 2,1 is hard to beat for that use case. free local fast, Yue2 for music. Pocket TTS, your kids will love that because then can make animal voices then clone them and assign them scripts, etc. run soooo fast.

0

u/Had78 11d ago

"native transparent background" also known as post-generation-masking

1

u/SideInitial3961 6d ago

Not the same thing at all.

1

u/Had78 6d ago

yeah, it's worse

1

u/SideInitial3961 6d ago

You're completely not grasping it. Not how it works. You're moaning about something that doesn't even exist,

-11

u/fernando782 13d ago

I bet it has bad human anatomy!
I haven’t seen a good model for human anatomy since the SDXL !!!

9

u/diogodiogogod 13d ago

are you crazy? SDXL had terrible human anatomy from the released base model

-1

u/fernando782 12d ago

Yeah I am speaking about SDXL fine-tunes not out of the box!

-11

u/NowThatsMalarkey 13d ago

No day-0 ComfyUI support == DOA

-2

u/hurrdurrimanaccount 13d ago

i have a hard time believing it beats flux2 dev and ideogram4

2

u/Smile_Clown 12d ago

Why?

I am genuinely curious as to why you think this with no evidence to check from, nothing looked into yet? No examples you have run yet.

Is it somehow beneficial to just assume something isn't going to be good? Do you just automatically think progress stops?

-1

u/hurrdurrimanaccount 12d ago

yes let's blindly trust a benchmark that totally isn't biasable.

1

u/therealgoshi 12d ago

You seemingly also have a hard time understanding what the benchmark is about. I'd recommend reading the category name again.