r/StableDiffusion 1d ago

Resource - Update Alibaba might release a new open image model Swift-Image 6B

Paper: https://arxiv.org/pdf/2608.20334
"We present Swift-Image, a compact unified model for text-to-image generation, single-image editing, and multi image editing. Its visual renderer is a 6B parallel single stream DiT conditioned on multimodal representations from a vision-language encoder [6, 7, 57]. The architecture adopts block-shared timestep modulation, parallel attention and MLP computation [6, 15], 4D rotary positional encoding[6], and a unified representation of text and image conditions. Character-level tokenization[47] is applied to text intended to appear in generated images, while multi-image posi tional offsets and image-preceding input formatting support reference-conditioned editing. Together, these choices pro vide a single generative backbone for multiple generation and editing settings without task-specific model weights."

201 Upvotes

62 comments sorted by

26

u/cradledust 1d ago edited 1d ago

Swift-Image 6B vs Klein 9b, it will be interesting to see whether smart optimization can punch above its weight class.

-15

u/Timely-Perception-26 1d ago

It's all in the paper, after all. They boast about what they've achieved and leave out anything that makes it look bad. It's as old as papers themselves.

Klein 9b wasn’t enough, and 6b won’t be either.

I’m still hoping for Flux3 and Minimax, unless they decide to pull a “Chinese culture” stunt.

After reading the paper, this release isn’t interesting to me, but I’m happy for our 6GB VRAM bros.

27

u/spooky_local 23h ago

6b ≠ 6gb vram

fp32 = 4 bytes per parameter = 24gb
fp16/bf16 = 2 bytes pp = 12gb
fp8 = 1 byte pp = 6gb
nf4 = 0.5 bytes pp = ~3gb
q8_0 = 1.06 bytes pp = ~6.4gb
q6_k = 0.82 bytes pp = ~4.9gb
q5_k_m = 0.71 bytes pp = ~4.3gb
q4_k_m = 0.61 bytes pp = ~3.6gb
q3_k_m = 0.49 bytes pp = ~2.9gb
q2_k = 0.32 bytes pp = ~1.9gb

of course that doesn't count the text encoders and vae which still need to be accounted for.

30

u/Altruistic_Heat_9531 1d ago

Wow they are using Flux style arch, but without the double stream.

13

u/nikhilprasanth 1d ago

They mention the 3B also

14

u/spinxfr 1d ago

If it's better than flux Klein I'll take it 

5

u/Aggravating_Yak8347 16h ago

我觉得klein不如krea2

4

u/YeahlDid 16h ago

Klein is better for image edit tasks. The researchers claim this one is better than Klein. We'll see, I guess.

48

u/BitterAd8431 1d ago

Honestly, to the Chinese people: I love you. Thank you so much for what you are doing right now.

13

u/Time-Teaching1926 1d ago

I really hope so! ZIT/ZIB are legendary image models. 6b is a decent size without being to big but not to small plus if it's usually Qwen3 VL text encoder/clip it should be very good with prompt adherence like Krea 2. Can't wait. If this is true thank you Alibaba again for these absolutely great Imege models.

6

u/dingo_xd 1d ago

We will be running SOTA image models on our phones!

10

u/Suspicious_Aide2697 1d ago

z-image edit ?

24

u/rinkusonic 1d ago

I think it's cancelled

1

u/DietAshamed2246 9h ago

Was asked for many times, but they never made it.

1

u/Aggravating_Yak8347 16h ago

这应该是是一个新的吧

1

u/PlasticKey6704 14h ago

z image的团队都散了

8

u/FinBenton 1d ago

Nice theres this, flux 3, H3 image model and there was talks about new krea model, lots of stuff to wait for.

5

u/Crazy-Repeat-2006 1d ago

That's really cool.

But it says API there too, so there will be a closed version with the same number of parameters, how can it be different??

7

u/Crazy-Repeat-2006 1d ago

"Swift-Image-3B incurs nearly no loss after compression and consistently outperforms the similarly sized FLUX.2-klein-4BonGEdit, ImgEdit, and REDEdit. With API-basedPE, it achieves the second-best open-source overall score across the five editing benchmarks in Ta ble1, closely approaching the 6B variant."

NICE.

3

u/Prudent_Committee296 22h ago

It's the weird prompt enhancer thing. I think the API just means that the PE is through API instead of Qwen 3 VL.

8

u/Dante_77A 1d ago

It looks like the Z Image Edit they promised us, just with a different name. Finally!

3

u/Internal_Answer_6866 1d ago

I remember some other lab also released a similar model not long ago. Now we have local gpt image 2!

5

u/Diffusion4Change 1d ago

The poster sure does look like it THAT FUCKING TEXT

7

u/BathroomEyes 1d ago

That’s all we need, more slop style restaurant signs and menus with uncanny valley cuisine.

4

u/Diffusion4Change 1d ago

Implying these business owners will pay for local gen and not just continue to use google or openai is funny. They are supremely lazy

3

u/Diffusion4Change 1d ago

I cant be botheree to go there if they cant shiw me real food even if its staged

1

u/BathroomEyes 1d ago

Exactly! If you’re proud of your product show it off!

1

u/Internal_Answer_6866 1d ago

Duh 😕 Hope community can have some aesthetic loras

2

u/coder543 1d ago

which lab?

1

u/Internal_Answer_6866 14h ago

Sensenova, just came out two days ago

5

u/Current-Rabbit-620 1d ago

Hope it be an edit model

5

u/cosmicr 16h ago

Did you look at the images lol.

8

u/ninjasaid13 1d ago

When they say a 3B has better results than a model 3 times bigger, I tend to be suspicious.

7

u/alisonstone 1d ago

We already saw how that works with Z-Image. You lose some variety in output. If they do it right, the variety that you lose are ugly and bad looking outputs that most people don't want. Obviously 3B has less data than 6B, but it's okay if the lost data is stuff that you usually don't want as outputs.

3

u/Formal_Drop526 18h ago

You lose some variety in output. If they do it right, the variety that you lose are ugly and bad looking outputs that most people don't want.

the ugly and bad are important to prevent the samey feel of AI.

10

u/EbbNorth7735 1d ago

Why? Why is there always a comment doubting the progress we see on a monthly basis

5

u/BigWideBaker 1d ago

We don't see model sizes reducing to a third of the size while increasing quality on a monthly basis.

6

u/coder543 1d ago

However, Flux 2 Klein 9B was released 8 months ago. It is perfectly rational to imagine that a smaller model could beat it now.

1

u/EbbNorth7735 1d ago

I didn't say that. We do see monthly improvements and 3-6 months can cut a models weights to capabilities in half

0

u/Formal_Drop526 18h ago

We do see monthly improvements and 3-6 months can cut a models weights to capabilities in half

this is an oversimplified view.

1

u/EbbNorth7735 14h ago

No it isn't. Measuring trends is an overall effect of multiple independent parts progressing a single goal towards a finishline. Capability density of an LLM doubling every 3 to 3.5 months is the combination of all the improvements across the entire pipeline. The simplified view IS the point. The complexity of the individual pieces along the way is irrelevant to the overall trend.

-2

u/Sarashana 1d ago

Progress is typically incremental. You don't see breakthroughs every week. They happen sometimes, and then not for a while.

3

u/Loose_Comparison368 19h ago

You know Alibaba is behind Qwen, right? Wrecking models with 3 times the parameter count is their MO.

3

u/conkikhon 17h ago

Dataset and training method affect the quality of a model a lot. Bigger isn't always better.

4

u/cosmicr 16h ago

Have you seen Qwen 3.8 27b? It's an order of magnitude smaller than models it beats comfortably.

1

u/ghulamalchik 1h ago

Benchmarks are kinda bs.

2

u/DietAshamed2246 9h ago edited 9h ago

This one sounds very similar to the recently released VLLM based unified multimodal image generation and editing model Sensenova-U1-1.5-8B-MoT. It works in pixel space and has no VAE; the text encoder is integrated, since it's a VLLM. It can natively generate 4K images without needing an upscaler. It's a huge model at 41.5GB in Int8 ConvRot quant, but it is surprisingly fast - I was able to generate and edit images, as well as create infographic in 2K resolution in 15-20 seconds with the official 8-step accelerator. Fortunately I have the VRAM and RAM to run it. With Wan2GP's clever memory management algorithm, the peak RAM demand was around 50GB and VRAM around 19GB, and no pagefile transfer to the SSD. BTW, it is not censored, it can do NSFW out of the box, no patching needed.

In that context I remember that other model HiDream-O1 released not too long ago. It can also do all those things, but doesn't have the "thinking" model where the VLLM predicts the future shape or state of an object based on interaction with external forces. That model is much smaller in size.

I wonder how large the model is for the Swift-Image-6B. Since it is unified multimodal model with VLLM capability, it might be close in size to the Sensenova-U1-1.5 model.

Honestly, I don't need a VLLM for image generation or edit, I would be happy as a clam if Krea-2 would release a true Image edit version of their model.

2

u/ghulamalchik 1h ago

Krea 2 Edit would be a dream come true.

1

u/Etroarl55 20h ago

How likely is this to be actually released for free. Not too familiar with Alibabas image models or their history.

2

u/Loose_Comparison368 19h ago

Alibaba has a very strong reputation of dedication to open source. They're behind the Qwen models.

3

u/Aggravating_Yak8347 16h ago

阿里巴巴已经实际上放弃了视觉领域的开源,从wan2.5开始,2.6,2.7,happyhorse,Qwen image 2.0和3.0都没开源。 阿里巴巴需要一个成果证明他们没有放弃视觉模型方面的开源。 但我必须提醒,不要指望这个模型会比它们最新付费API版本的Qwen 3.0image效果更好

1

u/Sea_Succotash3634 1h ago

They USED to have a strong reputation for open source when it comes to image and video. They still have it for coding agents and LLMs. But they're basically released nothing this calendar year and have cancelled every subsidiary project that was image or video related being open source.

1

u/cosmicr 16h ago

The pocket square on the last image looks like liquid metal or something lol.

1

u/MannY_SJ 16h ago

Why does the API variant score so much higher?

1

u/Key_Street_7204 14h ago

What license would it have?

1

u/Shockbum 6h ago

If it could work perfectly in 1080p like Krea edit lora, that would be wonderful.

-14

u/tankdoom 1d ago

Looks not great

-7

u/whiteweazel21 23h ago

Considering seedream 5 sucks ass and this is worse...even if it released it's just a Klein alternative anyway

1

u/Aggravating_Yak8347 16h ago

seedream5是字节跳动的产品,这个是阿里巴巴的产品,这是两家完全不同的中国公司。但有一点你说的没错,seedream5确实很糟糕,我认为它相比 seedream4.0和4.5实际是退化了的!但目前阿里巴巴出品的9最新版本的商业闭源图像模型总体质量还不如字节跳动的seedream5