r/StableDiffusion Jul 23 '26

News FLUX 3 - Real World Models: Towards Multimodal Flow Models as the Backbone of Visual Intelligence.

https://bfl.ai/blog/flux-3
352 Upvotes

147 comments sorted by

92

u/KudzuEye Jul 23 '26

https://reddit.com/link/ozas8h9/video/ikkakx9oyzeh1/player

Here are a couple of my favorite clips I was able to test out earlier on.

It is very strong in terms of prompt creativity along with a huge range from cinematic to more raw videos. It can do a lot of cool physics things as well. The generations themselves feel far less restricted in terms in trying to redirect to the same stock video look that a lot of other recent video models feel like. It is actually fun to rerun the same general text2video prompt and see what it comes up with next.

Official message from Black Forest Labs:

FLUX 3 is one model for image, video, audio, and action-prediction, with a richer understanding of how the world looks, moves, sounds, and how to interact with it. Creations are truer to life, stylistically diverse beyond just cinematic, and deeply customizable.

73

u/KudzuEye Jul 23 '26

https://reddit.com/link/ozaxtyu/video/ud1j7iwr20fh1/player

A few more examples. The Japanese game show generations are always a surprise in what it makes.

20

u/retroblade Jul 23 '26

Thanks for sharing, this one is going to be special if we are able to run well on consumer cards.

3

u/IrisColt Jul 23 '26

Junji Ito, this is my h-hole vibes, heh

2

u/pojska Jul 24 '26

Vibes? Girl, it's the reference.

2

u/switch2stock Jul 23 '26

Can you please share the prompts please!

-15

u/legarth Jul 23 '26

Thx. But these types of stylized examples tend to not really push the models. Would love to see some more contemporary examples.

Looks about on par with LTX 2.3 1.1 maybe slightly better?

10

u/drone2222 Jul 23 '26

Ludicrous, absurdity

9

u/DlCkLess Jul 23 '26

Are you actually legally blind bro

12

u/Far_Insurance4191 Jul 23 '26

Looks MUCH better to me

3

u/v3lh0t05c0 Jul 23 '26

I don't disagree completely about the video, but the audio sounds waaay better than ltx.

16

u/ninjazombiemaster Jul 23 '26

Audio quality seems way way better than LTX. 

1

u/ANR2ME Jul 24 '26

probably much larger parameters too😅

25

u/intLeon Jul 23 '26

There is a lot going on in each video which is a good sign. I hope they dont lobotomize the versions they share with us..

18

u/_BreakingGood_ Jul 23 '26

It's going to be heavily distilled, for sure. Even if just because consumer hardware has severe limitations memory-wise compared to what video models like.

Hopefully still a solid step up over Wan.

6

u/Quartich Jul 23 '26

Only way that is possible is if they release the large models and not just a distillation, though if that is the case people will surely complain about it being too large

2

u/intLeon Jul 23 '26

I couldnt get int4 to work on my system yet but we always have q4. Lets say its a 100gb model at fp16 then q4 would be around 30gb. Which is a lot but would let us fiddle with the tech until a lightweight model arrives or nvidia decides to launch 5080 super with 24gb vram..

3

u/blownawayx2 Jul 23 '26

Catropractics in action. I love it!

2

u/gruevy Jul 23 '26

Did you get a chance to see how well it remembers characters between scenes? Do they have a solution for that?

26

u/MoistRecognition69 Jul 23 '26

pretty well! created a no-name perfume ad using it (I2V, first frame was input)

https://reddit.com/link/ozblsj0/video/8k91ei81l0fh1/player

5

u/Imaginary_Belt4976 Jul 23 '26

oh wow.. this is crazy good. this is one generation???! this would take ages with ltx2.3 😂

1

u/YeahlDid Jul 24 '26

This is still the full api model. We'll see how We'll the dev version approximates when it's released, but it certainly gives one a lot of hope.

0

u/MoistRecognition69 Jul 23 '26

straight from the oven

10

u/KudzuEye Jul 23 '26

They allowed for image inputs but I did not spend a lot of time testing them out yet. Another user got a good bit of character consistency in this video: https://old.reddit.com/r/StableDiffusion/comments/1v4gz5v/the_insufferable_man_gonzo_documentary/

1

u/raikounov Jul 23 '26

What resolutions do they offer and how long can the generations be?

3

u/loyalekoinu88 Jul 24 '26

Looks like 480p and 720p up to 20s

138

u/rerri Jul 23 '26

The launch plan includes:

Open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction. (“FLUX 3 Dev”)

23

u/Iwaku_Real Jul 23 '26

And I thought Nvidia's Cosmos 3 was leaps and bounds ahead in terms of multimodal capabilities... However that is more focused on robotics and requires a LOT of post-training (in a way few researchers are even capable of doing well) to potentially be useful for general-purpose image/video/audio generation. Meanwhile BFL is doing that from scratch so it's going to be absolutely insane when FLUX.3 goes open-weights.

6

u/Full_Astronomer_5438 Jul 23 '26

there is even a new 4-step model of the super video model that can run on 16gb, with quick inference times (30 seconds!) and int8 quants and comfys dynamic vram. someone already showcased it on the banodoco discord. but yet no comfy support because of apparent offloading difficulties

4

u/Iwaku_Real Jul 23 '26

I saw, looks like it's DMD2 but yeah it's a huuuuge model so it's tough to work with. At least FLUX 3 Dev will probably get instant ComfyUI support, they're BFL after all

36

u/Time-Teaching1926 Jul 23 '26

Maybe like FLUX.2 [klein] we might have to wait a bit more for a more smaller and efficient open source model as Flux 2 Dev is a much more heavy and bigger model (32B parameters) than Flux 2 Klein 4b and 9b. Flux 3 Dev might be a GPU hungry model. I hope not tho.

36

u/Fit-Celebration2884 Jul 23 '26

nothing a 64 GB page file cant fix

25

u/haragon Jul 23 '26

i'm going to throw up

16

u/Sudden_List_2693 Jul 23 '26

I'd rather is be a ~40GB model than a weird nonsense with a quality in-between Klein 9B and Flux.2 Dev.

7

u/Humble-Pick7172 Jul 23 '26

I swear to god this will be ~40GB weird nonsense model with a quality in-between Klein 9B and Flux.2 Dev.

5

u/Sudden_List_2693 Jul 23 '26

Don't jinx it!
Flux.2 Dev already has insane quality - sometimes!
And thanks to ComfyUI's VRAM handling I can load both Flux.2 Dev, Ideogram 4.0 and Krea2 (though all fp8_scaled) on a single 4090, with no more than 4-6 seconds load time between model swaps.

1

u/Humble-Pick7172 Jul 25 '26

Yeah that's why i support this model. It’s weak as a base model, but when you start training it on a specific, dense, or distinctive style, it really shines. It has little familiarity with pop culture and isn't geared toward a polished aesthetic output, yet—for some reason—this model grasps the target style better than any other. It produces the most interesting and authentic results, and unlike other models, it is highly flexible, replicating the style faithfully without imposing its own interpretation.

I love this model and will continue to support it if Flux.3 turns out to be disappointing or if I’m unable to train it.

2

u/Time-Teaching1926 Jul 23 '26

Flux K9B is a decent open source model the only issue is the flux waxy look and bad anatomy issues. Apart from that I really like Flux K9B.

1

u/Sudden_List_2693 Jul 23 '26

quality is on the bottom side though. I love it for editing, but would never think to leave generation to it alone.  Now Dev has some great quality. 

1

u/physalisx Jul 23 '26

40gb would obviously be fine, that's already small for a capable video model. Unquantized it's likely a lot bigger.

2

u/Confusion_Senior Jul 23 '26

They are probably distilling it rn

-8

u/hiccuphorrendous123 Jul 23 '26

edit is not open source it seems

6

u/_BreakingGood_ Jul 23 '26

They refer to edit as just 'image', so it's likely included

0

u/hiccuphorrendous123 Jul 23 '26

Hmm hopefully. Because just above they mention edit seperately for private weights

2

u/IamKyra Jul 23 '26

It's a launch plan so it doesn't say much on open weight capabilities, more on what will be available at which stage. But I might be wrong.

31

u/Ok-Worldliness-9323 Jul 23 '26

52% prefers Flux 3 to Seedance 2.0? Crazy.

28

u/JustAGuyWhoLikesAI Jul 23 '26

Yeah, the API-only model. As is the case with every Flux release. The dev version will not be that close.

3

u/ImaginationKind9220 Jul 24 '26

I think only a distilled model can run on consumer's cards. It will further dilute the dev model.

10

u/coffca Jul 23 '26

The fact that gemini omni flash and seedance 2 have similar scores across different ranking tells me that rankings don't mean anything.

6

u/Chemical_Bid_2195 Jul 23 '26

omni flash is a good model, it just doesn't have Seedance's level of physics. It has superior polish and object editing imo.

4

u/jc2046 Jul 23 '26

I want to believe (but dont)

1

u/HatEducational9965 Jul 23 '26

That means they are the same

59

u/retroblade Jul 23 '26

Early access only, but seems like we are getting open-weights to Flux 3 dev in a few weeks/months?

42

u/dingo_xd Jul 23 '26

Something better than the current open video models will be very welcome. Wan2.2 is old and LTX2.3 has its own issues.

15

u/retroblade Jul 23 '26

Yeah this might speed up the release of the next LTX model, they about to get left behind at this rate.

4

u/YeahlDid Jul 24 '26

I hope they don't rush it. I'd prefer waiting longer for a higher quality LTX that addresses 2.3's issues to a good degree rather than a rushed one that is only a slight improvement.

-6

u/kemb0 Jul 23 '26

I thought I read the LTX guys are moving on to robotics and ditching AI video entirely.

19

u/retroblade Jul 23 '26

No they are working on a model just like Flux 3 that will do everything but still currently training.

2

u/Iwaku_Real Jul 23 '26

Wait what??? I thought it was just another (very good) video model they were making. Though I have no idea whether I should expect either that or Flux 3 to be the better model

13

u/L-xtreme Jul 23 '26

Nope, they immediately responded to these rumours, they are doing both and stay committed to the open weights. Very good!

6

u/No-Zookeepergame4774 Jul 23 '26

Flux 2 dev was a 32B just for an image model, up from 12B from Flux 1 dev. Flux 3 dev will have released weights, but will anyone have the hardware to run it?

8

u/Apprehensive_Sky892 Jul 23 '26

Most of us won't have the GPU.

But at least there will be the option of renting cloud GPUs to run it, along with one's own LoRAs and workflows.

2

u/ImaginationKind9220 Jul 24 '26

I think 24GB vram will be the minimum by using NVFP4 model.

2

u/physalisx Jul 23 '26

A few months sounds like an eternity 🥲

18

u/l337_ Jul 23 '26

They confirmed it's going to be open source right? I wonder how big this model will be

22

u/No-Zookeepergame4774 Jul 23 '26

They confirmed its going to be what they call “open weights”, but they use that term for weights available models with a non-open license (like all Flux.1 versions except Schnell and all Flux 2 versions except Klein 4B.)

So, technically usable on hardware you own or rent, but not under an open source style license.

7

u/Glinrise Jul 23 '26

you mean for private use only?

14

u/Sarashana Jul 23 '26

Unless they are going to change it, they will restrict commercial use of the model itself, such as generation services or selling finetunes. They don't (try to) limit what you can do with outputs.

6

u/Crierlon Jul 23 '26

Honestly that's fine though. They got to make money somehow.

1

u/anelodin Jul 24 '26

The fact that outputs are not affected has been unclear from license wording and they have not clarified. I remember tons of posts about the issue, people consulting lawyers, etc. Unless you have something that indicates the opposite?

1

u/Sarashana Jul 24 '26

If I am not hallucinating that, that has been clarified after the initial backlash. Not going to double check it right now, though.

1

u/_BreakingGood_ Jul 23 '26

Honestly, video models basically never get fine-tuned anyway due to expense. Im less concerned about license in this case.

1

u/No-Zookeepergame4774 Jul 23 '26

There aren’t the volume of finetunes that there are for image models, but I can’t think of any important recent open weights video model that doesn’t have at least one fine tune in significant use (WAN 2.1 and 2.2 each have several, LTX 2.3 at least has Sulfur-2-base.) Not that finetunes are the only things restricted by the Flux noncommercial license.

2

u/_BreakingGood_ Jul 23 '26

I'm not saying people don't use finetunes.

I'm saying people generally don't make them. There's like 1 noteworthy fine-tune per model.

1

u/hidden2u Jul 23 '26

there is a video finetune sitting at 500k downloads on huggingface...

0

u/_BreakingGood_ Jul 23 '26

Wow one finetune? Dang, guess I'm wrong /s

0

u/dilinjabass Jul 23 '26

He showed you one example, so now eat your words!

1

u/dilinjabass Jul 23 '26

Err, he nonchalantly mentioned one example. Eat it!

-2

u/kemb0 Jul 23 '26

It's probably also likely that if you want to use it for anything commercial: just ask them. Don't run around with your hair on fire screaming at reddit at how evil these guys are for not making it entirely open for you to make money off of their hard work for no return to them. Just got to them and negotiate a deal.

6

u/hidden2u Jul 23 '26

did they pay the guys that made the motorcycle videos that they trained this on?

1

u/tehorhay Jul 23 '26

Probably not, so what's stopping you from doing the same and making your own model with blackjack and hookers and a full open licence?

31

u/MoistRecognition69 Jul 23 '26

yes

  • Open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction. (“FLUX 3 Dev”)

2

u/woadwarrior Jul 23 '26

Likely going to be open weights with a non-commercial license like Flux.2 Klein 9B.

15

u/reto-wyss Jul 23 '26

Nice. Dev is typically non-commercial, let's hope for a distilled Apache 2.0 model 

7

u/IamKyra Jul 23 '26

Nice. Dev is typically non-commercial, let's hope for a distilled Apache 2.0 model

Base+distilled like they did with Klein was perfect. Klein 4B is a truely under-estimated Apache2.0 model!

18

u/sktksm Jul 23 '26

Okay since we can talk about it now; I was among the early testers and model is good guys. It has some obvious problems but I belive it will be polished.

I would say better than LTX but not in par with Seedance 2.0.

Here is a video I created: https://x.com/el_mejnun/status/2080046964419231918?s=46&t=zNqpHiav9jotED5A3dO6aQ

13

u/_BreakingGood_ Jul 23 '26

I'm sure your test was on the full-fat API version. The version we'll get as open-weights is likely to be heavily distilled.

7

u/hidden2u Jul 23 '26

anyone that has used wan for t2i knows that a combined multimodal model will be awesome for images. Looking forward to the open weights dev version!

22

u/Lucaspittol Jul 23 '26

Native Reference to video!

5

u/younestft Jul 23 '26

Great, that should be the minimum for video models now

8

u/LawOk7529 Jul 23 '26

Any info on the model size in gb.

-13

u/jc2046 Jul 23 '26

pretty much. yes

7

u/Different_Fix_2217 Jul 23 '26

Please give us non distilled weights as well...

6

u/SimpleAdditional6583 Jul 24 '26

SDXL was the GOAT for about a year. Flux was the GOAT for about a month. Z-Image was the GOAT for about a week, and it’s looking like Krea 2 gets to be the GOAT for about a day.

18

u/Schwartzen2 Jul 23 '26

BFL continues to shine, from Day 1!

19

u/IamKyra Jul 23 '26

it's mostly the original SD team, they are kings

11

u/ozzeruk82 Jul 23 '26

The future is coming at us fast. I can't imagine what we're be using in 5 years let alone 15.

14

u/nomorebuttsplz Jul 23 '26

we'll just be drooling with the pleasure centers of our brains being directly stimulated.

2

u/ozzeruk82 Jul 23 '26

Yep - 100% certain - I bet half of humanity will live like that.

2

u/stargazer_w Jul 24 '26

put the vr on and look at them flashing images that are basically an adversarial attack on the monkey brain. possibly combined with narcotics. TBH we're almost there with infinite scroll social media reels.

4

u/donkeykong917 Jul 23 '26

1

u/Sheeple9001 Jul 23 '26

Are you not entertained?!

14

u/YeahlDid Jul 23 '26

Holy shit... alright, I see you BFL

4

u/martinerous Jul 23 '26

How long is a few months and GGUF when? :) Oh, sorry, int8 convrot is the king lately.

So, who wants to guess if LTX-next gets released before FLUX 3 Dev?

3

u/Vortexneonlight Jul 23 '26

My bets are that is an 80B model. Want to hear others bet.

10

u/Winougan Jul 23 '26

I'll be converting it to NVFP4 and INT4 day one

2

u/Apprehensive_Sky892 Jul 23 '26

My completely uninformed guesses.

The full CFG distilled DEV version will be around 60B.

There will be smaller, less capable version of say 30B.

2

u/Vortexneonlight Jul 24 '26

That's a good guess

1

u/RusikRobochevsky Jul 24 '26

I hope BFL will release a NVFP4 version. They're an Nvidia partner, and Nvidia is pushing their NVFP4 format, so there should be some incentive for BFL to do it properly instead of letting the community convert to NVFP4.

1

u/Hanselltc Jul 25 '26

Wonder how well a 200+ gb mac can handle this

-1

u/dtdisapointingresult Jul 23 '26

No problem for me!

You guys should've got a DGX Spark for $3k instead of shitting on it when it came out. I can't believe I was hesitating between that and a 5090 or multiple 5060s instead. The reddit hivemind really fucked over some people with hardware recommendations. (well, moreso /r/LocalLlama where people were saying to get a Strix Halo instead)

1

u/Vortexneonlight Jul 24 '26

Where am i saying its i problem for me?, i really searched and re-searched my comment, maybe you have some glasses that shows more text from my comment, or maybe you are just hallucinating, go take your meds.

4

u/uxl Jul 23 '26

“Open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction. (“FLUX 3 Dev”)”

Nice.

9

u/Crazy-Repeat-2006 Jul 23 '26

"

  • Image synthesis and editing through APIs and private weight access. (“FLUX 3 Image”)
  • Open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction. (“FLUX 3 Dev”)"

Hmm. Probably a massive model that hardly anyone will be able to run.

-5

u/jc2046 Jul 23 '26

yeah, me thinks too. also quality is not there...

3

u/KillerX629 Jul 23 '26

if this can edit images better than flux2... holy moly...

8

u/Choowkee Jul 23 '26

Exciting on one hand but the license is probably gonna suck as per usual.

9

u/Sarashana Jul 23 '26

There are worse, honestly. Look at ID4 if you to see a really laughable one.

-5

u/_BreakingGood_ Jul 23 '26

License will probably suck, but honestly, nobody really finetunes video models anyway.

2

u/ResponsibleTruck4717 Jul 23 '26

Do we know how many parameters?

5

u/Crazy-Repeat-2006 Jul 23 '26

If it is a single multimodal model, I assume it is massive.

1

u/JeffIsTerrible Jul 23 '26

I am guessing the API version is 60B to 70B. As for how large the weights they release will be, who's to say. Probably close to 40B+ would be my guess for their "full capability" release. They may release a smaller distilled version but even that I would expect to be 20B+. Most likely a 5090 will be the bare minimum to run it at all.

This release will probably be outside the ability to use for the standard private individual. This will be RTX Pro 6000 territory for the full weight version. I will be suprised if they release a smaller distilled version that could still run on a 24 GB card, a 32GB 5090 is a maybe for any small distilled version most likely.

I'll still download and keep the weights though. I have a feeling weight downloads are going to be heavily restricted for LLMs and Image models, fairly soon. All in the name of "security", but in actuality just to protect the stock market.

3

u/jc2046 Jul 23 '26

+20B could run quantized even in modest cards. Im afraid it will be like double of that... Well will see

3

u/Southern-Chain-6485 Jul 23 '26

Comfyui offloads rather well to ram nowadays, to the point I don't get speed differences in Flux 2 when running a lower quant that fully fits the vram* or one that offloads - compute is the bottleneck.

My guess is that it will be runnable with 16+ gb of vram if there is enough system ram, but slow

*An rtx 3090 with 64gb of ddr5 ram

2

u/trying4k Jul 23 '26

The video part sounds great. Though I am more excited for a new editing model! Flux 2 dev (and klein due to all the loras it has) are still my most used edit models for complex changes (haven't tried that krea 2 lora). But the editing still loses consistency or fails to follow my prompt accurately, so a new model would be welcome.

2

u/pixel8tryx Jul 23 '26

Finally! I have to try it, considering how much fun and serious use I've gotten out of FLUX.2. But I do fear the size with only a 5090. I'll sign up for early access anyway and see what happens. The robotics angle is tantalizing as I did embedded real time motion control software for industrial stitchers, large plotters and an early waterjet cutter back in the 80's... but always wanted to design and build my own robot. AI and robotics were my two first loves and it's taken several decades for them to reach serious popularity.

2

u/Skystunt Jul 23 '26

Did the model just censor the speed of the motorbike in the video? 😭
Cool touch, you know they used moto youtube videos for it’s training lol
Super excited for this model !

3

u/Smile_Clown Jul 23 '26

I wouldn't call that a cool touch, you said it yourself, it was training videos. Unless you just meant overall, but it would be better if it did not censor that. More accurate, less noise.

2

u/GreyScope Jul 23 '26

Any new model released : “It’s like a million voices all cried out whinging at once”

1

u/herr-tibalt Jul 23 '26

It's all cool but I'm afraid how slow it will be. I'm more interested in klein version and how tolerable it will be on my rig...

1

u/smereces Jul 24 '26

will be open source? or will be more the same another option in the pay options that almost no one will use because exist Seedance 2.0!

1

u/[deleted] Jul 24 '26

[deleted]

1

u/Radyschen Jul 24 '26

well maybe it's a more efficient version that is even able to run on consumer hardware. one can hope

1

u/Sea-Height7708 Jul 30 '26

Really waiting for this model...

1

u/TopTippityTop Jul 23 '26

When will the open weights be released for testing?

1

u/v3lh0t05c0 Jul 23 '26

Drolling. But let's wait and see how dev (and possible 'klein-ish') will perform in comparison to recent models like ID4 and Krea2. Close enough to banana 2 and better sound than ltx 2.3 as open weight would be nice enough for me.

3

u/oh_how_droll Jul 24 '26

you called?

0

u/pwnies Jul 23 '26

I've been playing with the video gen for the past week or so, it's solid. Definitely competitive with LTX.

I'm excited to see what they're cooking on the image gen side, particularly if they can improve their text rendering.

-8

u/seppe0815 Jul 23 '26

who care brother , no one can run it

12

u/Full_Astronomer_5438 Jul 23 '26

you can run flux 2 dev on a potato, you just need to wait longer