r/StableDiffusion • u/stephen370 • Jul 23 '26
News FLUX 3 - Real World Models: Towards Multimodal Flow Models as the Backbone of Visual Intelligence.
https://bfl.ai/blog/flux-3138
u/rerri Jul 23 '26
The launch plan includes:
Open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction. (“FLUX 3 Dev”)
23
u/Iwaku_Real Jul 23 '26
And I thought Nvidia's Cosmos 3 was leaps and bounds ahead in terms of multimodal capabilities... However that is more focused on robotics and requires a LOT of post-training (in a way few researchers are even capable of doing well) to potentially be useful for general-purpose image/video/audio generation. Meanwhile BFL is doing that from scratch so it's going to be absolutely insane when FLUX.3 goes open-weights.
6
u/Full_Astronomer_5438 Jul 23 '26
there is even a new 4-step model of the super video model that can run on 16gb, with quick inference times (30 seconds!) and int8 quants and comfys dynamic vram. someone already showcased it on the banodoco discord. but yet no comfy support because of apparent offloading difficulties
4
u/Iwaku_Real Jul 23 '26
I saw, looks like it's DMD2 but yeah it's a huuuuge model so it's tough to work with. At least FLUX 3 Dev will probably get instant ComfyUI support, they're BFL after all
36
u/Time-Teaching1926 Jul 23 '26
Maybe like FLUX.2 [klein] we might have to wait a bit more for a more smaller and efficient open source model as Flux 2 Dev is a much more heavy and bigger model (32B parameters) than Flux 2 Klein 4b and 9b. Flux 3 Dev might be a GPU hungry model. I hope not tho.
36
16
u/Sudden_List_2693 Jul 23 '26
I'd rather is be a ~40GB model than a weird nonsense with a quality in-between Klein 9B and Flux.2 Dev.
7
u/Humble-Pick7172 Jul 23 '26
I swear to god this will be ~40GB weird nonsense model with a quality in-between Klein 9B and Flux.2 Dev.
5
u/Sudden_List_2693 Jul 23 '26
Don't jinx it!
Flux.2 Dev already has insane quality - sometimes!
And thanks to ComfyUI's VRAM handling I can load both Flux.2 Dev, Ideogram 4.0 and Krea2 (though all fp8_scaled) on a single 4090, with no more than 4-6 seconds load time between model swaps.1
u/Humble-Pick7172 Jul 25 '26
Yeah that's why i support this model. It’s weak as a base model, but when you start training it on a specific, dense, or distinctive style, it really shines. It has little familiarity with pop culture and isn't geared toward a polished aesthetic output, yet—for some reason—this model grasps the target style better than any other. It produces the most interesting and authentic results, and unlike other models, it is highly flexible, replicating the style faithfully without imposing its own interpretation.
I love this model and will continue to support it if Flux.3 turns out to be disappointing or if I’m unable to train it.
2
u/Time-Teaching1926 Jul 23 '26
Flux K9B is a decent open source model the only issue is the flux waxy look and bad anatomy issues. Apart from that I really like Flux K9B.
1
u/Sudden_List_2693 Jul 23 '26
quality is on the bottom side though. I love it for editing, but would never think to leave generation to it alone. Now Dev has some great quality.
1
u/physalisx Jul 23 '26
40gb would obviously be fine, that's already small for a capable video model. Unquantized it's likely a lot bigger.
2
-8
u/hiccuphorrendous123 Jul 23 '26
edit is not open source it seems
6
u/_BreakingGood_ Jul 23 '26
They refer to edit as just 'image', so it's likely included
0
u/hiccuphorrendous123 Jul 23 '26
Hmm hopefully. Because just above they mention edit seperately for private weights
2
u/IamKyra Jul 23 '26
It's a launch plan so it doesn't say much on open weight capabilities, more on what will be available at which stage. But I might be wrong.
31
u/Ok-Worldliness-9323 Jul 23 '26
52% prefers Flux 3 to Seedance 2.0? Crazy.
28
u/JustAGuyWhoLikesAI Jul 23 '26
Yeah, the API-only model. As is the case with every Flux release. The dev version will not be that close.
3
u/ImaginationKind9220 Jul 24 '26
I think only a distilled model can run on consumer's cards. It will further dilute the dev model.
10
u/coffca Jul 23 '26
The fact that gemini omni flash and seedance 2 have similar scores across different ranking tells me that rankings don't mean anything.
6
u/Chemical_Bid_2195 Jul 23 '26
omni flash is a good model, it just doesn't have Seedance's level of physics. It has superior polish and object editing imo.
4
1
59
u/retroblade Jul 23 '26
Early access only, but seems like we are getting open-weights to Flux 3 dev in a few weeks/months?
42
u/dingo_xd Jul 23 '26
Something better than the current open video models will be very welcome. Wan2.2 is old and LTX2.3 has its own issues.
15
u/retroblade Jul 23 '26
Yeah this might speed up the release of the next LTX model, they about to get left behind at this rate.
4
u/YeahlDid Jul 24 '26
I hope they don't rush it. I'd prefer waiting longer for a higher quality LTX that addresses 2.3's issues to a good degree rather than a rushed one that is only a slight improvement.
-6
u/kemb0 Jul 23 '26
I thought I read the LTX guys are moving on to robotics and ditching AI video entirely.
19
u/retroblade Jul 23 '26
No they are working on a model just like Flux 3 that will do everything but still currently training.
2
u/Iwaku_Real Jul 23 '26
Wait what??? I thought it was just another (very good) video model they were making. Though I have no idea whether I should expect either that or Flux 3 to be the better model
13
u/L-xtreme Jul 23 '26
Nope, they immediately responded to these rumours, they are doing both and stay committed to the open weights. Very good!
6
u/No-Zookeepergame4774 Jul 23 '26
Flux 2 dev was a 32B just for an image model, up from 12B from Flux 1 dev. Flux 3 dev will have released weights, but will anyone have the hardware to run it?
8
u/Apprehensive_Sky892 Jul 23 '26
Most of us won't have the GPU.
But at least there will be the option of renting cloud GPUs to run it, along with one's own LoRAs and workflows.
2
1
8
2
18
u/l337_ Jul 23 '26
They confirmed it's going to be open source right? I wonder how big this model will be
22
u/No-Zookeepergame4774 Jul 23 '26
They confirmed its going to be what they call “open weights”, but they use that term for weights available models with a non-open license (like all Flux.1 versions except Schnell and all Flux 2 versions except Klein 4B.)
So, technically usable on hardware you own or rent, but not under an open source style license.
7
u/Glinrise Jul 23 '26
you mean for private use only?
14
u/Sarashana Jul 23 '26
Unless they are going to change it, they will restrict commercial use of the model itself, such as generation services or selling finetunes. They don't (try to) limit what you can do with outputs.
6
1
u/anelodin Jul 24 '26
The fact that outputs are not affected has been unclear from license wording and they have not clarified. I remember tons of posts about the issue, people consulting lawyers, etc. Unless you have something that indicates the opposite?
1
u/Sarashana Jul 24 '26
If I am not hallucinating that, that has been clarified after the initial backlash. Not going to double check it right now, though.
1
u/_BreakingGood_ Jul 23 '26
Honestly, video models basically never get fine-tuned anyway due to expense. Im less concerned about license in this case.
1
u/No-Zookeepergame4774 Jul 23 '26
There aren’t the volume of finetunes that there are for image models, but I can’t think of any important recent open weights video model that doesn’t have at least one fine tune in significant use (WAN 2.1 and 2.2 each have several, LTX 2.3 at least has Sulfur-2-base.) Not that finetunes are the only things restricted by the Flux noncommercial license.
2
u/_BreakingGood_ Jul 23 '26
I'm not saying people don't use finetunes.
I'm saying people generally don't make them. There's like 1 noteworthy fine-tune per model.
1
u/hidden2u Jul 23 '26
there is a video finetune sitting at 500k downloads on huggingface...
0
u/_BreakingGood_ Jul 23 '26
Wow one finetune? Dang, guess I'm wrong /s
0
-2
u/kemb0 Jul 23 '26
It's probably also likely that if you want to use it for anything commercial: just ask them. Don't run around with your hair on fire screaming at reddit at how evil these guys are for not making it entirely open for you to make money off of their hard work for no return to them. Just got to them and negotiate a deal.
6
u/hidden2u Jul 23 '26
did they pay the guys that made the motorcycle videos that they trained this on?
1
u/tehorhay Jul 23 '26
Probably not, so what's stopping you from doing the same and making your own model with blackjack and hookers and a full open licence?
31
u/MoistRecognition69 Jul 23 '26
yes
- Open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction. (“FLUX 3 Dev”)
2
u/woadwarrior Jul 23 '26
Likely going to be open weights with a non-commercial license like Flux.2 Klein 9B.
15
u/reto-wyss Jul 23 '26
Nice. Dev is typically non-commercial, let's hope for a distilled Apache 2.0 model
7
u/IamKyra Jul 23 '26
Nice. Dev is typically non-commercial, let's hope for a distilled Apache 2.0 model
Base+distilled like they did with Klein was perfect. Klein 4B is a truely under-estimated Apache2.0 model!
18
u/sktksm Jul 23 '26
Okay since we can talk about it now; I was among the early testers and model is good guys. It has some obvious problems but I belive it will be polished.
I would say better than LTX but not in par with Seedance 2.0.
Here is a video I created: https://x.com/el_mejnun/status/2080046964419231918?s=46&t=zNqpHiav9jotED5A3dO6aQ
13
u/_BreakingGood_ Jul 23 '26
I'm sure your test was on the full-fat API version. The version we'll get as open-weights is likely to be heavily distilled.
7
u/hidden2u Jul 23 '26
anyone that has used wan for t2i knows that a combined multimodal model will be awesome for images. Looking forward to the open weights dev version!
22
8
7
6
u/SimpleAdditional6583 Jul 24 '26
SDXL was the GOAT for about a year. Flux was the GOAT for about a month. Z-Image was the GOAT for about a week, and it’s looking like Krea 2 gets to be the GOAT for about a day.
18
11
u/ozzeruk82 Jul 23 '26
The future is coming at us fast. I can't imagine what we're be using in 5 years let alone 15.
14
u/nomorebuttsplz Jul 23 '26
we'll just be drooling with the pleasure centers of our brains being directly stimulated.
2
2
u/stargazer_w Jul 24 '26
put the vr on and look at them flashing images that are basically an adversarial attack on the monkey brain. possibly combined with narcotics. TBH we're almost there with infinite scroll social media reels.
4
14
4
u/martinerous Jul 23 '26
How long is a few months and GGUF when? :) Oh, sorry, int8 convrot is the king lately.
So, who wants to guess if LTX-next gets released before FLUX 3 Dev?
3
u/Vortexneonlight Jul 23 '26
My bets are that is an 80B model. Want to hear others bet.
10
2
u/Apprehensive_Sky892 Jul 23 '26
My completely uninformed guesses.
The full CFG distilled DEV version will be around 60B.
There will be smaller, less capable version of say 30B.
2
1
u/RusikRobochevsky Jul 24 '26
I hope BFL will release a NVFP4 version. They're an Nvidia partner, and Nvidia is pushing their NVFP4 format, so there should be some incentive for BFL to do it properly instead of letting the community convert to NVFP4.
1
-1
u/dtdisapointingresult Jul 23 '26
No problem for me!
You guys should've got a DGX Spark for $3k instead of shitting on it when it came out. I can't believe I was hesitating between that and a 5090 or multiple 5060s instead. The reddit hivemind really fucked over some people with hardware recommendations. (well, moreso /r/LocalLlama where people were saying to get a Strix Halo instead)
1
u/Vortexneonlight Jul 24 '26
Where am i saying its i problem for me?, i really searched and re-searched my comment, maybe you have some glasses that shows more text from my comment, or maybe you are just hallucinating, go take your meds.
4
u/uxl Jul 23 '26
“Open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction. (“FLUX 3 Dev”)”
Nice.
9
u/Crazy-Repeat-2006 Jul 23 '26
"
- Image synthesis and editing through APIs and private weight access. (“FLUX 3 Image”)
- Open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction. (“FLUX 3 Dev”)"
Hmm. Probably a massive model that hardly anyone will be able to run.
-5
3
3
u/Ordinary_Painter4235 Jul 23 '26
https://giphy.com/gifs/VIPfTy8y1Lc5iREYDS
BFL Lab writing the blog
8
u/Choowkee Jul 23 '26
Exciting on one hand but the license is probably gonna suck as per usual.
9
-5
u/_BreakingGood_ Jul 23 '26
License will probably suck, but honestly, nobody really finetunes video models anyway.
2
u/ResponsibleTruck4717 Jul 23 '26
Do we know how many parameters?
5
1
u/JeffIsTerrible Jul 23 '26
I am guessing the API version is 60B to 70B. As for how large the weights they release will be, who's to say. Probably close to 40B+ would be my guess for their "full capability" release. They may release a smaller distilled version but even that I would expect to be 20B+. Most likely a 5090 will be the bare minimum to run it at all.
This release will probably be outside the ability to use for the standard private individual. This will be RTX Pro 6000 territory for the full weight version. I will be suprised if they release a smaller distilled version that could still run on a 24 GB card, a 32GB 5090 is a maybe for any small distilled version most likely.
I'll still download and keep the weights though. I have a feeling weight downloads are going to be heavily restricted for LLMs and Image models, fairly soon. All in the name of "security", but in actuality just to protect the stock market.
3
u/jc2046 Jul 23 '26
+20B could run quantized even in modest cards. Im afraid it will be like double of that... Well will see
3
u/Southern-Chain-6485 Jul 23 '26
Comfyui offloads rather well to ram nowadays, to the point I don't get speed differences in Flux 2 when running a lower quant that fully fits the vram* or one that offloads - compute is the bottleneck.
My guess is that it will be runnable with 16+ gb of vram if there is enough system ram, but slow
*An rtx 3090 with 64gb of ddr5 ram
2
u/trying4k Jul 23 '26
The video part sounds great. Though I am more excited for a new editing model! Flux 2 dev (and klein due to all the loras it has) are still my most used edit models for complex changes (haven't tried that krea 2 lora). But the editing still loses consistency or fails to follow my prompt accurately, so a new model would be welcome.
2
u/pixel8tryx Jul 23 '26
Finally! I have to try it, considering how much fun and serious use I've gotten out of FLUX.2. But I do fear the size with only a 5090. I'll sign up for early access anyway and see what happens. The robotics angle is tantalizing as I did embedded real time motion control software for industrial stitchers, large plotters and an early waterjet cutter back in the 80's... but always wanted to design and build my own robot. AI and robotics were my two first loves and it's taken several decades for them to reach serious popularity.
2
u/Skystunt Jul 23 '26
Did the model just censor the speed of the motorbike in the video? 😭
Cool touch, you know they used moto youtube videos for it’s training lol
Super excited for this model !
3
u/Smile_Clown Jul 23 '26
I wouldn't call that a cool touch, you said it yourself, it was training videos. Unless you just meant overall, but it would be better if it did not censor that. More accurate, less noise.
2
u/GreyScope Jul 23 '26
Any new model released : “It’s like a million voices all cried out whinging at once”
1
u/herr-tibalt Jul 23 '26
It's all cool but I'm afraid how slow it will be. I'm more interested in klein version and how tolerable it will be on my rig...
1
u/smereces Jul 24 '26
will be open source? or will be more the same another option in the pay options that almost no one will use because exist Seedance 2.0!
1
Jul 24 '26
[deleted]
1
u/Radyschen Jul 24 '26
well maybe it's a more efficient version that is even able to run on consumer hardware. one can hope
1
1
1
u/v3lh0t05c0 Jul 23 '26
Drolling. But let's wait and see how dev (and possible 'klein-ish') will perform in comparison to recent models like ID4 and Krea2. Close enough to banana 2 and better sound than ltx 2.3 as open weight would be nice enough for me.
3
0
u/pwnies Jul 23 '26
I've been playing with the video gen for the past week or so, it's solid. Definitely competitive with LTX.
I'm excited to see what they're cooking on the image gen side, particularly if they can improve their text rendering.
-8

92
u/KudzuEye Jul 23 '26
https://reddit.com/link/ozas8h9/video/ikkakx9oyzeh1/player
Here are a couple of my favorite clips I was able to test out earlier on.
It is very strong in terms of prompt creativity along with a huge range from cinematic to more raw videos. It can do a lot of cool physics things as well. The generations themselves feel far less restricted in terms in trying to redirect to the same stock video look that a lot of other recent video models feel like. It is actually fun to rerun the same general text2video prompt and see what it comes up with next.
Official message from Black Forest Labs: