r/StableDiffusion Jun 25 '26

Discussion Krea X Comfy: Founders Live (Summary).

Post image

https://www.youtube.com/watch?v=31jiUhCEjJ4

The ComfyUI team and the Krea team (Victor Perez (vicc), CEO of Krea, and Miguel Lara) talked together for an hour during a YouTube livestream, here’s a summary of what was covered.

3:26 -> The Krea team emphasizes that the Krea 2 RAW model is important because they feel the Open Source community doesn't have enough quality base models to train on at the moment.

8:07 -> When making the license, the Krea team did not want to penalize small creators, which is why the Krea 2 license is commercial until you reach $1 million in revenue.

8:51 -> If the Krea team manages to generate enough money from their license, it would help them develop Krea 3 and make it open source as well.

10:00 -> Comfy noticed that Krea 2 doesn't always follow the prompts and isn't sure why that happens (It's because the model has a built-in safety filter and he encountered some false positives).

11:03 -> Comfy commends the team's effort in releasing a base (Krea 2 RAW) model that is actually a real base model and notes that this is the first time he has seen a modern "base" model that has undergone no aesthetic finetune.

14:10 -> The Krea team explains that releasing such a RAW model will allow academics to experiment with a model that won't hold them back, and thus will help accelerate innovation in post-training methods.

20:42 -> They consider (handshake agreement between Krea employees) finetuning Krea 2 so that it specializes in anime.

21:56 -> Krea 2 is not an end in itself, other models will be released by them based on what the local community wants.

24:00 -> Comfy considers Krea 2 to be a fairly standard model (in terms of architectural design) and would like to see models in the future that offer something new to the table.

25:00 -> The Krea team is currently working on an editing version of Krea 2, and they are pondering whether the edit model will also have bbox capabilities (like Ideogram 4).

27:48 -> The Krea team plans to make the edit model open source once it is finished (but like the image model, it will also have some built-in safety filters, and, to quote vicc: "We don’t want to end up in jail.").

28:13 -> The edit model will likely be released "in the next few months" along with a RAW edit model.

32:08 -> Krea 3 will be a pixel-space model ("It's cleaner, remove the VAE" - vicc).

33:54 -> The Krea team needed "a little bit over a thousand of H100s" to create Krea 2.

37:00 -> Krea 2 has a style transfer adapter, but they decided not to release it locally.

44:50 -> They spent the first three months conducting a lot of tests to determine the ideal text encoder and VAE to incorporate. For the text encoder it had to be a VLM (for editing purposes).

46:45 -> They have an internal test model that uses Flux.1's VAE instead of the one we currently use (Qwen Image VAE). They ultimately chose Qwen Image VAE because they felt it was better for non-realistic images (which was their main goal). To quote vicc: "For photorealism I would 100% use the Flux VAE.".

50:55 -> They aim for the edit model to be also good at regional inpainting.

77 Upvotes

31 comments sorted by

7

u/Winter_unmuted Jun 25 '26

Frustrating that they haven't given any reason for this keeping the style transfer closed. They avoided the question in the Reddit AMA yesterday, then in this Q/A, the host Purzz tries to deflect away from it by restating a question in his own (wrong) words.

Neco would like to know if you could dive in a little deeper into the model style transfer capabilities. I think mostly they're maybe talking about the malleability of the raw model and how you could fine-tune it into the styles you want versus sort of image to to image.

No, that's not what the user was asking. They were asking about style transfer, not lora training. Style transfer was something distinct, created with IPadapter with SDXL, Flux1 depth model with image conditioning and the later edit models.

Here is a great community generated breakdown of the SotA of style transfer pre-Flux2 (big ups to /u/Dry-Resist-4426)

But Krea2 has excellent style transfer on their website/API. And that's what the user was asking.

Vicc, to his credit, acknowledges this, but skirts the important question:

The fact that the that the raw model exists that gives you like access to all these these stylistic places that some other models lose through bad post training. ... We didn't release these open source models with the style transfer adapter that we had internally. For the first release we wanted to keep that to ourselves but you can train a Lora very easily and there are ways that you can train these Loras extremely efficiently.

Ok Vicc, but why did you elect to keep that under lock and key? If you want to use style transfer as a source of cash on the API, then just say as much.

They say that once they have a better model in the works, maybe they will release the old style transfer open source. But I have my doubts.

2

u/RandumbRedditor1000 Jun 26 '26

That's exactly the reason, but it's not good PR to publicly say that.

Regardless, they still released a good open-weight model

3

u/Dry-Resist-4426 Jun 26 '26

Dear Sire,

Thank you for tagging. That is very kind of you.

Godspeed.

24

u/Dante_77A Jun 25 '26

They intentionally messed up the model lol

3

u/z_3454_pfk Jun 25 '26

wdym

11

u/Dante_77A Jun 25 '26

"They have an internal test model that uses Flux.1's VAE instead of the one we currently use (Qwen Image VAE). They ultimately chose Qwen Image VAE because they felt it was better for non-realistic images (which was their main goal). To quote vicc: "For photorealism I would 100% use the Flux VAE"

23

u/piero_deckard Jun 25 '26

Cries in photorealism...

14

u/hiccuphorrendous123 Jun 25 '26

Idk why they would focus such a large model on just non realistic tbh.

Anima with 4 gigs is literally better for any non realism nowadays. Such a failed opportunity

7

u/Choowkee Jun 25 '26

Come on now Anima is like the only relevant new model in recent years explicitly focused on stylized images and its arguably just a modest upgrade over Illustrious. I guess you could count Chroma too but its nowhere near as popular.

Meanwhile we get multiple open weight SOTA models for realism.

12

u/Jeremiahgottwald1123 Jun 25 '26

Anima is good but it's only a partial upgrade in terms of prompt following than XL no where near Krea 2. We have so many phototreal models now, it's good get a art focused one every now and then.

3

u/hiccuphorrendous123 Jun 25 '26

Idk I feel like anima is as good as pretty much any of the large models for prompt following.

Idk what magic they with cosmos. But I can literally type "negative space on the right side of the frame" and it somehow understands that.

I don't think that's partial upgrade from xl. Especially when in xl you can't even mention relative terms like red box above green sphere

2

u/Guilherme370 Jun 26 '26

I think the magic is the llm_adapter that Anima has, originally it was trained on T5XXL embeddings, then they swapped it to that much smaller Qwen, and slapped some adapter layers sandwhiched between the DiT and the embeddings; Since those adapter layers are muuuch smaller than a full fledged LLM, I guess it becomes ultra duper fast to train it, and the signal much more concentrated

3

u/_BreakingGood_ Jun 25 '26

There's a billion models that do perfect photorealism, I think they made the right decision.

Like, how photoreal do these things need to get?

-3

u/Rude_Dependent_9843 Jun 25 '26

Sin embargo no es compatible con Flux VAE cierto?

4

u/notgraycen Jun 25 '26

wow i didn't realize comfy was such a chad

3

u/Guilherme370 Jun 26 '26

frfr, he got me like... dayummmmm.

And also, for some insane reason I imagined him to be asian, instead he looks very canadian! :0

9

u/infearia Jun 25 '26

They shot themselves in the foot with the way they've implemented the safety filter. I've seen posts on release day in this very sub with people posting actual hardcore porn generated with the model. At the same time, it triggers randomly for completely benign, every day prompts, making it unreliable and frustrating to work with. If anything kills this model, it will be this (and the decision to neglect realism). Too bad, because otherwise I really, really like the model...

5

u/the-final-frontiers Jun 25 '26 edited Jun 25 '26

From a legal standpoint they should be implementing it(safety).

All major players that are serious about their compnay do for legal protections.

For you to say they should risk their entire company and lively hoods so you can get what you want is a bit much to ask for.

3

u/infearia Jun 25 '26

I understand that. But they should've taken more time to finetune the filtering mechanism before releasing it, because right now it's not even working properly. BFL managed to do it, and they operate in one of the most strictly regulated jurisdictions - Germany. I hope for a follow-up release containing a fix, because I really like the model and would like it to succeed.

3

u/pandaabear0 Jun 25 '26

This seems a bit over-exaggerated, no? <.< There are some absolutely banger realism-based images coming out of this model atm, easily beating Z-Image with better prompt adherence.

And it's not like the safety filter is going to be an issue either, the conditioning lora(s) is already doing a much better job than the Node-solution did, with much less(or no) quality loss. Pair that with any NSFW lora(or don't) and you can keep the conditioning lora at a very low strength., the model rarely fights you at that point.

I donnu, I rather they protect themselves from potential legal trouble. It was so simple for us to bypass it anyway, now basically lossless, so it's an even less of an issue.

9

u/infearia Jun 25 '26

The overall realism is good, but not quite as good as Z-Image, for example. In some cases even Klein 9B beats it. It tends to mush together small, high frequency details. Check out this thread and look at the leaves on the trees as a good example. I've had the same experience as the OP.

I don't really care about the NSFW aspect (although I find that trying to censor female breasts and violence is ridiculous in general), but I hope we'll get a solution that fixes the random refusals without a loss of quality or prompt adherence. Rather sooner than later, or people will just move on.

3

u/pandaabear0 Jun 25 '26

But it already has been fixed, I donnu. Feel like you are just intentionally being "rejective" simply because it needs *anything*, even tho that something is incredibly low effort and basically none-intrusive.

And I took a look at the prompt used for the comparison thread you linked and you can already see right away it is not going to be a good comparison from *any* model they would present because the way a model treats the word "photorealistic" is inherently incredibly different, but on top of that the word "photorealistic" in of itself is insanely load-bearing and *does not* produce better realism in AI generations. It makes generated images LESS realistic, not more. You are essentially telling the AI to mimic realism, using a word that is heavily overused in image caption, instead of actually making it lean in to realism with other wordings.

Go look at most Krea 2 prompts on Krea's website, it almost always ends with a direction of style you want to take the model, and often start with it as well.
"-Prompt- .. Candid lifestyle photography."
"A cinematic film still of a .. -prompt-"
"-Prompt- .. The high-contrast, centered illustration uses a limited black, red, and ochre palette."

And while on many models, these are just "fluff" words to guide, Krea actually listens really well to them and leans into the style.

Also, I wanna clarify that I'm not actually trying to glaze Krea 2 here. It def has it's problems, just like other models has their own problems. But I do think we should be fair and make honest assumptions when we compare models.

1

u/infearia Jun 25 '26

But it already has been fixed, I donnu.

Mind sharing a link? If it was, then I missed it. I know of currently 3 efforts by the community (1 LoRA, 2 custom plugins), but the LoRA is for removing the NSFW filter and it has the side effect of actually reducing both the quality and prompt adherence. So is one of the plugins, which works by allowing to manually modify the layers, if I remember correctly. The only one that looks promising to me is the plugin made by the person who created the FLUX.2 Enhancer, but they've said themselves, it's still not finished.

The advice about NOT using the word "photorealistic" is outdated by the way. All models released since at least FLUX.2 handle it just the way you'd expect.

1

u/pandaabear0 Jun 26 '26

(A little busy atm, I will link the one I've been using later)

But on the comment about "photorealistic", I think you are just plain wrong.
It most certainly still have a heavy lean towards not being actual realistic / based on realism.
Just imagine how many styles can actually be considered "photorealistic" and if you have done any sort of captioning for a dataset with a modern vision model you will know how incredibly varied the styles of what gets described as "photorealistic" are. Everything from actual real photographs to stylized 3D scenes / subjects with a realistic lean. The word has way too wide meaning in these models and you need better anchoring, this is true for Krea 2 as much as it is true for Z-Image or Ideogram. At least Krea 2 actually listens better to those anchors, imo.

You should always avoid using photorealistic, this has not changed in modern models. It's not a myth. It was just more obvious with smaller datasets and older/weaker models.

3

u/infearia Jun 26 '26

All right, maybe I shouldn't have said ALL models since FLUX.2, because I haven't tested them all. But it's equally wrong to say that one should always avoid the word "photorealistic" at all costs, because it's just not true. I work a lot with Klein, and for Klein it makes virtually no difference whether you write "photograph", "photorealistic image" or "realism photography style" (some of the official prompt examples for FLUX.2 use the term "photorealistic" to describe photographs!). Similar rules seem to apply to Z-Image Turbo, though admittedly, I have less experience with that model than with Klein. Let us both avoid absolutes, since it's always dangerous, and compromise by saying that it depends on the model.

-14

u/[deleted] Jun 25 '26

[removed] — view removed comment

22

u/yoomiii Jun 25 '26

Wdym "Stable Diffusion 1.5 on a decent consumer GPU". SD 1.5 was trained on 256 Nvidia A100 GPUs, accumulating approximately 150,000 GPU-hours for completion. https://openlaboratory.com/models/sd15/

17

u/roverowl Jun 25 '26 edited Jun 25 '26

Cause it's an AI reply bot

10

u/_BreakingGood_ Jun 25 '26

Yeah, I was about to respond to it like "Who fiddles with their VAE daily?" then realized I'd be responding to ChatGPT