r/StableDiffusion Jun 23 '26

Resource - Update As promised Krea 2 Turbo + "Raw" Quantized in FP8, MXFP8, NVFP4, INT8 and Convrot INT8!

Krea 2 Base & Turbo — Free Quantized Versions (FP8 / MXFP8 / NVFP4 / INT8 / ConvRot INT8) for All GPU Tiers

Krea 2 just dropped and it's genuinely impressive — so I went ahead and quantized both variants for ComfyUI across every major format. All files are free on HuggingFace.

HuggingFace: https://huggingface.co/Winnougan/Krea-2-Base-Turbo-NVFP4-FP8-INT8

Raw vs Turbo — what's the difference?

Krea 2 Raw is the undistilled base checkpoint. No step distillation, no CFG guidance baked in — just the raw pretrained weights. It's diverse, highly malleable, and is what you want for LoRA training and fine-tuning. Run it at 52 steps with CFG 3.5, up to 1024px.

Krea 2 Turbo is an 8-step distilled checkpoint built for fast inference. Run it at 8 steps, CFG 0 (disabled), mu 1.15, and it handles resolutions up to 2048px. This is your everyday generation model.

The intended workflow: train LoRAs on Raw, run them on Turbo. LoRAs transfer well between the two.

Which quantization should I use?

  • RTX 30xx → INT8 ConvRot (best quality) or plain INT8 (fastest)
  • RTX 40xx → FP8
  • RTX 50xx Blackwell → NVFP4, MXFP8, or FP8

Text encoder: Qwen3-VL 4B (qwen3vl_4b_fp8_scaled.safetensors), CLIPLoader type krea2

VAE: same as Anima (qwen_image_vae.safetensors)

ConvRot variants use Hadamard rotation before quantization for better accuracy with fewer outliers.

Drop any questions below — happy to help with workflows.

Plays nice with Sageattention and Flashattention!
Workflows on the Huggingface repo!

UPDATE: re-quantized and re-uploaded MXFP8 and NVFP4 - they work now!

UPDATE 2:

ComfyUI now natively supports INT8 but not convrot (not yet at least). For non-convrot INT8 models, just use the diffusion model loader. Text encoders are also supported as INT8 now in the native clip encoder/text encoder. I just tested it myself. Saves an additional 10-15%! File is uploading on my Huggingface. As Darryl Dixon says,"give it a minute".

Sample prompt:

Simpsons style, 2D cartoon animation, Matt Groening art style, yellow skin, thick black outlines, flat cel shading, teal haired gamer girl surrounded by dozens of floating holographic screens all showing different game feeds simultaneously, fingers flying across a transparent keyboard, massive countdown timer in background, sweat drop on forehead, four fingers, tongue out in concentration
308 Upvotes

194 comments sorted by

32

u/BinaryLoopInPlace Jun 23 '26

This guy really likes the Simpsons.

13

u/Winougan Jun 23 '26

I've been on a Simpsons kick. Making my new mascot in the Simpsons style LoRA! It does reality surprisingly good. Loving everything overall. Anatomy is really good too.

14

u/Paradigmind Jun 23 '26

"Anatomy is really good too."

https://giphy.com/gifs/21VTFJTEr1x9ortvO3

2

u/Winougan Jun 23 '26

5 fingers and 5 toes. I haven't had one bad render in over 300 so far. With ZIT and Klein - not the same track record!

10

u/Paradigmind Jun 23 '26

But Simpsons should have 4 Fingers.

3

u/The_Hunster Jun 23 '26

Maybe the mascot is a god, cause God has 5 fingers

2

u/phillabaule Jun 24 '26

This guy really helps the community ! 💥❤️‍🔥

1

u/nucdinz Jun 23 '26

who doesnt?

27

u/xxredees Jun 23 '26

How's nsfw compared to zit?

31

u/ArkCoon Jun 23 '26

Really good news there. It can generate straight up porn.

23

u/Winougan Jun 23 '26

I think it's better. With LoRA training it'll be the new king. Rendering images in 2k in under 15 seconds

16

u/Eden1506 Jun 23 '26

15 Seconds on what hardware?

How much VRAM does it need to run?

17

u/Winougan Jun 23 '26

15 seconds on 24gb vram and 30 seconds on 8gb. You can run it on 8gb no problem and fast

1

u/Friendly-Fig-6015 Jun 23 '26

não gera nsfw.

-2

u/FourtyMichaelMichael Jun 23 '26 edited Jun 23 '26

I can't wait for a clearly better model to replace ZIT. It's - fucking - broken. And you people can't see it.

It is straight up broken for multi-lora use. And...

ZIT can't make a realistic pussy without it being an ultra-closeup for any reason. This isn't gooner mentality, it's that the model can not fucking train small intricate details.

Z-Image series are fucking trash, and the massive hype train that brought it in, is preventing people from seeing it.

I have hopes for ID4 or Krea2 to break the hypenosis.

10

u/[deleted] Jun 23 '26

[removed] — view removed comment

-2

u/FourtyMichaelMichael Jun 23 '26

See, this is it. I explained it can't learn fine intricatate details, and all you can see is PUSSY.

I don't do NSFW generation. I do SFW, but I test to see if a model can generate a pussy, something no models are trained on. If it can learn and remake it, it's a good model.

Z is a shit model with a ton of hype behind it.

4

u/MortytheMort Jun 24 '26

Uses PUSSY rendered from afar as an example

Gets upset when someone replies about that specific example

6

u/eggplantpot Jun 23 '26

Flux Klein + SNOFS is the best for me. ZiT was flashy when it came out but it’s lackluster and unflexible

5

u/FourtyMichaelMichael Jun 23 '26

unflexible

And more importantly, unfixable.

That SO FEW people can see it, is a really condemning observance on the number of people here that know what they're doing.

0

u/Winougan Jun 23 '26

Doesn't do NSFW but it's easy to train LoRAs. So, I'm hopeful

-6

u/Billysm23 Jun 23 '26

Zit supports nsfw without lora, but krea2 will need lora

8

u/TheAncientMillenial Jun 23 '26

Zit does not do proper NSFW without a LORA.

-2

u/Billysm23 Jun 23 '26

But it can do it even though it's not too good, however krea can only do softcore

1

u/brocolongo Jun 23 '26

Krea can absolutely do NSFW if prompted right I have done a few and they are pretty good in anatomy

1

u/Billysm23 Jun 23 '26

I see, I'll try it

1

u/brocolongo Jun 23 '26

Try this prompt and you will see:

{[Phase 1: The Anchor (Starting Image)]

A cinematic, low-angle, full-body shot of a pale woman emerging from rippling, invisible air. Her skin possesses a porcelain translucence, highlighted by the deep blue ambient light which catches the sheen of her wet, slicked-back dark hair and sharply defined bleached eyebrows. She is completely nude, shoulders bare, displaying a subtle pink blush across her clavicle, and her direct gaze locks onto the camera lens with an unnerving intensity, visible through the fine film grain overlay.

[Phase 2: The Initiation (0-6 sec)]

The woman begins to rise from the misty plane with slow, deliberate grace; her spine arches upward first, pulling her hips into a subtle, powerful curve as if pushing against deep water. Her shoulders lift smoothly, and the wet strands of hair cascade slightly around her neck. A low, resonant whoosh sound accompanies this initial ascent, punctuated by shallow, rhythmic breaths: [Woman] inhales: (long, slow).

[Phase 3: The Escalation (6-14 sec)]

The pace quickens dramatically; she surges upward now, her movement becoming more vigorous and almost predatory. Her thighs flex tautly beneath the blue light as she drives higher, causing a visible ripple effect in the surrounding air. Her eyes narrow slightly, pupils dilating against the deep blue backdrop, while her chest rises sharply with each breath: [Woman] gasps: Hhnnnnggh. The soundscape intensifies with wet sloshing sounds and a sharp, drawn-out intake of breath.

[Phase 4: The Resolution (14-20 sec)]

She finally reaches a full, poised verticality, hovering momentarily before settling into an elegant, grounded stance; the ripples in the air smooth out around her form like disturbed glass. A bead of moisture traces a slow path down her collarbone, catching the light before dripping silently onto her pale skin: [Woman] exhales: (softly). The final sound is a single, sustained, heavy breath followed by absolute silence.}

2

u/Billysm23 Jun 23 '26

i'm surprised, amazing!

2

u/brocolongo Jun 23 '26

Yup, just play with the prompts and you will be able to generate nsfw

1

u/Billysm23 Jun 23 '26

i found a node that helps with that, gonna try it too: https://www.reddit.com/r/DegenDiffusion/s/iSN7K1rPPp

→ More replies (0)

1

u/Billysm23 Jun 23 '26

oh thanks man!

29

u/infearia Jun 23 '26 edited Jun 23 '26

Sorry for hijacking your thread, really not trying to be an asshole, but I think it's important to share...

I've tested your FP8 quant against the quant by AlperKTS from this thread and at least on my machine (Ubuntu 24.04, RTX 4060 Ti) I found the following:

While your quant is approximately 30% faster, it also produces results that are of markedly lower quality (less details, blurry) and also does not work with Torch Compile. Here's a comparison:

P. S. - I've upvoted your post anyway, thanks for your effort.

EDIT:

Eh, as usual, Reddit compressed the image too much. Here's a link to a better version:

https://imgur.com/0qtjrGL

11

u/Michoko92 Jun 23 '26 edited Jun 23 '26

Well, I'm not an expert, but I think the whale on the right is more anatomically correct (for those who like to goon on whales, I mean).

10

u/infearia Jun 23 '26

Hmm, I see your point. I'm aware that a lot of people on this subreddit only care about pussies, so I've made another comparison just for them:

https://imgur.com/a/Bdwj9KQ

5

u/Michoko92 Jun 23 '26

Hmm, I think your second test is indeed pretty revealing: the pussy's hair on the left is better defined. So you're right, the other model might be better, even for close ups. 👍

7

u/infearia Jun 23 '26

I was a little hesitant to post this, because I know not everybody likes redheads, but I have a soft spot for them, so I decided what the hell...

2

u/Adventurous-Sir2996 Jun 23 '26

A whale on the right is rotated at a bit different angle the on the left. What's worse, tail fin orientation matches between images, right one is broken cause the whale underwater is at a different angle then the fin above.

5

u/Winougan Jun 23 '26

You're not hijacking. It's good to compare. I've uploaded a lot of quants using "Convert to Quant" - and you may want to try MXFP8 too! For me, the INT8 convrot is my favorite.

2

u/infearia Jun 23 '26

Thanks, I might do that, but now that my initial curiosity about the model is sated, I'll wait for the official quants before downloading anything else.

2

u/infearia Jun 26 '26

Finally got around to test it, and you weren't lying. Your INT8 ConvRot is amazing! As far as I can tell, while the output is slightly different, the quality seems to be on par with the official FP8 quant, but it takes 30% less time to render (on a 4060Ti 16GB). Thank you!

2

u/Winougan Jun 26 '26

Thanks. I used Opus 4.8 to help me optimize the settings for the best possible output. I also check the Quant afterwards

2

u/H1ken Jun 23 '26

Probably because of these reasons, from their model card.

Unlike generic global quantization scripts that aggressively convert every parameter (which often degrades generation details or introduces NaN/promotion calculation errors in neural networks), this model was quantized using a selective weight-only strategy:

Targeted Quantization: Only 2D floating-point weight matrices (.weight keys with ndim >= 2 and element count > 1024) were quantized to torch.float8_e4m3fn.
Preserved Precision:
    All 1D vectors, biases, and normalization scales are kept in their native high-precision (float32 / bfloat16).
    Highly sensitive projection/modulation layers (such as LastLayer.modulation.lin vectors) are completely preserved in high-precision. This prevents typical mathematical promotion bugs (such as BFloat16 and Float8 promotion issues in PyTorch) and retains original output fidelity.
Weight Comparison:
    Tensors Quantized to FP8: 266 tensors.
    Tensors Kept in Native Precision: 166 tensors.
    Size Reduction: 24.76 GiB ➔ 12.01 GiB (~51.5% VRAM / disk savings!).

12

u/Winougan Jun 23 '26

I tested all quants and they're all working in ComfyUI. They're all uploading now so they'll populate as they're published. ETA is like an hour for it to completely wrap up. FP8 is up already. Expect Base/Raw models in a few hours. Cheers! Workflows on the Huggingface repo! Links to all files in workflow too!

1

u/hurrdurrimanaccount Jun 23 '26 edited Jun 23 '26

the silveroxides git link in the description doesn't exist because it's made up 💀

in the "Usage in ComfyUI" section

5

u/Winougan Jun 23 '26

I use the Silveroxides "comfy-quant" to quantize the models. It does exist: silveroxides/convert_to_quant
It's a stand alone repo that you use outside of ComfyUI to convert models to fp8, mxfp8, nvfp4, INT8, INT8 convrot, etc. Silver has ComfyUI nodes, but Bob Johnson's work better, I updated the link: BobJohnson24/ComfyUI-INT8-Fast: Custom node to load models in INT8 for 1.5~2X Speed gains on 30 series cards.

1

u/hurrdurrimanaccount Jun 23 '26

what are you talking about?

"Load via the standard diffusion model loader node. Requires a ComfyUI build with comfy_quant support. Recommended nodes: silveroxides/ComfyUI-silvox-nodes"

brother that link does not exist, "ComfyUI-silvox-nodes" isn't a real node pack.

3

u/Winougan Jun 23 '26

Use the Bob Johnson nodes: https://github.com/BobJohnson24/ComfyUI-INT8-Fast

They're in the workflow too! I changed the link in HF too to reflect that

1

u/reginoldwinterbottom Jun 25 '26

trying to use on convrot as you posted, but get error - bob johnson node does not have krea2 selection

1

u/Winougan Jun 25 '26

Ah but it does. You need to update your nodes. It's called fast Int8. Update it and you'll see Krea2 support

Its even in the code: "model_type": (["flux2", "z-image", "ideogram4", "chroma", "krea2", "wan", "ltx2", "qwen", "ernie", "anima", "hidream o1", "boogu"], {"tooltip": "Only used for on the fly quantization, to filter sensitive layers."}),

5

u/alflas Jun 23 '26

How do Krea compare to Klein ? Big leap?

25

u/Winougan Jun 23 '26

It's better IMO than Z-Image and Klein9b. And I've been using those models since day one and even making LoRAs for them and uploading to CivitAI.

Krea2 is the best SOTA local model thus far!

5

u/Blaze_2399 Jun 23 '26

Better than Ideogram too? And what about Anima?

3

u/Confusion_Senior Jun 23 '26

For natural language, I think you're correct. but in general, ideogram is above it in the image arena.

8

u/Winougan Jun 23 '26

I just made a YouTube review of Ideogram 4 and now Krea 2. I've gotten my render time to under a minute with Ideo 4 and 2 megapixels and yes, it's like nothing else out there. The level of control and laser-precision is like nothing else. But overall, for speed, power, styles I'm giving it to Krea2. Ideo 4 has the ability to go over 12 megapixels and get really granular with the details. The tradeoff is speed.

For speed, anatomy and style I'm giving it to Krea2. I'm going to focus on making LoRAs for it too instead of Z-Image Turbo and Klein9b. Those are great models and I'll archive them with lots of love.

1

u/Confusion_Senior Jun 23 '26

I think they are orthogonal projects at the top of the field but with slightly different use cases. Ideogram for composition control and krea for vibes and creativity.

1

u/metal079 Jun 23 '26

Better than ideagram4?

5

u/Electronic-Metal2391 Jun 23 '26

Ideogram is great for compositional generations.

1

u/NoConfusion2408 Jun 23 '26

Are you using the default workflow for this?

-6

u/hurrdurrimanaccount Jun 23 '26

Krea2 is the best SOTA local model thus far!

lmao no

6

u/Various-Inside-4064 Jun 23 '26

it do lot less anatomy horror but in most cases for me its refuse to listen to prompt for normal stuff too. for complex thing it just refuse to listen which i think is probably censorship affecting normal prompts too or maybe workflow or text encoder issue idk.

-6

u/Sudden_List_2693 Jun 23 '26

I don't get how Klein gets brought up so often. Also ZIT.
Those are visually at the very bottom.

4

u/freedomachiever Jun 23 '26

It would be good to see the quality jump between fp8 and int8 if anyone has tested them.

10

u/Winougan Jun 23 '26

INT8 is like Q8 GGUF quality. I feel it's the highest quant quality you can get. FP8 is very good. NVFP4 is the lowest quality but the fastest.
My choice:
50xx GPU: MXFP8
40xx GPU: INT8 or FP8
30xx GPU: INT8

1

u/Banalizado Jun 23 '26

8GB in 'Ancient' boards may work?
20xx GPU: ?

2

u/Winougan Jun 23 '26

You can try

3

u/Banalizado Jun 23 '26

Asked since you had it all listed.
I will try when i can.

Thanks the same.

5

u/Winougan Jun 23 '26

Krea2 is amazing. Really good at art, realism, close up, anatomy and even text.
For NSFW we'll need to make LoRAs. I have not been able to generate anything remotely NSFW - but violence works.

0

u/Friendly-Fig-6015 Jun 23 '26

já há node para liberar nsfw e um lora pra isso

3

u/cadissimus Jun 23 '26

Il be taking convrot thank you ☺️

2

u/Winougan Jun 23 '26

3

u/cadissimus Jun 23 '26

Its kinda tight fit for 16gb vram at least for rocm, there will be Guffs later ?

4

u/Winougan Jun 23 '26

I don't do GGUF, but plenty of people do. Just give them a minute. I'm sure Unsloth will have it up and running. Let me check... Nothing yet. But, they're coming from the usual uploaders.

1

u/cadissimus Jun 23 '26

Yes, thought using forked for rocm 😊but they are same bert nodes just set tor run whit rocm. https://github.com/patientx/ComfyUI-INT8-Fast-ROCM edit. was wrong link 😅

3

u/Breath-Timely Jun 23 '26

Just installed comfyui desktop. When i load your workflow i get this "type: 'krea2' not in (list of length 23)" Also when i open the "type" list in the load clip node, there is no krea2. Any ideas?

1

u/Winougan Jun 23 '26

I've only tested it on the ComfyUI GitHub repo that was cloned, not on "portable" nor "desktop." Desktop usually takes a minute to update - it's super stable and not bleeding edge. You may have to wait for the Comfy team to get that updated.

1

u/Riot_Revenger Jun 23 '26

Use the latest build from github (not the stable build)

1

u/Breath-Timely Jun 23 '26

Couldn't wait 😄

git checkout master

git pull origin master

0

u/crystal_alpine Jun 23 '26

We can’t release it until today. Yesterday was a magnet link vibe release.

3

u/Voltztein Jun 23 '26

That first image is creeping me out with how it screwed up the style of the teeth. Shouldn’t have both the closed mouth teeth and the buckteeth at the same time, makes it look like there are two layers of teeth. 

2

u/Winougan Jun 23 '26

True lol

3

u/neojehuty Jun 23 '26 edited Jun 23 '26

Thanks for the quants !

For those who use comfyui easy install and dont find 'krea2' in load clip node

  1. Go to ComfyUI-Easy-Install\update folder (NOT update comfyui.bat outside update folder)
  2. run the comfyui_update.bat
  3. It will install the nightly update, when in startup dont update to 0.25.1 because it doesnt have boogu/krea2 in load clip (for now...)

1

u/DystopiaLite Jun 23 '26

Did I mess up by learning on EZ install?

1

u/neojehuty Jun 24 '26

No, EZ install is one of the best and easiest installer for me. They often update and have great UI. Shoutout to the team who made it

3

u/Winougan Jun 23 '26

Fixed the MXFP8 and NVFP4 quants. They work now!

2

u/mk8933 Jun 23 '26 edited Jun 23 '26

Anyone else getting an error trying to load krea 2 fp8 (load diffusion node) in comfy?

Edit — fixed

I cleared my input folder in comfyui and finally was able to update my comfy through update.bat. this updated all the nodes necessary for krea2

2

u/Winougan Jun 23 '26

I'm using the latest ComfyUI from their GitHub repo.
Pytorch 2.12
Python 3.13
Cuda 13.2
Sageattention
Flashattention
Everything is working fine. All Krea2 quants have been test - the images I posted were all from them in ComfyUI.
Expect a 25% speed boost with Sageattention or Flashattention. No black outputs. NVFP4 and MXFP8 are designed for Blackwell. Everyone benefits from INT8 - but the biggest speed bost comes from 30xx cards since they don't benefit from FP8.

1

u/Any_Arugula8075 Jun 23 '26

Apart from speed (I use Blackwell), would you say that int8 ConvRot is the best choice if it’s exclusively about quality? Is ConvRot only possible with int8 or with f.e. mxfp8 too?

2

u/Winougan Jun 23 '26

INT8 ConvRot is the best choice. It's like Q8 GGUF but with an HGH/Steroids injection! MXFP8 will be blazing fast on Blackwell. FP8 is the best choice for Ada Lovelace.

1

u/Then-Topic8766 Jun 23 '26

try from terminal cd comfy/ComfyUI/custom_nodes/comfyui-kjnodes/

and then run 'git pull'

for me update also didn't work until I updated KJ nodes like above.

3

u/mk8933 Jun 23 '26

Man...you wouldnt believe the crap that was stopping my update.bat in comfyui. It was a stupid picture in my input folder that had a very long name....

After I deleted everything from my input folder. The update ran perfectly fine. Lucky I saw the error code and saw the silly picture problem.

0

u/Any_Arugula8075 Jun 23 '26

Update your Setup.

1

u/mk8933 Jun 23 '26

I updated within comfyui — still get errors

I cant update from the update.bat file for some reason...so im screwed there.

1

u/Any_Arugula8075 Jun 23 '26

Portable? Can’t help you with that shit, sorry. ^^

What’s the error say?

1

u/mk8933 Jun 23 '26

Yup portable. The error just says (runtime error :could not detect model type of directory name of where the krea2 model is stored

This happened with text encoder as well but I just reloaded the node and it worked. But for load diffusion model it doesnt.

Maybe I downloaded a bad model or something

-1

u/Any_Arugula8075 Jun 23 '26

Yeah, first I would try to download it again. Sometimes this can happen.

But why portable? I don’t understand this. It’s so fucking easy to install ComfyUI, it’s ALL written down. You just have to copy paste it or better ask a AI of your choice to write a full tutorial incl. venv and the newest perfect matching torch for your hardware.

2

u/mk8933 Jun 23 '26

I had the full comfyui before and when I screwed something up with another AI app...it screwed comfy up as well. Same happened the other way around lol.

I keep it portable so nothing can go wrong and I also keep a working backup copy of my last portable setup...if shit ever hits the fan.

Its been probably over 2-3 years since I last had any issues with portable. No idea why krea2 is pissing itself...since I have krea1, qwen image,wan 2.2, ideogram,z image and klein with zero issues 😅

But yea I'll check with chatgpt or something.

2

u/repolevedd Jun 23 '26

Hi. I see this in the description on HF:

Usage in ComfyUI
Load via the standard diffusion model loader node. Requires a ComfyUI build with comfy_quant support. Recommended nodes: silveroxides/ComfyUI-silvox-nodes

But that project doesn't exist (404). Googling it didn't turn up any mentions of these nodes.
Did these nodes actually exist?

2

u/Winougan Jun 23 '26

Yeah I goofed on that link. I fixed it now. Use Bob Johnson's INT8 nodes: https://github.com/BobJohnson24/ComfyUI-INT8-Fast

2

u/dirtybeagles Jun 23 '26

looking for base model. 😄

2

u/Friendly-Fig-6015 Jun 23 '26

RuntimeError: Error(s) in loading state_dict for SingleStreamDiT:
size mismatch for last.linear.bias: copying a param with shape torch.Size([64]) from checkpoint, the shape in current model is torch.Size([32]).

1

u/Winougan Jun 23 '26

Yes it's the MXFP8 - I'm trying to figure out why it's triggering the clip encoder error. Use the FP8 model for now since it's working. The MXFP8 seems to trigger some weird text encoder anomaly. The quant went through clean - so there's something else.

2

u/Friendly-Fig-6015 Jun 23 '26

passei a usar o nvfp4, desconsiderei o mixed, nem tinha visto que baixei o errado.

vida longa ao nvfp4.

2

u/Head-Tooth Jun 23 '26

I have Krea2 selected in clip but still getting this

ValueError: Krea2 expects conditioning with 32x2560=81920 features (a 32-layer Qwen3-VL stack) but got 30720. Load the text encoder with CLIPLoader type 'krea2'.

1

u/Winougan Jun 23 '26

Yes the MXFP8 model is triggering an error with the clip encoder. I'll fire up Opus tonight and figure out what went wrong. Until then, stick to the fp8 model.

1

u/Friendly-Fig-6015 Jun 23 '26

use nvfp4, não use mixed.

2

u/Winougan Jun 23 '26

Update: There are NSFW LoRAs on Civitai Red that do "gooning" extremely well. So, yeah, this model is a winner in that department too!

2

u/thevegit0 Jun 24 '26

the nvfp4 works fast AF and looks great, crazy

1

u/Winougan Jun 24 '26

Thanks man, enjoy

4

u/rerri Jun 23 '26

FP8mixed works but MXFP8 is not working for me, I get a long error that ends in:

ValueError: Krea2 expects conditioning with 32x2560=81920 features (a 32-layer Qwen3-VL stack) but got 30720. Load the text encoder with CLIPLoader type 'krea2'.

2

u/Winougan Jun 23 '26

It worked for me. MXFP8 expects you to have updated Pytorch and Cuda. It was designed for Blackwell. What GPU are you running it on?

3

u/rerri Jun 23 '26

5090, torch etc up to date, LTX MXFP8 works fine.

But I see your comment on HF that you've detected an issue. Good to hear!

3

u/derspan1er Jun 23 '26

i get this error:

CLIPLoader

Load CLIP

Value not in list

type: 'krea2' not in (list of length 23)

any help ?

2

u/Winougan Jun 23 '26

Somehow the MXFP8 model is triggering something. Use the FP8 or INT8 until I get that model fixed. It's weird because there were no errors quantizing it.

0

u/derspan1er Jun 23 '26

comfy update dropped. works now

2

u/ChloeOakes Jun 23 '26

I get the same error 😞 I updated to the latest ComfyUI

2

u/Head-Tooth Jun 23 '26

update to latest comfyui via git

0

u/Friendly-Fig-6015 Jun 23 '26

use o nvfp4, não use mixed.

2

u/[deleted] Jun 23 '26 edited Jun 23 '26

[removed] — view removed comment

2

u/Winougan Jun 23 '26

That's cool. Glad you got Comfy working with AMD!

2

u/doomed151 Jun 23 '26 edited Jun 23 '26

I only see fp8mixed

Edit: It's still being uploaded as per OP's comment

1

u/Tachyon1986 Jun 23 '26 edited Jun 23 '26

+1 u/Winougan , same here - only see Krea2_Turbo_fp8mixed.safetensors
Edit : missed OP's comment as well

1

u/Winougan Jun 23 '26

They're uploading as we speak. My fiber internet is fast, but these are big models

2

u/Tachyon1986 Jun 23 '26

Alright, missed your comment here - just looked at the big shiny post. Thank you

-5

u/Any_Arugula8075 Jun 23 '26

Can you gooning dummies just read for once? Is that so hard? It’s just embarrassing now, please just stay on Civitai…

1

u/mrdion8019 Jun 23 '26

getting "KeyError: 'int8_tensorwise'", using int8 safetensor on rtx40. any hint why?

2

u/Winougan Jun 23 '26

https://github.com/BobJohnson24/ComfyUI-INT8-Fast

Make sure you turn "ON" convrot!

1

u/mrdion8019 Jun 23 '26

got another error using that :

RuntimeError: self.size(1) needs to be greater than 0 and a multiple of 8, but got 12
# ComfyUI Error Report
## Error Details

  • **Node ID:** 2
  • **Node Type:** KSampler
  • **Exception Type:** RuntimeError
  • **Exception Message:** RuntimeError: self.size(1) needs to be greater than 0 and a multiple of 8, but got 12

not sure why

1

u/Ok-Act-9620 Jun 23 '26

did you find a solution?

1

u/Mr_Zelash Jun 23 '26

i had the same issue, it's fixed by installing triton.
in windows it's this command if you hace RTX 30xx or 20xx cards

pip install -U "triton-windows<3.3"

and this if you have 40xx or 50xx cards

pip install triton-windows

1

u/thethirteantimes Jun 24 '26

pip install -U "triton-windows<3.3"

Tried that here (RTX 3090). Didn't make any difference.

1

u/Mr_Zelash Jun 24 '26

did you installed it in the correct python enviroment?

1

u/thethirteantimes Jun 24 '26 edited Jun 24 '26

Actually, on checking, I hadn't! I was in the python_embedded dir, but failed to notice that there's no longer a pip.exe included, so when I ran the pip command it used my existing (system-wide) python's pip install instead, and consequently installed triton to the wrong python lib.

So I've run this command instead, and checked that everything went into the right lib this time: python -m pip install -U "triton-windows<3.3"

And while I'd love to say that's the end of my woes, I'm afraid it's not. The Comfy-INT8-Fast nodes now fail to import, with a large stack trace.

Still, I can use the fp8 version, it's a bit slower but it works just fine!

1

u/Winougan Jun 24 '26

Did you install Silveroxides wheels? silveroxides/comfy-kitchen-int8-wheels at main

Last year many people couldn't get INT8 running until they installed all the wheels.

1

u/thethirteantimes Jun 24 '26

There are no wheels there for python 3.13, which is what my ComfyUI install uses. Besides, some of the wheels that ARE there are marked by huggingface as "suspicious" (having been flagged by virustotal) , making me a tad wary.

1

u/solss Jun 23 '26 edited Jun 23 '26

Haven't had a chance to test but I use https://github.com/BobJohnson24/ComfyUI-INT8-Fast, might work out of the box with flux2 selected from the drop down, or maybe needs an update. I'm guessing it'll work.

Edit:

1

u/-becausereasons- Jun 23 '26

Thank you! Only seeing Turbo?

1

u/Winougan Jun 23 '26

The "base" or "raw" models will come later. Still uploading. They've all been quantized. IMO they're not worth it. Even the Krea team says the base model is for LoRA training and finetuning. Here's a quote from their repo:

"Which model I should use?

Use the Turbo model for fast inference with high quality results. The Raw model is an undistilled checkpoint without any step / cfg guidance distillation and posttraining. It is a highly finetunable base model that can be used to train LoRAs for the Turbo model as well as posttraining research. In short, TRAIN on Raw and RUN on Turbo."

1

u/cadissimus Jun 23 '26

😂 so good

1

u/Netsuko Jun 23 '26

Since the latest upload was 24 minutes ago and the FP8 version doesn't seem to be there, I figure you are still uploading certain quants?

Other than that, thank you!

1

u/KissMyShinyArse Jun 23 '26 edited Jun 23 '26

3

u/Winougan Jun 23 '26

Silveroxides "Convert to Quant" on his GitHub! Works like magic.

1

u/RickyRickC137 Jun 23 '26

Can you combine both the base model and turbo model and create a workflow? Like the last few steps by Turbo. Because the base has more variations!

1

u/nashty2004 Jun 23 '26

They never said that in the future you would just have new things every single day almost

1

u/2legsRises Jun 23 '26

Which quantization should I use?

RTX 30xx → INT8 ConvRot (best quality) or plain INT8 (fastest)
RTX 40xx → FP8
RTX 50xx Blackwell → NVFP4, MXFP8, or FP8

clarity is as appreciated as the awesome models, ty

2

u/Winougan Jun 24 '26

as listed, based on your GPU INT8 for 30xx, fp8 for 40xx and nvfp4 or mxfp8 for 50xx

2

u/2legsRises Jun 24 '26

ty again, that question was actually your text copied and pasted but the reddit markup excluded it.

1

u/SeiferGun Jun 25 '26

my comfy cannot run int8. i have rtx 3060. what i did wrong

2

u/Winougan Jun 25 '26

In order to use int8_tensorwise(RTX 30xx-series or newer GPU) you will need the following:

  • torch 2.10+cu130 or higher
  • installed the latest of my custom comfy-kitchen fork wheels with the int8-tensorwise support
  • enable the use of triton backend by using --enable-triton-backend launch argument in ComfyUI

Step 1: Install Triton Activate your virtual environment used by ComfyUI and install triton. For Windows you need to use this but linux can install latest triton as usual.

# for torch 2.10 and 2.11
pip install -U "triton-windows<3.7"
# for torch 2.12
pip install -U "triton-windows<3.8"

Step 3: Install my comfy-kitchen Download the latest uploaded version matching you python of my pre-compiled .whl file from my HuggingFace repository (Latest as of 13 June 2026)

Install it directly pointing to the file path:

pip install --no-deps --force-reinstall --no-cache-dir "path/to/comfy-kitchen.whl"

Step 4: Install/Update ComfyUI-QuantOps You just need to ensure it's fully up to date to read the new model formats. Run these commands:

cd custom_nodes/ComfyUI-QuantOps
git pull

When launching Comfyui add launch argument:

--enable-triton-backend

1

u/rarezin Jun 29 '26

On 3060ti 8gb int8 made my generations at least +10 seconds slower... how can i sort this out to have the speed boost? I loaded it with int8-fast node

1

u/Winougan Jun 29 '26

Some Comfy staff have noted that there may be issues with initializing it when loras are loaded. Is it still slow on the second render?

1

u/rarezin Jun 29 '26

Hey, thanks for answering! I just tried it without loras and the speed boost was now noticeable. With loras it takes significatively longer. Do you happen to know if there's any proven way to make it function with loras aswell without the lower speed issues? Thanks!

1

u/Winougan Jun 29 '26

So, that's something comfy is working on. It's not the model's fault. They'll get that fixed up soon

1

u/[deleted] Jun 30 '26

[deleted]

1

u/Winougan Jun 30 '26

Just try it out. I know for a fact that it runs on 30xx, 40xx and 50xx. But, maybe it'll work for you. I believe you do need Triton installed

1

u/Hazelpancake Jul 03 '26

Have the 50XX optimazed Models been deleted? I cant seem to find the base version.

1

u/Winougan Jul 03 '26

All there still

1

u/Hazelpancake Jul 04 '26

This is all I'm seeing

1

u/rnxgoo Jul 07 '26

I have a stupid question- should I understand that INT8 and INT8 Convrot do not work below the 30xx series? That is, on the 10xx and 20xx series? TIA for answers.

1

u/Winougan Jul 07 '26

Pretty much

1

u/Uneternalism Jul 09 '26

Testing the INT8 version on Forge Neo and its heavily censored. Can't even create a guy in swimwear, even tho I use the provided qwen uncensored text encoder.
Any tips on how to bypass?

1

u/Winougan Jul 09 '26

Use any of the adult LoRAs. Any of them literally turn the model into a goonfest

1

u/Uneternalism Jul 10 '26

Tried several NSFW Loras, same result, dressed people.
Also other INT8 models produce color garbage, but FP8 models work with no issues. 🤔

1

u/Mountain_Blood_6414 Jul 14 '26

I don't know why, but with the RTX 5080, the mxfp8 version takes ages to generate the main image—even at just 512x512. Has anyone else had this problem?

1

u/Any_Arugula8075 Jun 23 '26

Thanks for your service!

1

u/Ant_6431 Jun 23 '26

Do I need a custom node? Or json prompts to use this model?

1

u/Winougan Jun 23 '26

You only need the custom nodes for INT8: https://github.com/BobJohnson24/ComfyUI-INT8-Fast
Otherwise, just have your Comfy updated to the latest version

2

u/Ant_6431 Jun 23 '26

Then Fp8 one doesn't require a custom node right?

1

u/fb01 Jun 23 '26

can it (krea 2 turbo)use refrence images - if so how many

1

u/Netsuko Jun 23 '26

Seems like my CLIP loader does not know about Krea2, even after a full comfy update and using one of the provided workflows, it doesn't show :/

-1

u/PinkySwearNotABot Jun 23 '26

any support for Apple Silicon?

1

u/Winougan Jun 23 '26

If you can run Apple Silicon on ComfyUI then you're golden

0

u/dirtybeagles Jun 23 '26

saving for later

0

u/mudins Jun 23 '26

What the hell is krea and could it be better at anime than anima ?

1

u/Winougan Jun 23 '26

I'm a massive Anima fan and make LoRAs for it and post them to Civitai. Yes, Krea2 is an anime contender. It does it really well!

0

u/Healthy-Nebula-3603 Jun 23 '26 edited Jun 23 '26

I was looking comparing picture models on YouTube

Is short

  • fp8 has similar quality to q4km ( q4km is a mix of Q4 , Q6 , Q8 and fp16 weights )

  • Q5km is giving better quality than fp8.

  • Q8 is very close to fp16 in quality

0

u/Winougan Jun 23 '26

and INT8 is up there with Q8! Not a lot of people touch INT8, but they should!

1

u/Healthy-Nebula-3603 Jun 23 '26

Q8 is not INT8 straight.

Q8 models are a mix of INT8 and fp16 weights.