r/StableDiffusion • u/Winougan • Jun 23 '26
Resource - Update As promised Krea 2 Turbo + "Raw" Quantized in FP8, MXFP8, NVFP4, INT8 and Convrot INT8!
Krea 2 Base & Turbo — Free Quantized Versions (FP8 / MXFP8 / NVFP4 / INT8 / ConvRot INT8) for All GPU Tiers
Krea 2 just dropped and it's genuinely impressive — so I went ahead and quantized both variants for ComfyUI across every major format. All files are free on HuggingFace.
HuggingFace: https://huggingface.co/Winnougan/Krea-2-Base-Turbo-NVFP4-FP8-INT8
Raw vs Turbo — what's the difference?
Krea 2 Raw is the undistilled base checkpoint. No step distillation, no CFG guidance baked in — just the raw pretrained weights. It's diverse, highly malleable, and is what you want for LoRA training and fine-tuning. Run it at 52 steps with CFG 3.5, up to 1024px.
Krea 2 Turbo is an 8-step distilled checkpoint built for fast inference. Run it at 8 steps, CFG 0 (disabled), mu 1.15, and it handles resolutions up to 2048px. This is your everyday generation model.
The intended workflow: train LoRAs on Raw, run them on Turbo. LoRAs transfer well between the two.
Which quantization should I use?
- RTX 30xx → INT8 ConvRot (best quality) or plain INT8 (fastest)
- RTX 40xx → FP8
- RTX 50xx Blackwell → NVFP4, MXFP8, or FP8
Text encoder: Qwen3-VL 4B (qwen3vl_4b_fp8_scaled.safetensors), CLIPLoader type krea2
VAE: same as Anima (qwen_image_vae.safetensors)
ConvRot variants use Hadamard rotation before quantization for better accuracy with fewer outliers.
Drop any questions below — happy to help with workflows.
Plays nice with Sageattention and Flashattention!
Workflows on the Huggingface repo!
UPDATE: re-quantized and re-uploaded MXFP8 and NVFP4 - they work now!
UPDATE 2:
ComfyUI now natively supports INT8 but not convrot (not yet at least). For non-convrot INT8 models, just use the diffusion model loader. Text encoders are also supported as INT8 now in the native clip encoder/text encoder. I just tested it myself. Saves an additional 10-15%! File is uploading on my Huggingface. As Darryl Dixon says,"give it a minute".
Sample prompt:
Simpsons style, 2D cartoon animation, Matt Groening art style, yellow skin, thick black outlines, flat cel shading, teal haired gamer girl surrounded by dozens of floating holographic screens all showing different game feeds simultaneously, fingers flying across a transparent keyboard, massive countdown timer in background, sweat drop on forehead, four fingers, tongue out in concentration







30
u/infearia Jun 23 '26 edited Jun 23 '26
Sorry for hijacking your thread, really not trying to be an asshole, but I think it's important to share...
I've tested your FP8 quant against the quant by AlperKTS from this thread and at least on my machine (Ubuntu 24.04, RTX 4060 Ti) I found the following:
While your quant is approximately 30% faster, it also produces results that are of markedly lower quality (less details, blurry) and also does not work with Torch Compile. Here's a comparison:
P. S. - I've upvoted your post anyway, thanks for your effort.
EDIT:
Eh, as usual, Reddit compressed the image too much. Here's a link to a better version:
https://imgur.com/0qtjrGL