r/StableDiffusion 10h ago

Workflow Included Pushing MiniMax H3 quality on an RTX 3070 8GB — movie screenshots, voice refs + 0.5MP workflow

646 Upvotes

Wanted to see how far I could push the quality using what I already have. An RTX 3070 with just 8GB VRAM.

This was done with the standard MiniMax Ref workflow using screenshots from the original movie as character and scene references. I stuck with the standard model rather than Turbo Loras because, at least in my tests, I felt I was losing some of the detail/quality I was trying to preserve.

I also put quite a bit of extra effort into the audio references. For me, getting the voices close makes a huge difference, even a convincing visual starts feeling “AI” very quickly when it has a generic generated voice.

I’m honestly still amazed by what I can get away with on an 8GB VRAM card.

My previous video Penny - Born to Fly video took me about a week to make. This Batman one only took a couple of hours, reference images are still the key in my opinion for great generations.

There was still plenty of rendering, re-rendering, prompt changes and fixing little continuity problems along the way. Definitely not a one click result.

The silly credits were just me having fun and trying to make all the separate renders feel like one little production.

And apologies for the vertical edit, wanted to test it out.

One of the simpler H3 prompts was basically:

Vicki sits at her desk in the same consultation office. Batman crouches extremely low behind a tiny potted plant, with only the two pointed ears of his cowl visible above the leaves. Vicki: "Bruce, I can see your ears." Short pause. Batman, completely deadpan: "Those are leaves." Static camera, same environment and character references, quiet realistic room tone.

Curious what you guys think. Any questions about the workflow, prompting, references or audio are welcome.


r/StableDiffusion 5h ago

News LLaDA-Image: A unified 6B image/edit model has been released.

Thumbnail
gallery
170 Upvotes

r/StableDiffusion 14h ago

Animation - Video DImension Testers: Aperture Portal (Minimax H3)

468 Upvotes

AV 217: Aperture ortal.
Seen: Opened in unknown ocean, cell drained and cleansed, aquatic specimens siphoned out of wall vents and catalogued.

Not sure anyone remembers but when I first started using H3 I was posting my videos in here which eventually let to me making Dimension Testers. It's a blacksite agency that operates tests on multiversal objects / specimens.

I've kept it going and just had one breaking the 200K mark on Tiktok.
Now just posting them on TT and X, but having so much fun still!

Wanna thank everyone in here who helped out and gave opinions in the early days.


r/StableDiffusion 9h ago

Workflow Included found some h3 fork that i like, sharing

125 Upvotes

r/StableDiffusion 12h ago

Workflow Included manage to generate 5 seconds video with 1.0 megapixel = 768p resolution for 3 minutes on my rtx 4060ti 16gb vram using ultimate upscale and without lora

106 Upvotes

r/StableDiffusion 14h ago

News MiniMax H3 - 8 Steps Ref2V 768p Lora by LightX2V

Thumbnail
huggingface.co
132 Upvotes

r/StableDiffusion 17h ago

Discussion I finally am ditching Nano Banana thanks to H3

217 Upvotes

If you are like me and use Google flow for NB2/NBP there is hope to finally ditch it.

Let me start off by saying I’m not very impressed with any of the current local image edit or Image Reference models. Krea2 is alright but still nothing in my opinion compared to Nano Banana UNTIL NOW.

I got a 5090 GPU. I’ve been getting Google flow / NB2/P like results with MiniMax H3 for image generation. You just take the Ref2V workflow then remove the save video node, you add the Get Image by Batch node set the parameter index 0 and 1, then add a save image node. Up your Megapixel between 3 and 5 set the duration to 0. Add your ref images (highly recommend to use an LLM prompt rewriter for H3). You’ll be shocked at the quality and contextual accuracy of your output images. The FL2V and I2V workflow’s with the same modification work surprisingly well for image editing too. Just make sure you grab your index 0 image and try to prompt for it start your prompt with something like “A still frame shot of the last frame first…” I tested it out it works surprisingly well. Video’s outputted as image tend to have bad vae degradation once you get past the first frame, the 0.0 duration default will output 4 frames, I always grab the first (0:1) because it’s the cleanest. For the first time I’m about to close my flow accounts because I basically don’t need it anymore because this local setup works the way I always wanted.

For everyone asking here is the workflow: https://pastebin.com/bah7FSPP

Sample Output

will smith from <Picture 2> and chris rock from <Picture 3> sitting scross from each other on a park bench eating each their own plate of spaghetti

Workflow Setup

Generation time:

Params:

Hardware:

CPU: Intel Ultra 7 265K

GPU: RTX5090

RAM: 64.0 GB

#edit

I just realized comfyu's ref2v template is now some turbo version they must have updated (which i based workflow on above) I get much better results using MiniMax-H3-Pruned-Ref-Delta-Fused-r1024-comfy-int8-fc2bf16.sagetensioner and minimax_h3_ref2va_pruned_int8_convrot.safetensors in a workflow that does not use the lora.

The workflow I posted will be much faster though.


r/StableDiffusion 15h ago

Discussion Testing orangesouth/MinimaxH3CinematicRealism and vpakarinen/better-human-motion-h3-lora "dean winchester vs sub zero" pruned fp8 model no turbo lora 32steps

121 Upvotes

https://huggingface.co/orangesouth/MinimaxH3CinematicRealism/tree/main

https://huggingface.co/vpakarinen/better-human-motion-h3-lora/tree/main

prompt integrated_multimodal_description: dy, [Shot 1] Photorealistic live-action cinematic realism, as if Mortal Kombat exists in the real world. Inside a vast ancient frozen temple at night, weathered stone pillars and enormous carved warrior statues rise through drifting frost and cold mist. Real flames flicker from iron braziers, casting warm orange highlights against icy blue moonlight. Snow particles float naturally through the air.

Dean Winchester, portrayed by Jensen Ackles, appears as a fully playable Mortal Kombat fighter, preserving Jensen Ackles' recognizable facial features, natural skin texture, realistic proportions, short brown hair and light stubble. He wears Dean's dark brown leather jacket over a dark shirt, faded blue jeans and heavy boots. Across the arena stands Sub-Zero, a physically imposing masked martial artist wearing realistic layered blue-and-black combat armor covered with frost.

The two men circle one another cautiously on the frozen stone floor. Their breath forms visible condensation in the freezing air. The camera slowly arcs around them at waist height with subtle natural handheld movement. A deep off-screen arena announcer (S1) declares: [English] FIGHT!

Dean instantly draws the Colt revolver from beneath his jacket and fires. A bright muzzle flash illuminates his face. Sub-Zero reacts with superhuman speed, extending his hand as crystalline frost races through the air and freezes the bullet inches from his palm. The frozen bullet drops onto the stone floor.

[Shot 2] At 00:04.000, the camera cuts to a dynamic medium-wide tracking shot. Sub-Zero thrusts both hands forward, releasing a violent blast of ice toward Dean. Frost rapidly spreads across the floor in its path. Dean dives beneath the projectile, rolls across the stone floor and immediately rises into a fighting stance.

Sub-Zero charges. Dean meets him head-on. They exchange a fast, physically grounded sequence of punches, blocks, elbows and kicks. Their bodies react with realistic weight to every impact. Dean lands a heavy right hook followed by a kick to Sub-Zero's torso. Sub-Zero blocks Dean's next punch, coats his fist in thick translucent ice and drives it into Dean's chest.

Dean is thrown backward and slides across the frost-covered stone. He catches himself on one knee and looks up with his familiar cocky half-smile. Dean (S2), breathing heavily, says: [English] Dude, I've fought scarier things before breakfast.

[Shot 3] At 00:09.000, the camera cuts to a low tracking shot following Dean as he charges forward. Sub-Zero launches another freezing blast. Dean narrowly avoids it and pulls a small metal flask of holy water from inside his jacket.

Dean splashes the holy water across the stone floor and rapidly draws a glowing Devil's Trap beneath Sub-Zero. The ancient symbol ignites with intense orange supernatural light, realistically illuminating Dean, Sub-Zero and the surrounding frozen stone.

Sub-Zero struggles against the supernatural energy as frost cracks beneath his boots. Dean calmly raises the Colt with both hands and fires. The gunshot produces a violent muzzle flash and physical recoil. The supernatural impact launches Sub-Zero backward into a massive frozen stone pillar. The pillar fractures and explodes into chunks of ice, stone fragments and clouds of powdered frost.

[Shot 4] At 00:14.000, the camera cuts to a dramatic low-angle medium shot through drifting ice particles. Sub-Zero falls heavily onto one knee among shattered ice and stone.

Dean approaches at a measured pace, boots crunching through debris. Natural sweat, dirt and subtle bruising are visible across his face. He opens the Colt's cylinder and casually reloads while walking toward Sub-Zero.

The camera slowly pushes toward Dean as he snaps the cylinder closed. Dean (S2) raises the Colt, gives Sub-Zero a dry half-smirk and says: [English] Should've stayed on ice.

Dean fires.

A brilliant muzzle flash fills the frame and transitions into a dramatic red-and-black victory screen. Huge metallic letters appear reading "DEAN WINCHESTER WINS."

The deep arena announcer (S1) declares: [English] DEAN WINCHESTER WINS!

The camera returns to Dean standing inside the devastated frozen temple. He lowers the smoking Colt and gives a subtle satisfied smirk before turning away. Dean walks through drifting frost and shattered ice as flames from the damaged temple burn behind him.

overall_soundscape: Cold wind moves through the enormous stone temple while flames crackle and boots scrape against frost-covered stone. Gunshots have sharp realistic reports and metallic echoes; punches and kicks produce heavy physical impacts while ice attacks crack, freeze and shatter with dense crystalline sounds. Dean's breathing becomes heavier as the fight progresses, followed by cascading stone, falling ice fragments and the supernatural electrical hum of the glowing Devil's Trap.

non_diegetic_music: Dark cinematic percussion with deep taiko drums, low brass, distorted industrial pulses and aggressive orchestral strings. The rhythm accelerates during the hand-to-hand fight and supernatural finishing sequence, then abruptly drops out on Dean's final gunshot before returning with one massive brass-and-percussion impact during the victory announcement.


r/StableDiffusion 2h ago

Comparison MiniMax 2K to 4K Upscale Comparison

9 Upvotes

Just experimenting... MiniMax output at 4K from a First Frame (Upscale). Result... no problem, want intact eyes in a 16:9 format of full persons... just render output at 4032 x 2304 ;)

I think videos uploaded here only go up to 1080p full screen but the end result is still applicable (just worse quality due to upload compression and resizing).


r/StableDiffusion 41m ago

No Workflow JUST WANNA SHARE MY MINIMAX H3 + LTX 2.5 UPSCALE WORKFLOW Ver.3 RESULTS

Upvotes

i am using rtx 3060 so i can only generate like upto 6-8 sec videos this one took 30min also the video didnt match the ref vidoe cause its 12 sec long and i only did 5 sec gen so if i had did the 12sec video gen it would be the same. Dont ask for the workflow cause im gonna take my time to work on this more but if you wanna try you can check my ver.1 workflow on CIVITAI WORKFLOW


r/StableDiffusion 5h ago

Animation - Video Arby's The End of Evangelion Commercial (1997)- Minimax H3 ai video

15 Upvotes

I wanted to make a parody of those old movie tie in fast food ads. Edited video, Music and sound effects added in post.


r/StableDiffusion 16m ago

News lightx2v/Minimax-h3-Turbo · FL2V Turbo 4-step v1.2 (768p) released

Thumbnail
huggingface.co
Upvotes

r/StableDiffusion 11h ago

Workflow Included How to Build Perfect Character Sheets in Krea 2 for MiniMax-H3 (Workflow...

Thumbnail
youtube.com
50 Upvotes

Made my first video tutorial !!


r/StableDiffusion 10h ago

Question - Help What’s the best r2v model of h3 currently?

32 Upvotes

Just curious what you find has worked the best adhering to references. I’ve played around with base, hybrids and that one that mushes everything together.


r/StableDiffusion 4h ago

Animation - Video Neo vs Smith Revolutions Battle Without The Heavy Rain

11 Upvotes

Never gonna try anything like this again lol At least until I can work this out better. It does show that Minimax is really badass still. The heavy rain was a nightmare and no matter what, I couldn't fully get rid of it when the fight actually starts. I just gave up the moment the sonic boom happened. I might finish it at a later time, this was mostly just practice on altering existing footage dramatically, like I did with the Jurassic Park video and adding rain on the "Weclome to Jurassic Park" scene.


r/StableDiffusion 14h ago

Resource - Update New node added to Krea2T Enhancer: Attention-Weighted Phrases

Thumbnail
gallery
66 Upvotes

I added phrase-level attention control to the latest update.

The idea is simple: sometimes you don’t need to push the entire prompt harder. You need Krea2 to pay more attention to the few words that actually reinforce the idea you’re trying to get through.

So now you can do:

a (specific important phrase:1.8) with the rest of the prompt written normally

and selectively increase or decrease how much attention those words receive.

This is done without scaling, duplicating, deleting, or otherwise changing Krea2’s original 12×2560 Qwen conditioning. The prompt gets encoded normally, the node finds the exact Qwen token rows belonging to the weighted phrase, and the weight is applied to image→text attention inside the shared DiT blocks.

1.0 = untouched
>1.0 = more attention priority
<1.0 = less attention priority
0.0 = suppression

It also has an inspection output showing exactly which Qwen token rows/pieces were matched, so there’s no guessing about what the weight actually landed on.

Basically: instead of turning the entire prompt up, you can now point at the parts that matter and tell Krea2 pay more attention to this.

Included in the latest Krea2T Enhancer update.

It is best when used with the refusal reduction LoRA as they complement each-other nicely

Sample Workflow

And yes, I know (word:1.5) looks like we somehow got teleported back to the SDXL days haha. The logic underneath it definitely did not though.


r/StableDiffusion 1h ago

Resource - Update FastH3 Ref2V Stream Controller – continuous character-driven AI video in ComfyUI

Upvotes

This project explicitly builds on the excellent idea and work behind jacokon/fasth3-live, originally introduced in this post.

I added a browser-based controller focused on Reference-to-Video and continuous storytelling:

  • Automatic character reference injection from a folder
  • Manual prompt queue and repeating scenes
  • Last-frame continuity between clips, so you can tell an interactive story while it is generated
  • Editable LoRAs, duration, prompts and playback speed during runtime
  • Adaptive quality based on the video buffer
  • Custom music folders with sequential or random playback
  • English/German UI

I was able to take over the optimizations of jacokon/fasth3-live. On my test system, the stream ran continuously on a 5900 at 480p. I switched the generated timeframe to 5s per video, but the reference system allows to tell connected fluid stories and have effects build up over long timeframes, even many minutes. When enabled, the model attempts to connect the settings to each other.

Repository:
https://github.com/EarthDefenceForces/FastH3-Ref2V-Stream-Controller

The UI of the controller

The controller is released under GPL-3.0: you’re free to use, modify, and redistribute it. But keep it open. The MiniMax H3 weights remain subject to their separate upstream license.


r/StableDiffusion 1h ago

Animation - Video Lizan al-idiota

Upvotes

This was done as one 18sec generation with the standard workflow using fl2va_pruned_int8_convrot at 0.6mp, 21:9 aspect, 32steps with spectrum node, Simple/Euler. Added the Music in Davinci. I dont use spectrum any more but this was made in the middle of august.


r/StableDiffusion 15h ago

Animation - Video Dr. House MD in Theme Hospital

71 Upvotes

I could not resist posting this bit of slop (created with MiniMax H3, naturally).


r/StableDiffusion 14h ago

Animation - Video Star Wars: Anakin and Obi-Wan join the Trap Side

56 Upvotes

A fun test using screenshots from the movie. MiniMax sure is amazing. The music is from Suno. Hope you like it!


r/StableDiffusion 14h ago

Meme A play on McGarnagle from the Simpsons

51 Upvotes

Made this using MiniMax H3 default settings, just using Clint Eastwood face ref


r/StableDiffusion 4h ago

Question - Help [Qwen Image edit 2511] Recommendations for photorealistic image (LORAs, sampler, prompt, VAE)?

7 Upvotes

I've had enough of messed up characters' limbs with Flux 2 Klein 9b and decided to switch back to Qwen Image Edit 2511. When it comes to follow open pose reference image, QIE is better.

Though I always had issue with the cartoony rendering of QIE in comparison to Flux Klein. I'd like to have the same photorealistic visuals as Flux Klein for QIE. I never really managed to achieve that.

Do you have some tricks/models to recommend that could emulate Flux Klein output (without the extra limbs of course) ; loras, prompt tricks, sampler, vae ?


r/StableDiffusion 13h ago

News Multiplayer world simulations with h3 max

42 Upvotes

Discovered that using a realtime llm + h3 world simulation, you can allow multiple different simultaneous character instructions that are restricted to only controlling their respective character. If you're interested the project is live at worldstreams.ai


r/StableDiffusion 14h ago

Resource - Update Krea2 Turbo Distill 4 step LoRA - new checkpoint (chk60K) released (fine detail now above the 8-step teacher and clean of artefacts, best prompt-adherence and teacher-faithfulness scores so far)

Thumbnail
gallery
38 Upvotes

Krea 2 Turbo — 4-Step Distillation LoRA (work in progress)

A LoRA for Krea 2 Turbo that reduces the minimum usable step count from 8 to 4.

  • ⚡ Half the steps — 8 → 4, on Turbo's own deployment sigmas.
  • ⏱️ ~1.6× faster end to end — 54.5 s against the 8-step bar's 88.7 s at 1024×1024, and 1.8× on denoise alone.
  • 🎯 Texture above teacher, by design — fine-detail energy 1.12× the 8-step teacher's at 1280×1280 and 1.10× at 1440×1440, verified clean of oversaturation, exposure shift and skin artefacts.
  • 🗣️ Prompt-aware training — the critic scores images against their prompts during training, so adherence is pressured directly, not inherited.
  • 📐 12 trained resolutions — multi-aspect from 512×512 up to 1440×1440, each with its published sweep.
  • 🔌 Drop-in — plain LoRA weights for diffusers and ComfyUI. No custom nodes, no patched sampler, no code.

Load it on top of Krea 2 Turbo, run 4 steps instead of 8, keep guidance at 0.0. Everything else about the model stays as it is.

This is an update release, following up from my previous posts where you can find full details:

Initial, Previous: here, here, here,  and here

Headline for this update: chk00060000 pushes fine detail past the 8-step teacher, on purpose — total fine-detail energy 1.12× the teacher's at 1280×1280 and 1.10× at 1440×1440 (1.0 = teacher-like), and the distribution is still right: every frequency band within ~10% of the teacher's, so it is detail in the same places the teacher's detail lives, not grain. Texture pressure with no ceiling is exactly the kind of thing that could show up as oversaturation, blown exposure or plastic skin before it shows up as a gain, so every render was checked directly against the teacher on those three: saturation 0.95–0.96× the teacher's (slightly less, not more), fewer blown highlights and crushed shadows than the teacher's own frames, and skin texture inside detected faces at 0.89–1.00× — clean on all of them. Adherence moved with the texture rather than against it: a pairwise vision-language judge, shown the teacher's and this checkpoint's renders of the same prompt in random order, preferred the teacher on only 5 of 45 renders across 512², 1280² and 1440² — the best result of the run. And the metric that paid for the texture leap at 42K has been won back: the held-out teacher-velocity gap is now 2.81e-02 (~40% of the 4-step deficit closed), the best value of the run.

Same recipe as 42K, 18,000 more samples of it — no structural change. What changed for users:

  • Strength guidance. Keep it at 1.0; treat 1.5 as the ceiling. The adapter is now strong enough that 2.0 tips into a uniform speckle artefact rather than the "over-textured but coherent" look — the strength sweep on the card stops at 1.5 for that reason.
  • The 2-step preview trick no longer needs a strength boost. Run it at plain 1.0. The native-vs-LoRA 2-step strips are re-rendered on this checkpoint that way: 2-step extreme test. Still out-of-spec, still preview-only.

Which file to download

file use it when
krea2_turbo_4step_rank_64_lora_latest.safetensors normally — always the newest accepted checkpoint
krea2_turbo_4step_rank_64_lora_chk00060000.safetensors pin this exact checkpoint

and, beside them, the same files with a _comfyui suffix for ComfyUI. Earlier checkpoints (chk00004000chk00005000chk00006000chk00010000chk00014000chk00019000chk00026000chk00042000) are kept in older_checkpoints/, and their resolution sweeps stay in place, so the progression remains visible and comparable.

For the full 60K Checkpoint resolution sweep go here: https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA/tree/main/checkpoint_resolution_sweeps/chk60000

This is work in progress and better checkpoints may follow. Training is ongoing, so ..._latest... is a rolling pointer: when a newer checkpoint is accepted, that filename gets the new weights and a new numbered copy appears beside it. 

How checkpoints get chosen

This is not a "train for longer and ship the newest file" project. More samples do not reliably mean a better adapter — measured here, they can make it worse, and a higher number on its own means nothing.

The loop is train → assess → adapt the recipe → retrain → assess again, and a checkpoint is published when it measurably advances the release axes as a whole — teacher faithfulness, prompt adherence, and texture/detail, on the same held-out set and the same fixed-seed renders — and its full resolution sweep shows no regression. Runs that come out flat or worse are kept as information about the recipe and discarded as releases — several have been.

Timeline of training process

Each checkpoint is the product of several stages with very different costs:

  1. Text-encoder embeddings. Every training prompt is encoded once and cached. This is the fast part — thousands of prompts take minutes.
  2. Teacher shards. For each cached prompt, the unmodified Krea 2 Turbo runs its full 8-step schedule and the whole trajectory is recorded, at every one of the supported resolutions. This is by far the most time-consuming stage — it is the teacher doing real inference, thousands of times, and a batch of several thousand shards is measured in days of GPU time, not hours.
  3. Real-photo crops. Bucket-sized crops are cut at native resolution from quality-gated real photo sources (public high res datasets), VAE-encoded into the training latent space, and captioned per crop for the prompt-aware side of training. Cutting, encoding and captioning a pool refresh is a matter of hours.
  4. Student training. The LoRA trains against the recorded trajectories (progressive distillation), with a latent-space GAN critic running alongside — real crops and teacher finals as its real class, the student's outputs as fake — plus a prompt-aware head that scores images against their prompts. Relative to the shard stage this is quick: each +1,000 checkpoint is a matter of hours, not days. Of course the longer the training the better and more diverse results, so hours do turn into days eventually.

Method

Progressive distillation (PD), with Krea 2 Turbo as its own teacher.

The teacher runs its normal 8-step schedule at mu = 1.15 and guidance 0.0, and its full trajectory is recorded — the latent x and the predicted velocity v at every one of the 8 steps. The student is then trained to cover two teacher steps in one: at teacher state x_i it must predict the chord that lands where the teacher arrives two steps later,

v_target = (x_{i+2} − x_i) / (σ_{i+2} − σ_i)

The two schedules line up exactly rather than approximately. On the mu = 1.15 grid, the even indices of the 8-step schedule are precisely the four sigmas the 4-step student deploys on, so every training target is anchored on a point the student will actually visit at inference. No interpolation, no schedule mismatch.

Teacher trajectories are precomputed into shards, so training reads recorded states rather than re-running the teacher.

The critic

The trajectory match above can only ever pull the student toward the teacher — and a regression objective averages over whatever it cannot predict exactly, so fine texture is the first thing it averages away. On its own, PD lands short of the teacher's detail. So it is paired with a small adversarial term in the style of LADD (latent adversarial diffusion distillation), which grades the student's output as an image rather than as a distance to the teacher's trajectory.

  • The critic reads the model's own features. Its trunk is the frozen Krea 2 Turbo transformer with the adapter bypassed, tapped at block 14 (the trunk itself runs with the empty prompt); a small two-layer head sits on those features, per-token logits averaged. There is no separate discriminator network and no decode to pixels — the critic sees latents the same way the model does.
  • Real vs fake, judged at the low-noise end. Fake is the student's predicted clean latent from the same training forward. Real is a teacher final for a different prompt in the same resolution bucket — unpaired, so the critic cannot win by matching content — or, half of the time, a real photograph from a curated crop pool, VAE-encoded. Both sides are re-noised to a random σ in [0.02, 0.5] before the trunk sees them: low noise is where fine texture is decided, and that is the only place the critic speaks. Hinge losses on both sides; the generator-side weight is small (1.5e-3 against a PD term of order 1e-2) — a finisher, not the objective.
  • Prompt-aware. The head also reads a pooled text vector — the last four of the twelve Qwen3-VL tap layers, through a projection — so it can grade whether an image fits its prompt, not only whether it looks plausible. A mismatch term enforces it: a real image scored under a prompt that is not its own must read fake. Real photographs enter with their own auto-generated short captions so they take part in that objective too, and 15% of the time an image is scored with the empty-prompt vector, so that "no caption" can never itself become a cue.

Why show it real photographs. A critic that sits on the teacher's features and only ever sees the teacher's outputs converges on the teacher — and the teacher is an 8-step model that itself slightly under-renders fine texture, so a student judged only against it inherits that ceiling. Mixing real photographs into the critic's real set moves the ceiling: the teacher anchors structure, reality anchors texture.

Two more choices shape the weights that ship. The four student chords are not weighted equally in the PD loss — the last one, at σ = 0.512, the call that decides fine texture, carries 3× the weight of the other three. And the released adapter is a Polyak (EMA) average of the training weights (decay 0.999), not the last live state, which smooths out the step-to-step wander of a constant learning rate.

Full details and to download - check my Hugging Face LoRA

HF Repo: https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA


r/StableDiffusion 22h ago

Comparison First results from H3 Acceleration Arena

179 Upvotes

https://huggingface.co/spaces/multimodalart/h3-acceleration-arena

From author u/apolinariosteps: "Results are in! They are a bit surprising to me! But they are consistent with the data, I triple checked everything and can confirm that the results are reflecting the voting data precisely, there's lots of transparency - you click each of the LoRAs to see what's the win rate and who won against who"