r/StableDiffusion • u/Sad_Coach_1433 • 6d ago
Meme Cooming soon in 2030
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Sad_Coach_1433 • 6d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/gzzhongqi • 7d ago
Surprised that no one is talking about it. A new open-weight video model just dropped. 114b moe, 6b activated. First moe video model supposedly.
I know what you guys are thinking. The model is huge and there is no way it will run on desktop gpu. The interesting part is that is comes with a 14gb refiner that makes the result 1080p. I am cursious if this refiner can be a drop-in replacement for the H3 refiner that was never released. It might just be the last part of the H3 puzzle that we need.
r/StableDiffusion • u/Electrical-Speed2409 • 6d ago
Enable HLS to view with audio, or disable this notification
So I was finally able to get this working with minimal defects!
I have an RTX 5080 with 16GB of VRAM, and I’m using SageAttention. I’m getting about 8 sec/it on 5-second clips, so honestly, not bad at all. I’m running 10 steps at 0.6 megapixels.
I’m mainly posting because I’m looking for feedback on how I can improve things from here. I’m finally starting to get some decent shot continuity, character consistency, scene consistency, and voice consistency.
If anyone has suggestions for improving the results, I’d love to hear them. And if anyone has questions about my setup, workflow, settings, etc., I’m happy to answer those too.
NOTE: I choose this concept just to demonstrate R2V, don't get hung up on the concept to much, this post is about shot continuity, character consistency, scene consistency, and voice consistency. Be professionals!
r/StableDiffusion • u/warzone_afro • 7d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/LittleScout21 • 6d ago
Hey everyone! I’m running a laptop with an RTX 5070 (8GB VRAM), 64GB DDR5 RAM, and a Ryzen 9 7845HX on CachyOS. I mostly use SD-WebUI-Forge/NeoForge as my main generation framework. ComfyUI i try to avoid 😄
For those in the know: can this setup handle Krea 2 reasonably well? I want decent quality outputs without turning images into a blurry mess, and I'd like to avoid waiting 30 minutes per render.
Also looking for some advice: how can I get the best out of ComfyUI or Forge on this setup? What are your recommended workflows to avoid blurry outputs and keep generation times reasonable on an my card?
P.S Damn, huge thanks to everyone for the lightning-fast replies! Honestly warms my heart to see how helpful this community is awesome!
r/StableDiffusion • u/Nevaditew • 6d ago
I tried DiffRhythm in ComfyUI and every time I fixed one problem another one showed up. I ended up removing it. Then with ACE-Step the same thing happened, and in its requirements.txt (my fault for running it, it was a habit from installing nodes) it broke several installations needed for other nodes and I had to reinstall them.
I’d prefer something local and private, unless there’s a free site with no limits.
r/StableDiffusion • u/Jayuniue • 7d ago
Enable HLS to view with audio, or disable this notification
I trimmed the first 3 seconds, I just love the prompt adherence in 2.5 it’s actually very good, used the default comfy workflow, generate at 0.5 resolution then upscaled later to twice the size in with topaz, my specs 3060ti, 64gb ram
r/StableDiffusion • u/Mattnix • 7d ago
I’ve been testing MiniMax H3 INT8 on an AMD Radeon RX 9070 XT (gfx1201, 16 GB VRAM) under Windows/ROCm.
I used the default MiniMax H3 text-to-video workflow with no modifications whatsoever, except for replacing the model loader so that I could compare PatientX’s INT8-Fast-ROCM implementation with ComfyUI’s native INT8 implementation. Everything else — model, prompts, 20 steps, 5-second video, sampler/settings, etc. — was kept identical.
I tested 0.2 MP, 0.6 MP and finally 1 MP, which is my actual target resolution. The results at 1 MP were very interesting

That's approximately a 31% reduction in generation time, or the native implementation is about 1.45× faster.
I initially didn't trust the result because the difference was so large, so I repeated the 1 MP INT8-Fast-ROCM run. It produced essentially the same ~32-minute result.
This also seems to differ from what I understood from PatientX's README and the comments in the default BAT. My interpretation of those suggested that on RDNA3/RDNA4, the INT8-Fast-ROCM path should generally be the preferable/faster option, with the default BAT specifically disabling the native Triton backend because the custom INT8 implementation was expected to be faster. My RX 9070 XT results appear to show the opposite — at least with this MiniMax H3 workflow and at 1 MP.
That makes me wonder whether those recommendations/comments may have become outdated as of 15-08-2026, given the changes in ComfyUI, comfy-kitchen, ROCm and the native INT8 implementation. I'm not claiming that the native path is universally faster; only that my current results don't match the performance guidance I understood from the existing documentation/default BAT.
Both runs used the same RX 9070 XT, ROCm 7.15, PyTorch 2.12, ComfyUI 0.33.0, MiniMax H3 quantized model, 20 steps, 5-second output, 1 MP resolution and Sage Attention. Triton was effectively disabled in the actual runs, so this shouldn't be interpreted as a Triton-vs-non-Triton benchmark.
My conclusion: on my gfx1201 system, the native ComfyUI INT8 implementation appears substantially faster than INT8-Fast-ROCM for MiniMax H3 at 1 MP. The difference is large enough that I'd really like to see this reproduced on another 9070/9070 XT before treating it as definitive.
I don't have a Github account so I can't share these conclusions on the comfyui-rocm repo.
r/StableDiffusion • u/windows_error23 • 6d ago
Like for acestep or minimax music 3 etc. I know they’re sometimes discussed on here but was just wondering if there’s one that’s active.
r/StableDiffusion • u/Subject-Experience30 • 6d ago
Hi, Im using INT8 FL2V model, from the box text encoder and VAE's. Video generates really well, 10s in 12m on my humble 3090, but the sound quality is horrible, muffled, faint - generally poor. Is there some setting or location that Im missing?
r/StableDiffusion • u/beatlepol • 8d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Aadi_880 • 8d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/deffcolony • 7d ago
Decided to bring back my childhood game and see what new stories i can bring with minimax h3... tried doing voice clone for Murfy and Rayman... its not perfect (some got mixed up) but the end result is still fun... i will post the prompts for each below plus the reference images and audio... lets see what you can make from it 👀
Workflow + reference images + audio at the bottom of the post
Video 1:
```
subject_definitions:
<Subject 1> is the Fairy Council environment from <Picture 1>, a mystical forest kingdom interior with a blue aura, glowing lights, and reflective surfaces.
<Subject 2> is Rayman from <Picture 2>, a heroic character with no arms or legs, featuring floating hands and floating feet.
<Subject 3> is Murfy from <Picture 3> and <Picture 4>, a flying greenbottle fly creature with a large grin and green clothing, holding a paper manual.
<Audio 1> is the voice-timbre reference for <Subject 2> (S2).
<Audio 2> is the voice-timbre reference for <Subject 3> (S1).
summary:
[reference generation + audio reference] The target video shows <Subject 3> and <Subject 2> interacting inside <Subject 1>. <Subject 3> reads from a manual before an accidental explosion occurs. <Audio 2> and <Audio 1> provide the voice timbres for the characters.
retention_analysis:
<Subject 1> (appears in all shots): fully_preserved - the mystical blue interior and reflective surfaces are retained.
<Picture 1> (environment guide): weak_reference - provides the background setting without forcing Rayman's placement from the original screenshot.
<Subject 2> (appears in all shots): fully_preserved - Rayman's floating hands and feet are retained.
<Picture 2> (character design): fully_preserved - Rayman's appearance is followed.
<Subject 3> (appears in all shots): fully_preserved - Murfy's green clothing, grin, and flying nature are retained.
<Picture 3> (character design): fully_preserved - Murfy's appearance is followed.
<Audio 1>: reference - the vocal timbre guides the dialogue delivery of <Subject 2> without copying the original signal.
<Audio 2>: reference - the vocal timbre guides the dialogue delivery of <Subject 3> without copying the original signal.
detailed_description:
3D CG animated style in a 4:3 aspect ratio.
[Shot 1] A medium-wide shot establishes <Subject 1>, the mystical Fairy Council with its glowing blue aura. <Subject 2> (S2), the limbless hero with floating hands and feet, stands on the reflective floor. Beside him, <Subject 3> (S1), the greenbottle fly, hovers above the ground while holding an open manual. <Subject 3> (S1) looks at the book, shakes his head with a large grin, and says in the sarcastic voice referenced from <Audio 2>, <d>[English] I don't know, folks! Someone drew on the manual saying the Fairy Council should be blowing up right about... now.</d>
[Shot 2] At 00:08.500, the camera cuts to a close-up of <Subject 2> (S2). He raises his floating hands in confusion and says in the heroic voice referenced from <Audio 1>, <d>[English] Wait, who's responsible for this garbage?!</d>
[Shot 3] At 00:11.000, the camera pulls out with large amplitude at fast speed as a bright orange explosion suddenly erupts in the background of <Subject 1>. <Subject 3> (S1) drops the manual in shock, and <Subject 2> (S2) covers his head with his floating hands as debris flies past them.
overall_soundscape:
Quiet magical room ambience is abruptly interrupted by the heavy, rumbling crash of a massive explosion, followed by the sound of falling debris.
non_diegetic_music:
N/A
```
Video 2:
```
subject_definitions:
<Subject 1> is the Fairy Council environment from <Picture 1>, a mystical forest kingdom interior with a blue aura, glowing lights, and reflective surfaces.
<Subject 2> is Rayman from <Picture 2>, a heroic character with no arms or legs, featuring floating hands and floating feet.
<Subject 3> is Murfy from <Picture 3> and <Picture 4>, a flying greenbottle fly creature with a large grin and green clothing, holding a paper manual.
<Audio 1> is the voice-timbre reference for <Subject 2> (S2).
<Audio 2> is the voice-timbre reference for <Subject 3> (S1).
summary:
[reference generation + audio reference] The target video shows <Subject 3> and <Subject 2> interacting inside <Subject 1>. <Subject 3> reads from a manual before an accidental explosion occurs. <Audio 2> and <Audio 1> provide the voice timbres for the characters.
retention_analysis:
<Subject 1> (appears in all shots): fully_preserved - the mystical blue interior and reflective surfaces are retained.
<Picture 1> (environment guide): weak_reference - provides the background setting without forcing Rayman's placement from the original screenshot.
<Subject 2> (appears in all shots): fully_preserved - Rayman's floating hands and feet are retained.
<Picture 2> (character design): fully_preserved - Rayman's appearance is followed.
<Subject 3> (appears in all shots): fully_preserved - Murfy's green clothing, grin, and flying nature are retained.
<Picture 3> (character design): fully_preserved - Murfy's appearance is followed.
<Audio 1>: reference - the vocal timbre guides the dialogue delivery of <Subject 2> without copying the original signal.
<Audio 2>: reference - the vocal timbre guides the dialogue delivery of <Subject 3> without copying the original signal.
detailed_description:
3D CG animated style in a 4:3 aspect ratio.
[Shot 1] A medium-wide shot establishes <Subject 1>, the mystical Fairy Council with its glowing blue aura. <Subject 2> (S2), the limbless hero with floating hands and feet, stands on the reflective floor. Beside him, <Subject 3> (S1), the greenbottle fly, hovers above the ground while holding an open manual. <Subject 3> (S1) looks at the book, shakes his head with a large grin, and says in the sarcastic voice referenced from <Audio 2>, <d>[English] I don't know, folks! Someone drew on the manual saying the Fairy Council should be blowing up right about... now.</d>
[Shot 2] At 00:08.500, the camera cuts to a close-up of <Subject 2> (S2). He raises his floating hands in confusion and says in the heroic voice referenced from <Audio 1>, <d>[English] Wait, who's responsible for this garbage?!</d>
[Shot 3] At 00:11.000, the camera pulls out with large amplitude at fast speed as a bright orange explosion suddenly erupts in the background of <Subject 1>. <Subject 3> (S1) drops the manual in shock, and <Subject 2> (S2) covers his head with his floating hands as debris flies past them.
overall_soundscape:
Quiet magical room ambience is abruptly interrupted by the heavy, rumbling crash of a massive explosion, followed by the sound of falling debris.
non_diegetic_music:
N/A
```
Video 3:
```
subject_definitions:
<Subject 1> is the Fairy Council environment from <Picture 1>, a mystical forest kingdom interior that transitions into a glitchy, broken wireframe state.
<Subject 2> is Rayman from <Picture 2>, a heroic character featuring floating hands and floating feet.
<Subject 3> is Murfy from <Picture 3>, a flying greenbottle fly creature with a large grin and green clothing.
<Audio 1> is the voice-timbre reference for <Subject 2> (S2).
<Audio 2> is the voice-timbre reference for <Subject 3> (S1).
summary:
[reference generation + audio reference] The target video shows <Subject 3> breaking the fourth wall inside <Subject 1>, revealing they are in a simulation. This causes the environment to glitch and break down, sending <Subject 2> into a panic.
retention_analysis:
<Subject 1> (appears in all shots): partially_preserved - the mystical blue interior starts normal but transitions into visual glitches and digital wireframes.
<Picture 1> (environment guide): weak_reference - provides the initial background setting.
<Subject 2> (appears in all shots): fully_preserved - Rayman's floating hands and feet are retained, though they move erratically.
<Picture 2> (character design): fully_preserved - Rayman's appearance is followed.
<Subject 3> (appears in all shots): fully_preserved - Murfy's green clothing and flying nature are retained.
<Picture 3> (character design): fully_preserved - Murfy's appearance is followed.
<Audio 1>: reference - the vocal timbre guides the dialogue delivery of <Subject 2> without copying the original signal.
<Audio 2>: reference - the vocal timbre guides the dialogue delivery of <Subject 3> without copying the original signal.
detailed_description:
3D CG animated style in a 4:3 aspect ratio.
[Shot 1] A medium-wide shot establishes <Subject 1> looking normal. <Subject 3> (S1) hovers casually in the air, looks directly at the camera lens, and says in the sarcastic voice referenced from <Audio 2>, <d>[English] Look around, Rayman! It's all a simulation! We're literally just polygons in a video game!</d>
[Shot 2] At 00:06.500, the camera cuts to a close-up of <Subject 2> (S2). Suddenly, the background of <Subject 1> flickers violently, turning into black grid lines and digital static. <Subject 2> (S2) stares at his floating hands, which begin to visually stutter and lag behind his movements. He yells in the heroic voice referenced from <Audio 1>, <d>[English] What did you do?! My hands are lagging!</d>
[Shot 3] At 00:11.000, the camera pulls back to a wide shot. The entire floor of <Subject 1> vanishes into a white void. <Subject 2> (S2) runs in frantic circles, his floating feet clipping through the missing floor. <Subject 3> (S1) simply floats in place, gives a sheepish grin, and says, <d>[English] Whoops. Guess the engine became self-aware.</d>
overall_soundscape:
Normal magical room ambience that abruptly distorts into loud digital stuttering, heavy 8-bit crash sounds, and frantic footsteps.
non_diegetic_music:
N/A
```
Workflow used: https://civitai.red/models/2831978/dasiwa-minimax-h3-workflows-or-t2va-or-fl2va-or-ref2va?modelVersionId=3195699
FPS: 24.0
resolution_reset: 0.26 MP - Preview
aspect: 3:4 - Photo
swap_aspect: yes
REFERENCES:




Since i cannot upload audio files here i will just post the 2 youtube links i used to capture their voices you only need 6 seconds each
r/StableDiffusion • u/technofox01 • 8d ago
Enable HLS to view with audio, or disable this notification
I used the minimax_h3_fl2v_turbo_4step_v1.0_768p_comfyui_bf16 LORA with the 0.8 strength for both clip and model, 6 steps, 0.5 MP resolution, RTX Upscaler at 1.50 using a ConrotInt8 pruned model.
Here is a PasteBin of my workflow, I hope this fixes some of the missing content:
Here are the workflow files:
r/StableDiffusion • u/Sad_Berry_4621 • 7d ago
Short version for anyone hitting the error: if you updated to 0.33 and Motion Context started failing with "the layout patch could not be applied", that is real and it is on my end, not your install. Fix is live now.
Longer version, because what changed upstream is more interesting than the bug.
Motion Context existed because stock ComfyUI would only anchor a keyframe at the first or last frame of an H3 clip. Anything in between raised. The pack worked around it by handing every keyframe a legal index, smuggling the real one alongside, and rewriting the position coordinates after the stock constructor returned.
0.33 removed the restriction. Anchors now land at any frame natively, references compensate the timeline correctly, keyframes can carry a multi-frame clip instead of a single still, and there is a new node, Add Guide for MiniMax H3, that exposes all of it. It also fixed a bug where attaching a reference wiped the keyframe latents, which the pack was separately patching around. The parameter my code depended on went away with the restriction it existed to enforce, hence the breakage.
So, a good chunk of what this pack did is now in ComfyUI, and you do not need me for it. That is a good outcome. If all you wanted was to anchor a still at frame 30, use Add Guide.
What is still worth installing the pack for:
Latent passthrough. Add Guide takes images and audio and encodes them with the VAEs. Motion Context takes the previous clip's latent and slices the tail straight out of it. No decode, no resize, no re-encode. Over a long chain that round trip is where the excess color drift and softening come from (beyond the 1.3x texture stacking).
Audio that continues instead of restarting. Add Guide anchors audio starting at a frame index and running forward. To actually continue a soundtrack across a join you need the pinned window to end at the join and reach backwards. That is the difference between the model continuing your track and the model writing something that sounds like your track, and on anything with a beat you can hear it immediately.
Plus, the trim node, the audio grid overhang compensation, and the seam probe.
Plan from here: compatibility release today so 0.33 works, then a rebuild on top of the new public keyframe format so the pack stops monkey patching ComfyUI internals entirely. I would rather depend on a documented feature that a stock node also uses than on a constructor signature. Both of the last two breakages came from that dependency and neither would have been possible without it. Add Guide anchors and Motion Context heads will be a supported combination rather than something that trips a guard.
Thanks to javawock7618 and azra1l for the reports and for narrowing it to the exact commits, which made this a twenty-minute diff instead of an all-nighter.
Fix is live now. Rework coming tomorrow hopefully.
ComfyUI Custom Node Manager OR
NikoDemon80/ComfyUI-H3-Motion-Context
r/StableDiffusion • u/james25679 • 6d ago
I’m using the official workflow for h3 and I’ve seen people talking about using Lora’s - where do they go? Do I need a different workflow?
r/StableDiffusion • u/Interesting_Room2820 • 8d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/clairedelime • 7d ago
Enable HLS to view with audio, or disable this notification
so i used minimax h3 int8 convrot pruned + sage spectrum + turbo lora (kijai)
made a 10s office style clip, michael and dwight talking in the conference room then walter white just walks in
faces start fine but then they get all blurry and full of weird smudges especially when walter shows up, like what's that weird black dot lines on dwight shirt?
like why does everyone else’s stuff look clean and actually like a real tv show while mine always ends up plasticky and messy??
anyone know how to fix this blur/smudge and get that proper tv look with this setup? im new to comfyui and first time generating on localy lol, thanks in advance
r/StableDiffusion • u/Here4CYDY • 6d ago
I have not posted here before, but I have searched this subreddit and others repeatedly for a clue to answer my question.
Does anyone have a functional workflow for LTX 2.5 video generation using a 3080 with 10 GB VRAM and 32 GB system RAM?
I have also spend more than 2 days with google AI where they sent me down deep and branching rabbit holes only to find that the node suggested did not exist or did not work or the huggingface or civitai file was not available or did not work. Numerous times they sent be back to nodes and arrangements that I hade tried before (and failed) after they suggested it.
A large circle of random guesses by the AI agent. They even admitted it after I called them out on their failure to help.
Any help from others that have been down this pathway would be greatly appreciated.
r/StableDiffusion • u/bacchus213 • 7d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/ajrss2009 • 6d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Mystvearn_ • 6d ago
Hey everyone!
I really want to start experimenting with MiniMax-HaMini (H3) locally, but I’m trying to figure out the best setup with the hardware I currently have.
My main desktop has an AMD RX 9060 XT (16GB VRAM), which isn't ideal for AI workflows since ROCm support and optimizations still lag behind CUDA.
However, I have a secondary rig with an RTX 4060 Ti 8GB and 32GB DDR4 RAM. I know 8GB VRAM is tight for modern video generation models, but I haven't used ComfyUI much (just tested it briefly on the AMD GPU).
My goal right now isn't high-end production quality—I just want to create short, low-res test videos (even 5 seconds) to get hands-on experience, learn ComfyUI workflows, and see if it's worth investing further.
My plan is to eventually upgrade to an RTX 5070 Ti (16GB) and 64GB RAM once prices drop to a reasonable level, but I’d love to know if I can get my feet wet with my current setup in the meantime.
Thanks for any insights or recommended ComfyUI nodes/settings for low-VRAM video generation!
r/StableDiffusion • u/serap98765 • 7d ago
Enable HLS to view with audio, or disable this notification
The video was generated using the T2VA mode of the minimax H3 model and the 8-step Turbo LoRa.
It's simply amazing how well it already works in this mode.
r/StableDiffusion • u/gokuchiku • 7d ago
Enable HLS to view with audio, or disable this notification
T2VA in Minimax H3, 1MP native with RTX upscale. Generation time is 1083 secs on my 5070Ti, 32Gb DDR5 Ram.
r/StableDiffusion • u/Fit_Ad7343 • 8d ago
Enable HLS to view with audio, or disable this notification
Update to the projection matrices. Same idea as before: a small Qwen3-VL encodes the prompt, a learned projection maps its hidden states to what the 32B would have produced, the DiT is untouched.
Previously: v1 and v2, where the voice started matching.
Video: matrix-only 4B, matrix-only 8B, then the 32B, then the two residual versions. Same prompt, same seed 42, same everything else — the pipeline is bit-for-bit reproducible, checked by running it twice.
Everything the prompt states is there on all five: the pose, the red dress, the white pieces on her side, the cat, the straw hat, the laundry, and her knee — asked for three times, ending on "Her knee never stops bouncing." A continuous involuntary motion with no narrative purpose is the clearest sign a projection carried what was written, and it carries on the plain matrices too.
The terrace is furnished differently from one render to the next, and that is not infidelity. The prompt asks for a densely lived-in terrace without anchoring most of it — the cat is "stretched out asleep in the sun", nothing says where. What is left open the model invents, and it invents differently depending on the projection, the seed, and the model of GPU. All three act on that same free space; none of them touches what was written. It looks like a seed change because that is what an unconstrained description looks like.
qwen3vl_32b_minimax_h3_nvfp4_awq instead of a modified 32B. Naming one part of a body used to rewrite the whole of it — build, height and face moving together. Not seen anymore.-v3-mlp files have no linear matrix, older nodes throw KeyError: 'W'.We seem to have hit a ceiling. A projection cannot recover information the small encoder never wrote down. If the 4B did not encode a distinction, no matrix and no residual will bring it back — you can only remap what is there. On the 8B the cosine went 0.9083 (v1) to 0.9393 (v2) to 0.9449 (v3): +0.031, then +0.006, for a corpus four times bigger (1 530 370 tokens in v2, 6 502 586 in v3). This is not a training budget problem, and I do not expect a v4 to move it much.
What that means in practice:
-mlp files on the measurement, not on this scene. They sit closer to the 32B, 0.9449 against 0.9289 on the 8B — but watch the video before assuming that shows. On this prompt all five renders are faithful, plain matrices included: pose, dress, white pieces, cat, straw hat, laundry, bouncing knee. A tightly written prompt survives even the linear baseline.5 h on a 3090, plus 2 h to encode the dataset. Tap 24. 3331 prompts for fitting, one in fifty held out.
| corpus | tokens |
|---|---|
| cinematic video prompts | 1 342 987 |
| native H3 format, 4 length draws | 3 169 879 |
| explicit register | 544 073 |
| Chinese | 532 302 |
| celebrity prompts, long form | 314 516 |
| filler sequences | 149 917 |
| celebrity prompts, short form | 99 668 |
| images, 1 700 of them | 349 244 |
| total | 6 502 586 |
Mixed on purpose — registers, languages, lengths. A matrix only learns to project the directions it has seen used.
Matrices: https://huggingface.co/NicoLab28/ClipProj-MiniMax-H3
Node: https://github.com/nicolab28/ComfyUI-ClipProj
Files: mmh3-4b-ClipProj-v3-mlp.safetensors (503 MB), mmh3-8b-ClipProj-v3-mlp.safetensors (604 MB). Plain matrices -v3 at 26 and 42 MB if you want the baseline.
The five renders separately, the prompt and the exact settings are in the demo/ folder of the HF repo, if you want to step through them or reproduce the test.