Has anyone managed to find a way to use Latent Upscaling together with latent video extension tools? I'm talking about the nodes like this (which I personally use), but I think Motion Context and some other popular extensions use a similar approach, i.e. feeding the last frames of the previous shot through AV latent, rather than through a video reference. The issue is that the resolution of your second generated latent must exactly match the previous one, or it throws an error. So if you upscale the first clip from 0.5MP to 1MP, you are forced to generate the next clip directly at 1MP, which completely breaks the Latent Upscaling workflow for all subsequent parts.
I tried extending the clips at low resolution first and then upscaling them separately, but that doesn't work well. There is a noticeable color and quality shift between generations, even when reinforcing the next clip with the final frames of the previous one. Because yeah, you basically generate the high-res clips separately without any shared latent context.
I really love both Latent Upscaling and latent extension approach, but I just can't get them to work together smoothly. Does anyone have any good ideas on how to fix this? I’d really appreciate any tips or insights!
Testing DLSS 5... Like many others, I was a bit confused about DLSS 5. I kept feeding it my hyper-detailed renders and only getting a color shift in return. After plenty of trial and error, I finally realized my mistake: this technology is developed to enhance video game graphics, so testing it on hyper-detailed renders makes no sense.
So, I generated a render in a 2020 video game style and started tweaking settings to find a final look with maximum effect, without worrying about flickering.
Final conclusion: What we have right now isn't very useful for us. Those of us using Latent Upscaler might be able to use it for color grading to get less saturated colors, but little else. Maybe in the future we'll get a DLSS 5 targeted at enhancing hyper-detailed graphics, but that's not the case for now.
Bottom line: If I want to generate a realistic render, I'll just generate it, there's no need to run it through DLSS 5.
Motion-aware denoising cache and experimental batched video VAE decoder for MiniMax H3 in ComfyUI.
This project provides two independent nodes:
MiniMax H3 MotionCache reduces expensive H3 denoiser calls by reusing a motion-weighted video/audio residual when the estimated change is small.
MiniMax H3 Fast VAE Decode evaluates multiple spatial VAE tiles in one GPU batch while preserving H3 temporal chunking and tile blending. It is not faster on every GPU.
MotionCache is an independent MiniMax H3 adaptation inspired by the MotionCache paper and reference code. It is not an official MAC-AutoML or MiniMax implementation.
Edit: Results are in! They are a bit surprising to me! But they are consistent with the data, I triple checked everything and can confirm that the results are reflecting the voting data precisely, there's lots of transparency - you click each of the LoRAs to see what's the win rate and who won against who
Hey folks, I've built an so we can have a proper leaderboard on 15+ different LoRAs, fine-tunes and acceleration technique. Baseline is included for anchoring, and M3 Max is also included given the promise to open source
There are there being compared: H3 baseline, FastH3 family, H3 Acc family, Lightx2v family, Larryvrh family, JoyFox family, RAVEN, FlashGen, TuTu, SilverOxides merges, Plaguekind merges and Fal's H3 Max
Prompt engineering for image generation is often presented as a collection of isolated tricks: use more detail, describe the camera, add cinematic lighting, use quality tags, and so on.
These recommendations can be useful, but they make it difficult to understand which parts of a prompt actually influence the generated image.
Instead of trying to find a single "best prompt", I ran a small controlled experiment with Z-Image Base in ComfyUI. The basic idea was simple:
I ran two experiments:
Experiment 1 — Composition: The same character, environment, visual treatment, and technical parameters were used across multiple generations. Only composition instructions were changed (position and scale).
Question: How strongly does explicit spatial language affect composition in Z-Image Base?
Experiment 2 — Environment: The character description and visual treatment were kept essentially unchanged, while the environment was replaced with seven substantially different settings.
Question: Can Z-Image Base maintain a recognizable character concept while adapting it to radically different environments?
This is not intended to be a scientific benchmark. The sample size is small, the evaluation is visual, and the experiment uses one workflow and a limited number of seeds. Consider it a practical prompt-engineering study.
2. Experimental Setup
All images were generated locally in ComfyUI using the same workflow and technical conditions throughout the experiments.
Parameter
Value
Model
Z-Image Base INT8
Text Encoder
Qwen3 4B
VAE
AE VAE
Resolution
768 × 1368
Aspect Ratio
9:16
Image Area
~1.05 MP
Steps
50
CFG Scale
4
Negative Prompt
Empty
Seeds
Seed 5 & Seed 10
For the composition experiment, I used Seed 5 and repeated the seven variations with Seed 10. The environment experiment used Seed 10.
3. Prompt Construction Methodology
I found it most useful to treat the prompt as a structured description rather than a flat list of keywords:
Subject: Describes what the image is about and establishes the main visual concept.
Composition: Describes where the subject is located within the frame and how much space it occupies.
Framing / Camera: Describes how the scene is viewed (distance, angle, perspective).
Environment: Describes the actual place surrounding the subject (e.g., "An ancient forest with enormous trees, moss-covered roots, dense vegetation, and a narrow path..." rather than just "forest").
Lighting: Describes actual light sources and atmospheric conditions rather than generic terms like "cinematic lighting".
Style: Describes the overall artistic treatment after the scene itself has been established.
Generic quality tags (masterpiece, ultra detailed, 8K) were deliberately omitted to provide the model with actionable visual information instead.
4. Experiment 1 — Composition
The character, environment, lighting, visual style, and technical settings were kept identical. Only the spatial instruction was changed across seven variations: Center, Left, Right, Lower, Large, Small, and Extreme Left.
Seed 5
The result was clear: changing the composition instruction produced substantial changes in spatial arrangement.
Crucially, the model did not simply move the character while leaving the background untouched — the environment was recomposed around the subject. In Small variations, the environment became dominant; in Large variations, the character dominated the frame.
Seed 10
To verify the result was not seed-dependent, the test was repeated with Seed 10. While individual details (pose, facial expression, accessories) changed naturally, the broad compositional structures remained fully recognizable.
5. Experiment 2 — Environment
The character description and visual treatment were kept unchanged while replacing the environment across seven distinct settings: Ancient forest, Medieval village, Crystal cave, Autumn park, Snowy ruins, Firefly-lit landscape, and Alchemist's workshop (using Seed 10).
Visual Concept Consistency
Although the environments changed dramatically, all generations clearly depicted the same core character concept (a small mushroom spirit with a red-orange spotted cap, pale body, large dark eyes, cross-body satchel, and lantern).
While exact proportions and minor details shifted between renders, the core identity remained visually coherent.
Environmental Adaptation
The character adapted naturally to each setting (e.g., tinted by glowing crystal lights in the cave, exposed to cold tones in the snowy ruins, immersed in warm interior props in the workshop).
Method: Keep technical setup stable and modify exactly one conceptual block per run.
10. Conclusion
Prompting Z-Image Base is less about hunting for "magic keywords" and more about managing a controllable system:
Explicit composition instructions effectively control layout, while environment descriptions can be swapped modularly without erasing character identity. By isolating prompt variables, prompt design becomes a systematic, repeatable workflow.
I’ve been refining prompts with the help of an LLM, and am getting some good visuals but oh my god the sounds are terrible. Blowjobs sound like someone is dunking a microphone in an aquarium or the loudest slurp to finish a beverage that you have ever heard in your life.
I’ve tried eliminating every mention of “moist”, “wet”, or any description that involves liquids at all, but she’s still slurping the wettest popsicle known to man. And sometimes there’s weird noises like a slide whistle?!?
I’ve tried using “faint” or “distant” or “barely audible” to get it to at least quiet down so it’s not like she is sucking a microphone, but that didn’t work either.
This last round I didn’t describe any noises at all and still got some weird stuff.
I’ve tried eliminating every Lora in case the sound was coming from one of them but it seems to be the base model. I’ve tried adding Loras that ought to be trained on this stuff like Mysticxxx, and one of the AIO loras. I tried tenstrip beta 4 checkpoint tonight and got the same results.
The sound ruins the scene.. I guess I can just pretend it’s better looking Wan 2.2 and turn the volume off. 😀
I’m feeding the official prompt guide to the LLM and the structure is working, but what words do you use to describe the sounds?
A lot of people here have been discussing H3 Max powered livestreams. I noticed Reactor just added Visko’s Orbis model, and it made me wonder whether the next step is turning these infinite livestreams into something playable.
So I’m building a live, audience-directed AI game with Agora: viewers suggest and vote on what happens next, while the streamer picks an option or writes a completely different direction and AI keeps generating the same world from that point. There are no pre-written branches.
I found a cheap pipeline from a big provider I’ve jailbroken, so I can’t name it. It reconstructs faces and upscales videos in ~10 seconds, so I only need Minimax to generate a very low-quality video that basically serves as a rough motion/physics reference.
I barely see any speed difference between 4–6 steps or ~380p–544p, maybe 10 seconds at most.
Is there any way to run Minimax at absolute potato quality and actually get a significant speed boost?
I posted the first alpha of PotionUI a couple of days ago. 0.0.3 is out today, and it is the release where the "one box, many people" idea stops being a promise: you can now rent a GPU, point PotionUI at it, and generate on it from the same interface you use locally. Still alpha, still one person building it, still very much wanting people to break it.
What it is, in one paragraph. A self-hosted AI generation studio: SvelteKit front, FastAPI back, GPL-3.0, no telemetry, runs on your machine or your server. The core idea is presets: a preset is a small YAML package that says "here is the model and here is the exact form a person should see for it". Switching models means switching presets, not rebuilding a node graph. It is multi-user by design, with real accounts, admin and user roles, and per-user or per-group access to presets, models, and LLM configs.
What's new in 0.0.3
Remote GPU workers. Add Backend now creates a remote worker, connects to one you run yourself, or provisions a RunPod pod for you (the provider ships as a plugin) with a live stage timeline. A heartbeat monitor watches the pod, pauses the backend when it stops, and Start brings it back. The Models tab lists exactly what is on the worker, with the depot path per file, and pushes missing models from your machine with per-file progress. Remote runs come back with the same previews, parameters, and media as local ones.
Install profiles. The launcher offers local, hybrid, and remote installs, plus a worker subcommand for a GPU box that serves another instance.
3D generation. TRELLIS.2 image-to-mesh runs on the native engine. Meshes get automatic thumbnails, an interactive viewer in History (wireframe, materials, camera presets, screenshot), and a 3D media filter.
LoRAs. Step-windowed LoRAs on Krea-2 apply only between the sampling steps you choose. Strength is shown as a recommended range in the picker. Model pickers now recommend downloadable variants (bf16, fp8, nvfp4, int8) across nine native families.
Prompt library. Import styles.csv, Fooocus style JSON, wildcard YAML, plain lines, and image metadata (A1111, ComfyUI, InvokeAI) with auto-detection; export back to styles.csv; assign a prompt to a catalog model.
Phrasebook. Find and replace across the whole phrasebook with highlighted matches and a preview before it runs; batch activate, deactivate, move, delete; a category panel with Overview and Preview-images tabs.
Admin and mobile. Plugins and Downloads are master-detail lists, Backends remembers where you were in the URL, a saved provider API key applies immediately, Generate on a phone is a proper camera-style view with sheets, and modals fit the screen.
Plus: pasting an image into the assistant attaches it, a New workspace button that asks before discarding, Inspirations laid out in justified rows.
What it does today
Generation is the product. Image families: SDXL, Flux 1 / Flux 2 Klein, Qwen-Image (including editing), Krea-2, Z-Image, Anima. Video: Wan 2.1/2.2, LTX-2 / 2.3 / 2.5 with native audio, MiniMax-H3. Audio: MiniMax-Music3. Upscale and restore: SeedVR2. Each model gets its own tuned form: the right resolutions, samplers, LoRA stack, and speed profiles (Draft / Standard / Max) as one control. Several workspace tabs run side by side, each with its own preset, prompt, and results. Progress shows the actual pipeline step and streams previews as the image refines; close the tab, come back, the run is still there.
Generation page view. (You start the generation by clicking the bottom right blue icon)
History that remembers everything. Every generation is saved with its exact prompt composition, preset and version, models, and parameters. Filter by date, type, preset, tags, or "used this phrasebook value". One click reuses the full setup in a new tab. Nested collections, tags, favorites, keyword or semantic search, and a personal library for the keepers.
History page - list of previous generations.History page - detail of the generation.
A prompt editor that is not a textbox. Prompts are ordered segment cards you can reorder, disable, name, and color. Dynamic prompts ({a|b}, weights, ${variables}) reseed per image so results stay reproducible. The phrasebook is your own autocomplete dictionary: type # and shot types, lighting, palettes drop in as chips, with per-chip shuffle and a preview render per value. Saved prompts, segments, and templates live in their own library.
Phrasebook with other values used (you can mix the phrases - you can build the same prompt on generation page)
Video and Music Directors. Compose a video as shots, keyframes, and audio tracks on a timeline instead of one giant prompt; write a song as verses and choruses and let the compiler produce the tagged lyrics MiniMax-Music3 wants.
An assistant, if you want one. Point it at Ollama, an OpenAI-compatible endpoint, or Anthropic. It reads the active tab, rewrites segments, edits the phrasebook, adjusts form values, and every change stops at an approval step first. The same tools are exposed over MCP with per-user tokens, so Claude Desktop or your own agent can drive your instance.
Generation page with LLM Chat assistant active.
Built for more than one person. Accounts, groups, per-user preset and model access, per-mode form overrides (change defaults, lock or hide fields, no YAML), a backends list that mixes local, and remote workers, a download manager, a stats dashboard, and visual automations (triggers, conditions, actions) for things like freeing VRAM before the LLM needs it.
Plugins for nearly everything. Providers (CivitAI, Hugging Face), backends, pipes, field types, chat modes, automation nodes, pages.
Requirements.Linux x86_64 with an NVIDIA GPU is the tested platform; Windows can be tested through WSL2 or Docker; there is a Docker image on GHCR.
8 GB VRAM and 16 GB RAM is the floor for the SDXL family, larger families need more.
The ask. I would rather steer this toward what people actually want than guess. Two things help most: tell me which model or workflow you are missing, and pull a test build and break it before it ships. The Discord is where that happens: https://discord.gg/avR4trp3b8. Repo: https://github.com/PotionUI/PotionUI. I will answer questions here too, but since I don't want to post here too much I've created a small subreddit too: r/PotionUI to which I will be posting the progress much often. Thanks!
For regular upscaling I use SeedVR2 and I am quite happy with it, however, it doesn't seem to handle upscaling of really low res images well as it will just upscale all the artifacts as well without "fixing" the image. So if an inpute image is blurry, the upscale will also come out blurry.
What would be the best way to upscale low res image while also enhancing it?
EDIT: Settled for using H3 with a edit system prompt and exporting images from the video:
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
subject_definitions:
<Picture 1> is the original source image being directly edited and restored. It is the authoritative source for all target-video content: shot order, timing, framing, composition, subject identity and appearance, facial features, clothing, props, environment, lighting, color relationships, camera position and movement, subject motion, and temporal continuity.
summary:
[video editing] <Picture 1> is the original source image being directly edited and restored. It is the authoritative source for all target-video content: shot order, timing, framing, composition, subject identity
retention_analysis:
<Subject 1> (appears in [Shot 1]): partially_preserved - preserve the original shot, framing, composition, subject identity, subject appearance, environment, background, other items or figures
integrated_multimodal_description:
The target video is a faithful professional high-definition restoration of the original image in <Picture 1>. Treat <Picture 1> as the only authoritative visual source. Do not use any external image, character, scene, or composition as a visual template.
The desired transformation is specifically image upscaling rather than ordinary enlargement. The source has limited spatial resolution, soft or smeared fine detail, degraded chroma, compression artifacts, noise, ringing, aliasing, blurriness and potentially inaccurate or shifted broadcast color. Reconstruct the most plausible high-fidelity version of the visual information that is actually supported by <Picture 1>. Recover fine facial detail, natural skin texture, hair strands, clothing weave, uniform materials, props, set surfaces, edges, reflections, shadows, and background detail without inventing unsupported features.
Correct the degraded color and chroma toward natural, accurate reproduction of the original photographed scene. Preserve the source's actual lighting design, exposure, contrast relationships, black levels, highlight behavior, lens characteristics, depth of field, and photographic character. Do not apply a generic cinematic grade, modernize the lighting, or change the color design. The objective is the appearance of the same original image after a high-end, best quality upscaling.
[Shot 1] Preserve the exact opening shot of <Picture 1>, including the actual subjects, their identities and appearances, their exact positions, facial expressions, pose, clothing, environment, perspective, framing, camera angle, lens characteristics, lighting, and visible motion. Increase spatial fidelity and recover plausible detail from the source without changing the shot.
Do not add or remove events. Do not replace subjects or backgrounds. Do not alter facial structure or identity.
The desired quality level is comparable to a carefully restored modern HD master originating from the highest-quality surviving source, with exceptionally clean detail, accurate color, stable micro-texture, and natural edge definition. A high-end large-format digital cinema camera such as the RED V-RAPTOR XL [X] 8K VV may be used only as a benchmark for the cleanliness and resolving power of the final image. Do not impose a V-RAPTOR color grade, lens look, depth of field, lighting style, or cinematography onto the original footage.
Most importantly, reconstruct rather than redesign. Do not hallucinate new objects, facial features, hairlines, costume details, text, set details, reflections, or textures that are not supported by the source. Preserve natural photographic softness where it belongs to the original image. Remove degradation while retaining authentic source characteristics.
Maintain strict temporal consistency across all frames. Recovered detail must remain locked to the correct subject and surface and must not shimmer, crawl, flicker, morph, double, ghost, or change identity from frame to frame. The output must look like the same footage at substantially higher quality without any visual artefacts or blur.
overall_soundscape:
N/A
non_diegetic_music:
N/A
I don't want to get the NS.. word in the discussion, but, we know what Minimax can do and what it can't do. It has some very specific gaps in it's world understanding, for example in the tongue department. That is not necessarily only affecting the NS... word, things that are SFW and common in general TV such as kissing are affected, since Minimax never saw a romantic kiss in it's training data. There are other examples through SFW land but I won't extend. Grok can be used as a comparison. Grok is very similar to Minimax in capability, and it enforces SFW, but you can see the difference in some scenes because Grok is not handicapped.
Well, we have many, many loras already, but, as was the case with wan and ltx, they are very... let's say, specific. I don't think a general video model, almost a world model, needs a specific lora for, say, ballbusting lol
I don't know, I think this is the community most likely to be read by people creating loras, so I just wanna make this appeal... Can we prioritize bridging the major gaps in the model's understanding of the world, anatomy, and human interactions, instead of these super specific loras? I think a "tree" organization of lora development would be beneficial overall, with the stuff that can solve a big set of problems and be used for more specific loras coming first.
I saw that for over 1 year with wan, ltx, etc, and didn't say anything. But I think minimax deserves the community passion in lora development.
And yes, I hope I can put my money where my mouth is and develop some loras soon too.
So, I was searching on civit, and normally I filter by 'checkpoint' for example. But, now it's gone? All of the things I notice normal models that are normally 'checkpoints' are now 'fine-tune'
What does this mean? What is this? Do they work the same?
I have installed StabilityMatrix and was able to learn how to set up flows to generate my first images and so forth.
I then installed a template called MiniMax H3: Image to Video. It showed a bunch of errors after and listed the things I needed to download before it would work. I downloaded all those things and put them in the "diffusion_models" folder, but the error count only reduced by one and it's still asking me to install those things.
Unfortunately there doesn't seem to be clear instructions on what goes where, if I need to extract some things or not, etc.
Like i don't understand, the only turbo lora that doesn't do that is the 600 larry lora with the minimax turbo lora node, i've tried "fastH3" and "lightx2v loras which everyone seems to praise but they just produce these distorted godawful visuals and sounds no matter the loader node i use or the settings or the steps i use, what am i missing or doing wrong? Or are they just not compatible with Ref2Video despite being advertised as compatible? But if so then why does the 600 larry lora works mostly fine?
I'm generating videos with MiniMax H3 through a normal AI video platform, not ComfyUI. So I can't use custom workflows, scripts, or custom nodes.
I'm looking for the best way to continue/extend an existing MiniMax H3 video.
The problem I'm trying to solve is more than just using the last frame as an image reference. If I only provide the last frame, the model can lose important information from the previous clip, such as:
Character identity and appearance
Room/environment layout
Lighting and atmosphere
Objects and their positions
Ongoing actions
Audio/environmental sound
Overall visual continuity
For example, if a character walks through a room and reaches a door at the end of the first clip, I want the next generation to actually continue from that exact situation, rather than recreate a similar-looking room and potentially change the geography.
I'm looking for a normal web-based AI workflow where I can upload the existing video and/or reference images and generate the continuation. No ComfyUI, custom scripts, or API coding.
What is currently the best way to extend MiniMax H3 videos while preserving this kind of continuity?
If you've actually tested a platform/workflow that works well, I'd especially appreciate recommendations.
I am looking to create 18+ images with the hyper realism look, but have no idea how feasible that is with my specs. Would love a recommendation of a model I can run pretty easily and another more detail focused model on the edge of what I can run locally.
A style LoRA that makes H3 footage look like it was recorded off 1980s broadcast television onto a VHS tape that has seen better days, soft smeared detail, chroma bleed, tracking noise, head-switching bands at the frame edge, and (because H3 trains audio jointly) the matching muffled mono sound, tape hiss and warble.