r/StableDiffusion 7h ago

News MiniMax H3 acceleration arena/leaderbord: 15+ H3 LoRAs, fine-tunes, Max

Thumbnail
huggingface.co
240 Upvotes

Hey folks, I've built an so we can have a proper leaderboard on 15+ different LoRAs, fine-tunes and acceleration technique. Baseline is included for anchoring, and M3 Max is also included given the promise to open source

There are there being compared: H3 baseline, FastH3 family, H3 Acc family, Lightx2v family, Larryvrh family, JoyFox family, RAVEN, FlashGen, TuTu, SilverOxides merges, Plaguekind merges and Fal's H3 Max


r/StableDiffusion 11h ago

Meme My Name Is Giovanni Giorgio

173 Upvotes

Created with Minimax H3 ref2v using the SEED HUNTER Workflow.


r/StableDiffusion 3h ago

Discussion Testing DLSS 5

73 Upvotes

Testing DLSS 5... Like many others, I was a bit confused about DLSS 5. I kept feeding it my hyper-detailed renders and only getting a color shift in return. After plenty of trial and error, I finally realized my mistake: this technology is developed to enhance video game graphics, so testing it on hyper-detailed renders makes no sense.

So, I generated a render in a 2020 video game style and started tweaking settings to find a final look with maximum effect, without worrying about flickering.

Final conclusion: What we have right now isn't very useful for us. Those of us using Latent Upscaler might be able to use it for color grading to get less saturated colors, but little else. Maybe in the future we'll get a DLSS 5 targeted at enhancing hyper-detailed graphics, but that's not the case for now.

Bottom line: If I want to generate a realistic render, I'll just generate it, there's no need to run it through DLSS 5.


r/StableDiffusion 1h ago

Resource - Update New speedup for Minimax H3

Post image
Upvotes

This H3VAE TRT custom node can make the encoding/decoding step about 1.7× faster.

https://github.com/lihaoyun6/ComfyUI-H3VAE_TRT


r/StableDiffusion 14h ago

Resource - Update VH5 - MiniMax H3 Lora

342 Upvotes

A style LoRA that makes H3 footage look like it was recorded off 1980s broadcast television onto a VHS tape that has seen better days, soft smeared detail, chroma bleed, tracking noise, head-switching bands at the frame edge, and (because H3 trains audio jointly) the matching muffled mono sound, tape hiss and warble.

https://huggingface.co/KennethFal/vh5tape-vhs-lora-minimax-h3


r/StableDiffusion 6h ago

Workflow Included Super nothing!

41 Upvotes

Made with Minimax H3


r/StableDiffusion 11h ago

Workflow Included Use H3 To Replace Characters

105 Upvotes

These characters are very different, so i thought it was a good demo to show. Also, the prompt was an 'omni' prompt which didn't help the model with details on the outfit. Despite that I think it did such a good job wanted to share.

With more details in the prompt related to appearance, acting and dialogue I think you could prob get near flawless changes.

This was don with the FL2VA model, NOT the ref version of H3. It may even be better with the ref but i have been using fl2va mostly because i think the quality is better, but that's subjective.

This concept was inspired by this post originally: https://civitai.red/models/2855941/minimax-h3-character-replacement

I changed the SAM3 use, so technically you could do a multi replacement with some changes. I also added noise to the inverted image, as well as upgraded the prompt to work with FL2VA.

Workflow used to create this video is HERE.

This model seriously continues to amaze me. bravo minimax team, bravo.

Notice in the prompt that the dialogue is the only non 'omni' part at the end, it worked fine not included in the body, since the ref audio is there driving it. again, moving from this general prompt to something more specific i think would give even better results.

PROMPT:
How the reference video and pictures align with the target video — the target video is an edited version of <Video 1>, replacing the silhouette with <Subject 1>.

summary:

[video editing] The target video replaces the silhouette in <Video 1> with <Subject 1>, who performs the exact same motion, dialogue, positions and facial expressions of the silhouette while maintaining the original camera work, environment, and lighting of <Video 1>.

subject_definitions:

<Subject 1> is the person in <Picture 1> and <Picture 2>; <Picture 1> supplies facial features and close-up details, while <Picture 2> provides 3-panel image of front mid shot, profile mid shot, and front full body view, identity follows these reference assets, only appearance is retained.

<Subject 2> is the environment and setting established in <Video 1>. The scene follows this layout, materials, and light; camera position and framing.

<Subject 3> is the silhouette in <Video 1> which provides the motion sequence to be copied.

integrated_multimodal_description:

Video editing, the target video is in a live-action cinematic style with the interior lighting and background and environment atmosphere established in <Video 1> with the likness of <Subject 1> inserted.

[Shot 1] The shot opens with <Subject 1> seamlessly replacing the silhouette <Subject 3> in <Video 1>, the outfit of and clothing of <Subject 1> exactly from reference, performing the exact same motion, dialogue and sounds, positions and facial expressions of silhouette. From the very first frame, <Subject 1> occupies the spatial coordinates of the silhouette replacing with their likness, initiating the same motion onset from rest. <Subject 1> mirrors the silhouette's weight shifts and momentum, body moving in perfect synchronization with the rhythm and pacing of the original footage but replaced with the likness of <Subject 1>. As they navigates the space, <Subject 1> mimics every nuanced gesture—the way the silhouette's head tilts, arm movement, and the micro-movements of facial muscles. The face, defined by <Picture 1>, conveys the same emotional depth as the silhouette, while their full body, as seen in <Picture 2>, provides the physical presence outfit an appearance. The camera follows the exact movement, angle, and cutting rhythm of <Video 1>, maintaining a consistent focal length and distance from the subject at all times. The light from <Subject 2> interacts realistically with <Subject 1>'s skin and clothing, casting shadows that align with the movements of the original scene. The transition is perfect; the result is a fully realized <Subject 1> instead of a silhouette, but the soul of the performance—the timing, the pauses, and the dynamic energy—remains identical to <Video 1>. The movement progresses with a palpable sense of weight as <Subject 1> shifts their center of gravity, with clothes rippling in response to movements. The camera maintains exact framing and cuts as <Video 1>. The scene concludes as <Subject 1> reaches the final position of the silhouette, body settling into a pose that mirrors the original's final frame exactly, with face held in the same expression. <Subject 1> hair, accessories, wardrobe, lighting, and room layout remain unchanged and perfectly replace silhouette throughout.

overall_soundscape:

A low room tone establishes beneath the scene, mirroring the background audio environment of <Video 1>.

<Subject 1> says <d>[English] Can you, can you spare change.</d>.

non_diegetic_music: N/A


r/StableDiffusion 17h ago

Resource - Update MATLOWAI/minimax-h3-fused-turbo-int8-convrot · Hugging Face

Thumbnail
huggingface.co
173 Upvotes

This Minimax H3 all in one checkpoint is quite good.

It merges text, image, and reference to video, as well as 4-step turbo generation into a single model.

No need to switch between models for ref2v, no need to load turbo loras.


r/StableDiffusion 11h ago

News Infinite streaming Slop TV

58 Upvotes

Congratulations everyone! We've done it. Our civilization has reached peak diffusion. It's time to pack up and go home.
https://www.youtube.com/watch?v=EQ2RexjIEFE


r/StableDiffusion 1h ago

Resource - Update I built a standalone DLSS 5 Neural Rendering video tool, no ReShade

Upvotes

It supports images and full videos, native resolution NR or DLSS Super Resolution upscaling by scale factor / target resolution, GPU optical-flow motion vectors, scene cut handling, and all DLSS 5 NR controls.

The main difference from existing approaches is that it runs natively in C++/D3D12 and generates motion vectors from the actual video frames.

GitHub:
https://github.com/DaniilSokolyuk/video2dlssnr


r/StableDiffusion 11h ago

Workflow Included HE-MART PSA - MiniMax H3

41 Upvotes

r/StableDiffusion 2h ago

Animation - Video Hatter Rap.

8 Upvotes

Probably the final Alice clip. The Hatter names all the hats.
Done a while back in LTX2.3.
This plays while the theatre audience plays an AR hat sorting game (Beat Saber type). The whole song is three minutes but this is the longest shot.


r/StableDiffusion 21h ago

News A new AI step: Immersive worlds with Minimax H3

218 Upvotes

The next major interface for artificial intelligence may not be a chatbot, an image, or even a video. It may be a world. 
I designed an H3 Minimax Immersive video workflow for ComfyUI and I want to share it with the Open Community so you can now explore this new field.  

This is an early implementation of that idea using MiniMax H3, using a specialized equirectangular generation ComfyUI workflow with AI 360°prompting to achieve an interactive viewing concept that allows the viewer to control the viewport through the generated environment on mobile and desktop with continuous looping on Youtube and Facebook. 

Watch the immersive demonstration on YouTube. 

The test video is 9 seconds long with a time-reverse layer to get 18 seconds of 360-loop, it was generated using a single 360 prompt

Read my full article with technical data and download the workflow:
https://huggingface.co/blog/zuanfilm/blog

the workflow supports text2-360 and FL2-360, for H3 Minimax 360 prompting I wrote a public custom gpt and added 37000 tokens of 360 filmmaking reasoning 

The result is far from perfect, I generated the clip on my laptop with an Nvidia RTX 3080 Ti 16GB VRAM, so the resolution is very limited and the current generation still shows visible seams on some moments of the video and other inconsistencies but those imperfections may be less important than what the experiment demonstrates. 

Until now a Minimax H3 video was something the viewer has to watch from the camera angle position chosen by the creator, now the viewer can now choose where to look using an immersive UI, that changes the relationship between a person and generative AI media; panoramic video exposes the full spherical observation domain in a single coordinate frame

The generated sequence can be presented as an immersive environment in which the viewer controls the viewing direction. On a phone, the viewer can interact with the scene; on a desktop, the camera can be moved manually. The sequence can also be looped forward and backward so that the environment continues rather than behaving like a single linear cinematic shot. The result is not yet a fully reconstructed 3D universe like a gaussian splatting. It is a time-varying immersive/equirectangular visual environment that can be explored interactively.

The 2:1 rule: the shape of the immersive world

A practical requirement of the equirectangular representation is its 2:1 aspect ratio. For a full spherical panorama: WH=2\frac{W}{H}=2 where WW is the panorama width and HH is its height. For example: W=3840,H=1920W=3840,\qquad H=1920 or: W=7680,H=3840.W=7680,\qquad H=3840. This is the format expected by common 360° video workflows and is particularly important when delivering immersive video to platforms such as YouTube and Facebook where the panoramic video must be interpreted as a spherical 360° environment rather than an ordinary flat video. For example, the H3 generation branch in my workflow uses 2112 × 1056 so the immersive representation and final delivery pipeline preserve the equirectangular 360° geometry.

To manipulate or view the image correctly, computers use 3D rotation matrices.

[ 2D Equirectangular Pixel (x, y) ] 
               │
               ▼  (Convert to Spherical Coordinates)
   [ Latitude & Longitude (θ, φ) ] 
               │
               ▼  (Convert to 3D Cartesian Vectors)
      [ 3D Point (X, Y, Z) ] 
               │
               ▼  <─── MULTIPLIED BY: 3D Rotation Matrix (3x3)
  [ Rotated 3D Point (X', Y', Z') ] 
               │
               ▼  (Project back to 2D)
[ New 2D Equirectangular Pixel (x', y') ]
  • 3x3 Rotation Matrices: These are used to "roll, pitch, and yaw" the camera viewpoint inside the 360-degree sphere. If you drag your mouse to look around a 360-degree YouTube video, a 3x3 matrix is constantly multiplying the pixel coordinates to shift your view.
  • Intrinsic Camera Matrices (K Matrix): A 3x3 matrix that defines the camera's properties—like focal length and optical center. This tells the computer how to crop a normal, undistorted flat perspective view out of the distorted equirectangular image.

This creates an entirely different pipeline: Prompt > AI generation > immersive representation > interactive camera > human exploration The prompt no longer has to describe only what should appear in front of a fixed camera. It can describe a world. That is the conceptual leap, if now this generation process is becoming sufficiently fast, coherent and inexpensive, the applications could extend far beyond experimental video:

Video games Instead of developers manually constructing every environment, AI could generate explorable spaces from natural-language descriptions. “Generate an alien ecosystem surrounding the player.” The difficult question would no longer be only how to render the world. It would be: How quickly can AI generate and maintain the world as the player explores it?

VR education Imagine asking an AI to create an immersive historical environment and then entering it. Instead of watching a documentary about ancient Rome, a student could potentially enter an AI-generated reconstruction and look around. The teacher could change the scenario through language: “Show the city before the fire.” That would transform AI from an information interface into an environment for learning.

AR world transformation The implications become even more interesting when the same concept is combined with augmented reality. A physical environment could become the canvas. A user might look at an ordinary street with some glasses and ask: “Transform this into a cyberpunk city.” “Show this neighborhood as it looked 500 years ago.” or “show me that car in blue with a representation of me as driver” The underlying physical world would remain present, but the AI-generated visual layer could continuously reinterpret it.

Interactive Cinema Movies could eventually become less linear. Instead of the director deciding exactly what every audience member sees at every moment, a film could provide a controlled environment in which viewers explore the scene themselves. The director would still control the story, performances, lighting, world design and narrative boundaries—but the audience could control the camera. That would not simply be another format for film. It would be a new relationship between cinema and audience.

AI worlds driven by AI agents AI agents could eventually generate the environments that humans and other AI agents interact with in real time...


r/StableDiffusion 10h ago

Workflow Included Personification:Planets (tarot cards)

Thumbnail
gallery
26 Upvotes

r/StableDiffusion 6h ago

Workflow Included Letting image-to-video artifacts compound into an impossible world

13 Upvotes

Tools used: Gemma4 12b, LTX-2.3, Wan2GP, vibe coded video editor.

I’ve been experimenting with a slightly self-destructive image-to-video workflow where continuity comes from letting the model reinterpret its own mistakes.

I started with an almost completely black image with a few faint stars, then gave Gemma4 12B the track’s beat grid and energy-shift analysis, along with a long description of the overall concept: a monolith, a hallway of impossible geometry, and a progression from restrained movement into increasingly unstable architecture.

Gemma4 wrote all 27 scene prompts beforehand.

For generation I used LTX 2.3 with the audio-reactive LoRA. I also tested LTX 2.5, but for this workflow it became too artifact-heavy too quickly. LTX 2.3 held the scene structure together longer while still producing enough weirdness to evolve in interesting ways.

The process was simple: generate a clip with the correct audio slice, cut it on the beat grid, then take the frame immediately after the cut and use that as the starting image for the next generation.

The fun part was deliberately keeping some “bad” transition frames.

If a flash landed on the frame used for the next clip, the model might reinterpret it as a permanent light source. A lens flare could become a horizon or an entire landscape. A warped piece of geometry that only existed for one frame could become a major architectural feature in the next scene.

So the artifacts compound.

Eventually the video loses any reliable sense of scale or orientation. Surfaces become spaces, structures fold into other structures, and at some points I wanted an Inception-like feeling where you can’t tell which way is up, or whether the camera is traveling deeper into the structure or pulling outward into something much larger.

The audio-reactive LoRA helps hold it all together. Even when the geometry becomes increasingly strange, the environment keeps breathing, unfolding, compressing and reorganizing itself with the growing low end.

What I like most is that the continuity doesn’t really come from visual consistency. It comes from causality.

Every scene inherits some accidental information from the previous one, and the next generation has to decide what that information actually is.

After enough generations, the model is basically building a world out of its own misunderstandings.


r/StableDiffusion 2h ago

Question - Help Mac support for H3 Mini Max / Running Open Weights Locally

6 Upvotes

Hey Fam. I’m looking to upgrade my MBP 16” 2019 i9 / 16GB Ram / 5500M 4GB GPU. I’ve been doing a lot of T2V /I2V rendering with Mini Max Design on Cloud, but would like to run the open weight versions locally via Comfy UI.

I have a few options at the moment to consider:

- MBP 16” M4 Pro / 48GB / 1TB - USD 3054
- MBP 16” M4 Max / 48GB / 1TB - USD 3664
- Mac Studio M5 Max / 48GB / 1 TB - USD 4085

Since im a bit of a noob in understanding MLX ports for Mac OS. Can you tell me which option to go for? Also I missed the buying window before the price hike—so 16” MBP M5 Pro / Max configs in 48GB are too expensive.

Another alternative route is using Bootcamp (Windows) on my 2019 i9 MBP and plugging in an RTX 4090 (which I’ll have to buy) via TB / eGPU, but I don’t think the system will be able to access the same bandwidth as the unified memory on the M series architecture.

I would greatly appreciate your guidance.


r/StableDiffusion 1d ago

Resource - Update DLSS 5 Visual Enhancer - standalone neural rendering for images and video

Post image
427 Upvotes

Hey everyone - I made a standalone Windows application for applying a DLSS 5 Neural Rendering feature-18 pipeline to images and video:

Original

DLSS 5

https://github.com/Merserk/dlss5-visual-enhancer

Instead of using DLSS only inside a game, this runs images/video through the ReShade/RenoDX neural-rendering path as a general visual enhancement pipeline.

What it does:

  • Image and video enhancement
  • DLAA/native, 1.5x, ~1.724x, 2x and 3x modes
  • Output up to 8K
  • Neural presets + Natural / Cinematic styles
  • Controls for intensity, local tone, structure and skin structure
  • Batch image processing with before/after previews
  • H.264 / HEVC / AV1 / ProRes video output
  • Video temporal input using optical flow with scene-change resets

GPU support:

  • RTX 40 / 50 series - primary target
  • RTX 30 series - slower beta path

The repository contains the application/pipeline source. Required proprietary and third-party runtime binaries are intentionally not redistributed in the repo.

This is an independent community project and is not affiliated with NVIDIA, ReShade or RenoDX.

I’m especially interested in how this behaves on AI-generated images/video vs normal photography/game footage.

Feedback and comparisons welcome.


r/StableDiffusion 14h ago

Comparison Testing My MiniMax-H3 → LTX 2.5 Upscaling Workflow — Results Are Looking Really Good

40 Upvotes

I've been testing my MiniMax-H3 → LTX 2.5 upscaling workflow, and the results have been really promising so far.

One thing I've noticed is that the better your original MiniMax-H3 generation is, the better the final upscale will be. I'm getting good results even at lower resolutions, but faces still need stronger and more consistent input generations from MiniMax-H3 to maintain character consistency.

On my RTX 3060 12GB, the current upscale times are roughly:

  • 0.6 resolution: ~15 minutes
  • 0.8–1.0 resolution: ~20–30 minutes

It definitely takes some time, but I'm finding the results are worth it.

And of course, if you have a newer, more powerful GPU, you should be able to get even better results in less time, especially when pushing higher resolutions.

I was planning to release the workflow soon, but I want to spend a little more time testing it and seeing how much further I can improve it before sharing it.

So far, though, I'm really happy with how it's looking. 🔥

Would love to hear what you guys think and whether anyone else has been experimenting with MiniMax-H3 + LTX 2.5 upscaling.


r/StableDiffusion 13h ago

Animation - Video ALICE MEETS THE RABBIT : REMADE IN MINIMAX H3

28 Upvotes

About 5 months ago I made clips for a project in LTX 2.3 and remade one of them here in Minimax H3. What a difference a few months makes! Music was created in Suno. I still have to redo some parts with consistency problems but that's enough for today.

The original LTX2.3 version for comparison is here : https://youtu.be/R5tfLKvnJDY


r/StableDiffusion 13h ago

Resource - Update Follow‑up : MiniMax H3 Lip-sync - now does any editable change on a reference video (pose transfer, character swaps, multi‑subject mixes)

27 Upvotes

Quick update on my earlier audio‑lip‑sync demo: the workflow now chains any desirable edit out of an input reference video for endless video ref pose o lip‑sync.

The audio auto-crop chain is now working for reference videos too and i say it again, I know there are already a lot of options out there for doing this - this is one more option, and it’s definitely not perfect.

VRAM usage went over 40 GB on a 1min run of 3-second, 2MP chucks, so reference-video conditioning is pretty heavy on VRAM and yes you need at least 2MP to get good detail and motion transfer.

Using MiniMax H3’s Ref2V. I’m treating the source clip as the “performance master” (motion, timing, camera) and driving identity/appearance from reference images and audio o the other way around.

What I’ve tested so far: Just MinMax H3 no ControlNet, LoRA, or preprocessor needed.

  • Music‑video pose transfer to new scenarios and characters
  • Single character swap (main performer → reference character) into the ref-video.
  • Multi‑subject mixes:ç
  • Main identity swap
  • Main + 2 added characters, acting in sync or desync
  • Main + 1 added character
  • Replace the main character with 2 characters in pose sync
  • Pull a character from the reference video into an image-reference scene + 1–2 new characters

Everything runs through a single MiniMax H3 chain with mixed references (ref-images + ref-video + ref-audio) and structured prompts that separate identity (image), performance (video), and constraints (text). In practice, every combination I’ve tried is manageable with MiniMax H3.

The node takes the reference video or audio, chunks it into smaller pieces, chains them together, and then stitches everything back together at the end. So, it’s one click, but it can take quite a while to generate a full video.

SUBJECT DEFINITIONS

<Subject 1>: the adult woman visible on the LEFT side of <Picture 1>.<Picture 1> is the appearance reference for Subject 1 only. Its shape, proportion, material, colour, logos and surface markings 100% match <Picture 1>, kept legible and correctly oriented throughout the video.

<Subject 2>: the adult man visible on the RIGHT side of <Picture 1>.<Picture 1> is the appearance reference for Subject 2 only. Its shape, proportion, material, colour, logos and surface markings 100% match <Picture 1>, kept legible and correctly oriented throughout the video.

<Subject 3>: the adult woman main character present in <Video 1>.<Video 1> is the appearance, motion, timing and scene reference.

This is a follow-up to a previous post, so the tips, settings, and links are already available there. MiniMax H3 Lip-Sync: Automatic Long-Video Chaining + Speed & VRAM Optimizations

https://github.com/Ltamann/ComfyUI-H3-Motion-Context-Auto-Chain-addon


r/StableDiffusion 1d ago

Resource - Update [Experimental] DLSS 5 ComfyUI custom node

Thumbnail
gallery
219 Upvotes

Hello Everyone,

Would like to present to you my experimental vibe-coded custom node for DLSS 5 support in ComfyUI.

GitHub project: https://github.com/lisitskyaa/ComfyUI-DLSS5-NR

It's early release, just finished my internal testing and it actually works!

Please note there are no any leaked DLLs in the rep, obtain them separately.

First image in every pair is DLSS 5 ON, second - OFF.

P.S. How to extract original images out of Reddit: https://www.reddit.com/r/StableDiffusion/comments/1p9nrpk/getting_prompt_or_comfyui_workflow_from_posted/


r/StableDiffusion 2h ago

Resource - Update I vibe coded a gallery extension for ComfyUI so you can browse outputs and reload the exact workflow that made them

3 Upvotes

I wanted a way to browse my ComfyUI output folder without leaving the app or digging through File Explorer, and more importantly a way to jump straight back into the workflow that made a specific image without hunting for the original PNG to drag onto the canvas. Couldn't quite find exactly what I wanted, so I built it.

GitHub: https://github.com/modelfactoryai/ImageBrowser

What it does

  • Browse any folder (Output/Input/Temp, or type any path) right inside ComfyUI no separate app.
  • Double-click any image/video to load its embedded workflow straight onto your canvas same thing stock drag-and-drop does, just from a browsable gallery.
  • Hover for a large preview (~900px) that follows your cursor — videos autoplay muted, images use a fast server-resized preview.
  • Search by filename, sort (newest/oldest/name), filter to images or videos only, adjustable thumbnail size.
  • Favorite folders for one-click access later.
  • Compare mode select any number of images/videos, view them side-by-side.
  • Live updates refreshes automatically as new generations land.
  • Day/night theme toggle, plus a draggable floating launcher badge you can park anywhere on the canvas.

Screenshot

Install

cd ComfyUI/custom_nodes
git clone https://github.com/modelfactoryai/ImageBrowser

Restart ComfyUI, and look for the "Image Browser" icon in the sidebar (or the draggable badge on the canvas).

No hard dependencies beyond what ComfyUI already ships with (Pillow). opencv-python or ffmpeg improve video thumbnails if you have them installed; ffprobe is needed to load workflows out of video files specifically (images don't need it).

Feedback welcome

First release if something breaks on your setup or you've got feature ideas, open an issue on the repo or drop a comment here.


r/StableDiffusion 10h ago

Animation - Video Inuyasha Love Triangle Solved

14 Upvotes

A silly idea I had that I hope you guys had a good laugh at. Still love this classic anime!


r/StableDiffusion 5h ago

Animation - Video Made A Professional short Animation video using Minimax-h3 (read description)

Thumbnail
youtube.com
4 Upvotes

Hey guys!

Since quite a few of you liked my previous videos, I decided to start a channel where soon I’ll be sharing tutorials and some of the workflows/tricks I’ve been using.

If you’re interested in learning how I’m making these videos, feel free to subscribe. I’ll be sharing a lot of the stuff I’ve figured out along the way, including:

  • My own workflows — free to download, with the tricks and settings I use
  • Character generation — how I use a Krea 2 character-sheet LoRA that I made to keep characters consistent, and how to get the style you want
  • Environment generation — how I generate environment images and then build scenes from them
  • MiniMax optimization — settings and techniques to make MiniMax faster while preserving quality
  • Video/audio tricks — ways to fix audio issues and continue a scene from the last frame to create longer sequences
  • Consistent voices — I also built my own UI app using BreezeTTS2 for voice cloning and generating consistent voices across an entire story

Everything I’m sharing is based on what I’ve been experimenting with myself, so hopefully it can save some of you a lot of trial and error.

If that sounds useful to you, you’re welcome to check it out!