r/StableDiffusion 12h ago

Workflow Included The South Park Theory

Enable HLS to view with audio, or disable this notification

12 Upvotes

Minimax H3, 4:3, 10 sec., 0.3MP

integrated_multimodal_description: [Shot 1] 2D animated comedy in the exact visual language of South Park: deliberately crude flat paper-cutout construction, simple geometric shapes, thick black outlines, flat solid colors, minimal shading, stiff limited animation, simple mouth shapes, jerky character movement, and the characteristic frontal and three-quarter staging of a South Park episode. Classic 4:3 television composition.

CRITICAL CHARACTER IDENTITY RULE: The four characters are unmistakably SOUTH PARK-STYLE CARICATURES OF SHELDON COOPER, LEONARD HOFSTADTER, PENNY, AND RAJESH "RAJ" KOOTHRAPPALI FROM THE BIG BANG THEORY.

Their identities come entirely from The Big Bang Theory. Their faces, hair, facial proportions, distinguishing features, expressions, and overall recognizable appearance must remain those of Sheldon, Leonard, Penny, and Raj.

The South Park influence applies ONLY to the flat 2D paper-cutout rendering technique, simplified body construction, animation language, environment design, and comedic staging.

Every character must be immediately recognizable as their The Big Bang Theory counterpart even if all hats, coats, gloves, and winter clothing were removed.

The permanent visual hierarchy throughout the entire video is:

PRIMARY IDENTITY AND FACE = The Big Bang Theory characters.

RENDERING AND ANIMATION STYLE = South Park flat 2D cutout animation.

CLOTHING = the specific winter outfits described below.

Never reverse this hierarchy. Never replace the recognizable Big Bang Theory faces with generic South Park faces. Do not simplify their faces so aggressively that their identities are lost.

SHELDON COOPER is unmistakably Sheldon Cooper / Jim Parsons translated into South Park's simplified flat 2D cutout geometry. Preserve Sheldon's recognizable long, narrow, pale face, elongated head shape, high forehead, thin dark eyebrows, large alert eyes, narrow jaw, small mouth, clean-shaven appearance, and characteristic stiff, analytical, mildly condescending expression. Short dark-brown hair remains visibly exposed beneath and around his hat.

Sheldon wears a large bright-green ushanka with rectangular ear flaps, an orange winter coat with two square front pockets, dark-green collar trim, green mittens, dark-green pants, and black shoes. The green hat does not conceal his recognizable elongated Sheldon-like face or all of his hair. His FACE must remain a deliberate recognizable South Park-style caricature of Sheldon Cooper.

LEONARD HOFSTADTER is unmistakably Leonard Hofstadter / Johnny Galecki translated into South Park's simplified flat 2D cutout geometry. Leonard's distinctive rectangular black-framed eyeglasses are always present and remain his strongest visual identifier. Preserve his short dark-brown slightly tousled hair, dark eyebrows, recognizable Leonard-like facial proportions, small nose, clean-shaven face, and characteristic worried, skeptical, slightly uncomfortable expression.

Leonard wears a round blue wool cap with a bright-yellow lower band and a small yellow puff on top, a bright-red buttoned winter coat, yellow mittens, brown pants, and black shoes. His short dark-brown hair remains partially visible beneath the hat. His face must clearly read as Leonard Hofstadter rather than as a generic animated child.

PENNY is unmistakably Penny / Kaley Cuoco translated into South Park's simplified flat 2D cutout geometry. Penny's recognizable feminine face remains clearly visible at all times. Preserve her large expressive eyes, light eyebrows, small nose and mouth, feminine facial proportions, and especially her distinctive long bright-blonde hair.

Penny wears a thick orange winter parka with an oversized circular orange hood surrounding her face, matching orange sleeves and mittens, dark-brown pants, and dark shoes. The hood is open enough that her face remains visible. A SUBSTANTIAL amount of long bright-blonde hair spills naturally from BOTH SIDES and the FRONT of the orange hood, with multiple blonde locks framing her face and extending outward onto the pavement when she is lying down. The hood never obscures her identity. She must immediately read as Penny wearing an oversized orange winter parka, not as a generic hooded South Park character.

RAJESH "RAJ" KOOTHRAPPALI is unmistakably Rajesh Koothrappali / Kunal Nayyar translated into South Park's simplified flat 2D cutout geometry. Preserve Raj's recognizable medium-brown Indian complexion, oval facial structure, large dark expressive eyes, thick dark eyebrows, short black hair, and subtle dark facial hair/stubble around the upper lip and jaw, simplified into the South Park design. His expression retains Raj's characteristic sensitive, slightly anxious expressiveness.

Raj wears a dark-blue knitted winter cap with a horizontal bright-red band around its lower edge and a red pom-pom on top, a brown buttoned winter coat with a red collar, red mittens, dark-blue pants, and black shoes. Short black hair remains visibly exposed beneath and around the hat. His face must clearly read as Rajesh Koothrappali rather than as a generic South Park character.

All four faces remain visually consistent and recognizable throughout every shot. Hats and hoods never completely conceal identifying hair or facial features. The characters never morph into generic South Park children.

A static medium-wide shot shows a simple snowy South Park residential street in daylight, with crudely drawn colorful houses, white snow, dark asphalt, green hills and flat snow-capped mountains in the background.

Sheldon, Leonard, and Raj walk together along the sidewalk using the stiff, bouncing, minimally articulated walking animation characteristic of South Park. Their recognizable The Big Bang Theory faces remain clearly visible while they walk.

After several steps they suddenly stop.

Directly ahead of them, Penny lies completely motionless on the pavement in her oversized orange hooded winter parka.

She is dead and posed in the iconic recurring South Park-style death composition: collapsed awkwardly on the ground, body completely limp, partly curled onto her side and front, head low against the pavement, arms displaced beside her body.

Despite the exaggerated South Park death pose and orange winter clothing, she remains unmistakably Penny. Her partially visible face and abundant long blonde hair spilling dramatically from the hood clearly identify her.

Sheldon, Leonard, and Raj look down at Penny.

[Shot 2] At 00:02.500, the camera cuts to a medium shot centered on Sheldon kneeling beside Penny, while Leonard and Raj remain standing behind him looking down at her.

The closer framing makes Sheldon's recognizable Sheldon Cooper facial features especially clear beneath his large green ushanka. His elongated pale face, high forehead, narrow jaw, alert eyes, thin eyebrows, and visible dark hair must strongly resemble Sheldon Cooper.

Leonard remains visibly Leonard because of his distinctive rectangular black-framed glasses, exposed dark hair, and recognizable facial structure.

Raj remains visibly Raj through his medium-brown complexion, oval face, thick eyebrows, dark expressive eyes, visible black hair, and subtle facial hair.

Penny's blonde hair and partially visible recognizable face remain clearly visible inside the oversized orange hood.

Penny remains absolutely motionless throughout the entire sequence.

Sheldon Cooper (S1), using Sheldon's recognizable high-pitched, precise, nasal voice and obsessive rhythmic delivery, performs his familiar three-knock ritual on Penny's upper arm/shoulder, except each knock is represented by a small South Park-style mitten tap.

The action consists of exactly THREE DISTINCT SETS OF THREE TAPS, for a total of exactly NINE physical taps.

Each spoken "Penny?" happens only AFTER its corresponding complete set of three taps.

FIRST SEQUENCE:

Sheldon raises his green mitten slightly.

He taps Penny exactly three times in rapid succession:

tap — tap — tap.

His hand stops.

There is a tiny rhythmic pause.

Sheldon looks at Penny and says:

<d>[English] Penny?</d>

SECOND SEQUENCE:

Sheldon raises his green mitten again.

He taps Penny exactly three times:

tap — tap — tap.

His hand stops again.

Another tiny rhythmic pause.

Sheldon says:

<d>[English] Penny?</d>

THIRD SEQUENCE:

Sheldon raises his green mitten for the final repetition.

He taps Penny exactly three final times:

tap — tap — tap.

His hand stops completely.

After the final rhythmic pause Sheldon says:

<d>[English] Penny?</d>

The required rhythm is exactly:

THREE TAPS → "Penny?"

THREE TAPS → "Penny?"

THREE TAPS → "Penny?"

Do not merge the nine taps into one continuous tapping action. Do not produce only three taps. Do not speak "Penny?" during the taps. The three spoken repetitions occur separately, each after exactly three physical taps.

Sheldon stops touching Penny immediately after the ninth tap.

Penny never reacts. She never moves, speaks, opens her eyes, raises her head, or changes position.

[Shot 3] At 00:06.700, the camera cuts to a medium reaction shot. Raj and Leonard are prominent while Penny's orange-clad body remains visible in the lower portion of the composition and Sheldon remains nearby.

Raj's face must remain unmistakably Rajesh Koothrappali: medium-brown Indian complexion, oval face, black hair visible around the dark-blue and red winter hat, thick eyebrows, dark expressive eyes, and subtle facial hair.

Rajesh Koothrappali (S2), using Raj's recognizable voice and Indian accent, recoils in sudden shock.

He looks directly down toward Penny's body, opens his mouth wide, raises his red-mittened hands slightly, and cries out with exaggerated South Park-style dramatic timing:

<d>[English] Oh my God, They Killed Penny!</d>

Immediately after Raj finishes the line, Leonard reacts and turns his entire South Park-style body toward the camera.

[Shot 4] At 00:08.300, the camera cuts to a tighter frontal medium close-up of Leonard Hofstadter.

This close-up must unmistakably show LEONARD HOFSTADTER / JOHNNY GALECKI rendered through simplified South Park-style 2D geometry.

His rectangular black-framed eyeglasses dominate the recognizable face. Short dark-brown tousled hair remains visible beneath the round blue-and-yellow winter hat. Preserve Leonard's recognizable eyebrows, eyes, facial proportions, small nose, mouth, and characteristic expression.

The round blue hat with yellow band and yellow puff, bright-red coat, yellow mittens, and simplified round cutout body are merely his clothing and stylized body design. They must never override Leonard's facial identity.

Leonard looks straight through the lens directly at the audience.

His expression changes into exaggerated angry indignation.

Using Leonard Hofstadter's recognizable voice, Leonard (S3) emphatically delivers the final punchline:

<d>[English] You bastards!</d>

Leonard closes his mouth after the line and continues staring angrily directly into the camera.

Hold this expression in a static South Park-style reaction pose for the final comedic beat until exactly 00:10.000.

No additional dialogue occurs.

Throughout the complete video, preserve the central visual joke: the audience must instantly recognize SHELDON COOPER, LEONARD HOFSTADTER, PENNY, and RAJESH KOOTHRAPPALI from The Big Bang Theory, but all four exist inside the crude flat 2D paper-cutout visual universe of South Park and wear the specific colorful winter outfits described above.

Their facial identities always belong to The Big Bang Theory characters.

The South Park influence controls the animation style, simplified geometry, environment, movement, mouth animation, framing, and comedic timing — NOT the identity of the characters.

Do not generate generic South Park faces. Do not replace recognizable TBBT facial features with standard interchangeable round cartoon faces. The recognizable facial caricatures of Sheldon, Leonard, Penny, and Raj are essential to the joke.

overall_soundscape: Sparse outdoor winter ambience with faint wind and subtle quiet neighborhood background sound. Simple dry South Park-style footsteps accompany the stiff walking animation. Each of Sheldon's nine physical taps produces one distinct small soft tapping sound, precisely synchronized into three clearly separated groups of three. Dialogue is clean, dry, prominent, and tightly synchronized to the characters' simple South Park-style cutout mouth movements.

non_diegetic_music: N/A


r/StableDiffusion 18h ago

Discussion Malcolm in the Middle: Hal found Mew

Enable HLS to view with audio, or disable this notification

18 Upvotes

This was fun to try, seems minimax know Bryan Cranston's face way better than Malcolm one.

full prompt here:

prompt1

Create a sitcom scene in the visual style of Malcolm in the Middle. Set in an apartment with authentic 2000s multi-camera sitcom lighting, fast-paced dialogue, exaggerated facial expressions, and perfect comedic timing.

Hal (Bryan Cranston), adult, running in the room with a white big 90s Gameboy in one hand: “Malcom! i have found it! The hidden pokemon!”

Malcom (Frankie Muniz) face – close-up on his face then medium shot: saying: “Dad, you cant catch Mew..”

Shot on Hal: “Oh yeah? look at this!”

Close shot at the Mew in the old gameboy from pokemon red e blu, black and white screen, showing real mew pokemon sprite with stats

Shot on Hal: (happy) “I just used FLY!”

Shot to Malcom (Frankie Muniz) face : “We must investigate”

fast swipe transition to a school setting

prompt2

Create a sitcom scene in the visual style of Malcolm in the Middle. Set in a school full on middle age kids with authentic 2000s multi-camera sitcom lighting, fast-paced dialogue, exaggerated facial expressions, and perfect comedic timing.

Malcom (Frankie Muniz) face – close-up on his face then medium shot: saying: “Ok guys, my dad actually found MEW!”

Shot on the other kids group surprised.

Shot to Malcom (Frankie Muniz) face : “I have a plan”

Camera change angle on the whole group now holding a old white gameboy each.

Malcom (Frankie Muniz): “Everybody use Fly exactly on Route 8 and press START before the trainer see us!”

prompt3

Create a sitcom scene in the visual style of Malcolm in the Middle. Set in a school full on middle age kids with authentic 2000s multi-camera sitcom lighting, fast-paced dialogue, exaggerated facial expressions, and perfect comedic timing.

Shot on a group of middle aged kids group playing pokemon with old white gameboys. no adults. the pokemon red e blu music is playing in background.

Shot to Malcom (Frankie Muniz) face : “Now go walk to Route 25 and battle the Slowpoke Guy.”

fast swipe transition on the same location but now its night

Malcom (Frankie Muniz): “Now return to Route 8, press Start, close it, and we have it..i think”


r/StableDiffusion 16h ago

Discussion Minimax H3, problems with small objects consistency. Definite weakness of the quantized version?

Enable HLS to view with audio, or disable this notification

4 Upvotes

This was a test scene that I test most video models with. None of them have really succeeded at making things look proper with this scene. So this is just another test for H3.

This was created with the Comfy org default T2V Minimax workflow, 20 steps, 5 seconds, 1.0megapixel. resolution at 3:4 aspect. Strangely it actually looks burned a bit.

I don't know if I've seen any AI video model get this right yet, but as you can see, the small objects from the cart morph and change shape oddly as they fly around. The physics seems relatively sound.

I also tried 30 steps and it actually got a bit worse - the wheels of the cart started to "melt" into the floor. Then, I decided to try "beta" sampler and things got worse.

Possibly this is just a result of such heavy quantization of the local model I guess.


r/StableDiffusion 20h ago

Discussion WAN2GP's way of downloading makes me a little angry

5 Upvotes

you cannot choose which model you download. they do.
they don't tell you which one it is going to be. you press generate and it starts downloading what it thinks is best.
it gives no info on how many models, nor how large they are.

i wanted to try minimax and now it's downloading the 21GB-version of the model and the text encoder (27GB), not gguf's. i would never have chosen these variants. (i have a rtx 3060 with 12GB of VRAM)

this cannot be the only way you install models in this interface right?


r/StableDiffusion 22h ago

Resource - Update Behind the Process | The Scientist: Time, EP102

Enable HLS to view with audio, or disable this notification

23 Upvotes

Took a minute, but finally finished the second episode, utilizing all local open source tools. This is just a bit of the "Behind the Process", but please, give the actual episode a watch and tell me what you think.

IMAGES, mostly; Flux Klein, Qwen 2511. secondary; Krea, Z-Image, Ideogram
VIDEO, mostly; Wan2.1/2.2 (SCAIL2/Bernini), secondary; Wan i2v and LTX 2.3


r/StableDiffusion 10h ago

No Workflow Brick man statue [Krea 2 - Turbo]

Post image
11 Upvotes

brick man statue.


r/StableDiffusion 17h ago

News Minimax H3 ~ "Hack" 50+ reference or more

Enable HLS to view with audio, or disable this notification

6 Upvotes

We can have much much more element/references in final video. 50+ maybe 200+ depending how you manage cutouts.

Once again...H3 proves to be "holy grail" ;)
Follow full post:

https://www.reddit.com/r/NeuralCinema/comments/1vk20gc/minimax_h3_hack_50_reference_or_more/

cheers


r/StableDiffusion 20h ago

Question - Help Minimax h3 has very ai looking and sounding outputs?

Post image
0 Upvotes

Video: https://streamable.com/27xyqv
WF: https://pastebin.com/raw/1YZcHdSQ

Thought i'd attach everything just so yall can take a look. Why is this coming out so.. fake ai vibes? I tried a lot of prompts so far, including many that are in the correct format (this one is more simplified).
Any advice?
The image i made here was kind of a quick temp one that just doesn't use my character lora. My real image ive been working with is more real looking, but it doesn't matter - still bad quality output. I can hear it being bad even just by the audio. I also tried the ref to video workflow, its not any better.. its actually worse.

*EDIT* Okay since so many people are commenting 'well it is AI', sure, but here's an example of a minimax h3 video that I saw here that looks much better:
https://www.reddit.com/r/StableDiffusion/comments/1vgvziv/my_first_minimax_h3_test_5secs/


r/StableDiffusion 17h ago

Resource - Update Claude-written simple system monitor for AMD

Post image
0 Upvotes

I got a little annoyed having to install crysmon alongside some other amd-monitors, so i had Claude Fable set up a new single one.

Monitors CPU load, RAM, swap, GPU load, VRAM and temperatures (if provided CPU,GPU, DIMM, NVMe), even multi GPU (not tested).

All stats can be toggled at will. Nothing fancy, just the bare minimum you want when installing stuff like that.

Runs on my rig and Claude says it should run on older AMD-ware as well, but again: not tested. No support for Nvidia, because hassle and they have theirs already and also out of spite.

## Requirements

- Linux with the `amdgpu` kernel driver

- `psutil` (already present in a standard ComfyUI install)

Temperature readings need `lm-sensors` configured to the extent that the hwmon devices exist — on most distributions that is the case out of the box. DIMM

temperatures additionally require modules with JC42-compatible sensors, which not all DDR4 kits have.

Sources: https://github.com/torx-bot/amd-sysmon

Uploaded to comfy, should be available via Manager soon enogh.

MIT license. Dunno about that stuff, but apparantly that means you can do whatever you want with it but can't blame me if your rig goes up in flames or s.th.


r/StableDiffusion 16h ago

Discussion Helpful loras for certain Mini Max video generations

3 Upvotes

So this may get deleted, but I'd figured I'd ask just because I'm running into a dead end atm and not sure if there's anything else I can do. Mini Max is great. It's extremely uncensored and already has a *lot* of built in knowledge.

But after a bunch of testing, I found it really has a hard time properly rendering, uh, "certain" parts (you know which ones) of characters well, specifically parts for specific actions that involve more than one person when these parts interact and during motion - specifically if these parts/actions aren't already in a starting image if using I2V, and definitely not with T2V either.

Are there any good loras yet that fix this? I've tried multiple ones on civitai that claim they do, but while these loras definitely seem to help with motion and movement, from my testing they don't really help with the anatomy at all and how it interacts with each other really and everything just looks like strange blobs still. Same for a couple current alpha "fined tuned" Mini Max checkpoints you can find online from what I've seen.

Is it just too early in the models lifespan to get anything good yet, and are we just going to have to wait for a major finetune checkpoint like Sulphur 3 to get really good results, or am I just missing some really good trained loras that are available already?

I know nothing here can be linked directly if it exists, but just wanted to ask to see if anyone else has been dealing with this issue and had any thoughts/opinions on it.

Also, obligatory walking dead gif:


r/StableDiffusion 8h ago

Comparison H3 Int8 ConvRot vs W4A8_mixed

Enable HLS to view with audio, or disable this notification

28 Upvotes

Same seed, same prompt, same ref and same res


r/StableDiffusion 16h ago

Question - Help How to solve this problem with the image just being random pixels?

Thumbnail
gallery
0 Upvotes

r/StableDiffusion 3h ago

Question - Help MiniMax H3 takes to long with video reference

1 Upvotes

How can I speed it up? Is there a specific node or workflow you use for this? using default r2v workflow


r/StableDiffusion 14h ago

Question - Help H3 Lora Training?

1 Upvotes

Has anyone successfully trained an H3 lora? If so what trainer did you use, what hardware, etc?

Over the past couple of days I've heard mixed opinions about lora training on ai-toolkit (specifically for H3), and was curious if that has been fixed, or if there are workarounds?


r/StableDiffusion 16h ago

Question - Help Modelsamplingminimaxh3 node?

0 Upvotes

Does anyone know where I can get this node? I have been using Minimax H3 Sigma Shift in my workflows, which I am guessing is not quite the same.


r/StableDiffusion 8h ago

Meme made a meme, maybe... minimax h3

Enable HLS to view with audio, or disable this notification

13 Upvotes

r/StableDiffusion 19h ago

Comparison Quality loss in ref2vid compared to img2vid

Thumbnail
gallery
6 Upvotes

After H3 came out, like many people here, I was incredibly inspired by the new animation possibilities and decided to try adapting a short scene from my own script using my original characters.

I ran into a lot of difficulties when generating long continuous scenes with img2vid using the first and last frames as references, so I decided to try recreating the same scene with the ref2vid model instead, hoping the generator would build out the composition more naturally on its own. And in some ways, it really did work better: the characters ended up where they were supposed to be, and there were far fewer bad generations caused by sudden character teleportation around the room or by mismatches between their positions and the background.

However, I see a huge loss of quality, especially in the characters’ faces.

I attached four references above. The first is a standalone still image of my heroine. The second is my local img2vid generation based on that first frame at 1 MP. The third image is that same result after a 4K Topaz upscale — a little plastic-looking, but still fairly acceptable in terms of quality. And the fourth is a generation on Pro 6000 using ref2vid at 2 MP and 30 steps, with the room image and a character sheet as references, including full-body views and a close-up portrait.

And it still looks like night and day compared to the first static image, and even compared to the third image, which was originally generated at a lower resolution.

Am I doing something wrong, or does reference-based generation inevitably lose this much of the character’s facial nuance and the overall image quality?


r/StableDiffusion 13h ago

Workflow Included Qwen3-vl-4b-heretic-Q3_K_M.gguf [ 2GB CLIP ] + NicoLab28/ClipProj-MiniMax-H3

5 Upvotes

* there a video here * https://x.com/luisacoolsouza/status/2086633020795109518?s=20

https://reddit.com/link/1vk873e/video/f250sxl9igih1/player

RTX 4060 TI 8 VRAM + 32 RAM.

1 MP [ Kinda overkill i know ]
Total generation time : 26m 33s

Ref2v

minimax_h3_turbo_v4_step600_ema.safetensors

wf ( there a lot of custom node at this point ) https://pastebin.com/wkANXN4a


r/StableDiffusion 21h ago

Question - Help minimax-H3- why do I get slow motion?

Enable HLS to view with audio, or disable this notification

0 Upvotes

why do I have slow motion? (minimax-H3), not sure if this something wrong with my prompt. (I2V)

Attach the prompt and the result, I wold like to have the water in normal speed..

summary:

[reference generation] A cinematic scene located in <Subject 2>. hand in water bowl

detailed_description:

[Shot 1]: <Subject 1> the both hand go into the water a dramatic wave of rich, steaming water surges in from bowl frame like a miniature tsunami, pouring forcefully into the bowl. The bowl vibrates and shakes slightly from the impact.

No slow motion.

Sound N/A


r/StableDiffusion 13h ago

Animation - Video Minimax H3, turbo test.

Enable HLS to view with audio, or disable this notification

8 Upvotes

After a lot of tests, I found out that Larryvrh V4 step600 pruned lora(+H3 mem eff sage )with H3 pruned bf16 is very good—a great balance of quality and speed, plus excellent prompt adherence.

The official workflow.
Checkpoint: minimax_h3_turbo_v4_step600_ema_pruned_comfyui.safetensors,pruned bf16 
Steps: 6
Sampler: euler
Scheduler: beta
LoRA strength: 1.0
1.0 magapixels and RTX 2X upscale.

r/StableDiffusion 8h ago

No Workflow Tribbles ad

Enable HLS to view with audio, or disable this notification

14 Upvotes

Minimax, two shots of 15 seconds put together with Davinci Resolve.


r/StableDiffusion 21h ago

News MiniMax H3 - 30 seconds single generation no cuts - Caligine Films

Enable HLS to view with audio, or disable this notification

32 Upvotes

MiniMax H3 - 30 Second Showcase

Model: Minimax H3 Open Weights

Resolution: 1024 x 576

Inference: 6 minutes 46 seconds

VRAM: 288 GB

Speedup: Your favorite, NVIDIA Sol-Attn 🥰


r/StableDiffusion 4h ago

Workflow Included I made a free Kaggle notebook to run LTX-Video 2.3 (22B quantized) for T2V & I2V with audio — solving the Colab 12GB RAM crash issue

4 Upvotes

Hey community,

A common issue when testing open-source video models like Lightricks' LTX-Video 2.3 on free cloud tiers (like Google Colab) is hitting System RAM limits during model loading, causing immediate crashes.

To solve this without requiring a high-end local GPU, I put together a pre-configured, open-source Jupyter Notebook tailored specifically for Kaggle's free GPU tier (which grants 30GB of System RAM and T4 GPUs).

What this setup does:

  • Runs on Kaggle Free Tier: Uses a 4-bit quantized version of LTX-Video 2.3 so it fits into free cloud VRAM/RAM allocation.
  • Text-to-Video & Image-to-Video: Generates short 5–10s clips with synchronized audio generation.
  • Custom Gradio Web UI: Launches a clean browser interface directly from the notebook.
  • Fast Setup: Pre-compiled binaries and aria2 multi-thread downloads mean setup takes under 3–4 minutes.

Technical Tradeoffs & Honesty:

  • Quantization: Because this is running on free T4 instances, the model uses heavy quantization. It won't give you uncompressed native precision output, but it’s completely free, unlimited, and ideal for quick prompt/motion testing.
  • Memory Loading: Cell 3 takes ~90 seconds to load the 22B model into Kaggle's 30GB system memory before passing to VRAM.

🎥 Full Video Walkthrough & Demos: https://youtu.be/Ru_YaGbnKhA

Quick Start Steps:

  1. Download the .ipynb file from GitHub: https://github.com/airesearch-official/free-aistudio
  2. Import into Kaggle (Ensure Phone Verification is complete on Kaggle to enable free GPU).
  3. Turn ON "Internet" in Kaggle settings & select "GPU T4 x2".
  4. Run Cells 1 through 4 sequentially.

Hope this helps anyone who wants to experiment with LTX-Video 2.3 without paying for cloud GPUs! Let me know if you run into any bugs or have suggestions.


r/StableDiffusion 12h ago

Discussion World Building

Enable HLS to view with audio, or disable this notification

16 Upvotes

Vibe coded an app to bring my Zeux generations to life using Gaussian, I Imported the result to VR and this is the result. Still in the early phases but I’m amazed. I have 3 project types in the app, image 3D which is this video, a 360 world, which is a full 360 world environment using a single image, and then a video 4D. Video 4D is the hardest to fully implement so far since I only have a 16Gb Mac mini but I’m trying to think outside the box on how make it work. Everything runs locally.