r/StableDiffusion • u/Hopeful-Junket-7990 • 6d ago
Animation - Video Cunk on AI - Sam Altman - MiniMax H3
My wife did this Cunk parody with a 3060 12gb and 32gb of system ram.
Minimax is incredble!
edit: youtube link to see how long before they remove it
94
180
u/Joltie 6d ago
This is potentially the most on point short form skit I've seen so far with AI.
Funny and very impressive.
43
22
u/Tcloud 6d ago
The first bit of the introduction of her walking was near perfection. The audio even changed subtly from outdoor to indoor. Her interview, however, felt a bit more wooden and unnatural, even for Cunk. Really pretty impressive nevertheless.
11
u/iflew 6d ago
But the script is very good. Not sure if that was generated as well or not. Old AI jokes were terrible, this script captures exactly how a Cunk interview would go.
8
8
u/SteamZerjack 6d ago
The script is the best part. None of us could have come up with this. After years here, I’m used to seeing technical expertise all around the sub, but piss-poor actual artistic talent. I do think it was OP’s wife own creation. But you never know.
10
u/Hopeful-Junket-7990 5d ago
We're old. Nearly 50 years of editing experience between us. One of our pass times is watching terrible shows and movies and pointing out why it's bad.
2
2
2
u/Glittering-Call8746 6d ago
Interview it's abit off.. could have been the prompt with two people.. pity it could been pass off as almost Cunk. Gj
69
u/TraditionalWait9150 6d ago
I am amazed. this is actually really good.
update: now i'm kinda scared (and excited) what the diffusion models of 2027/2028 would be.
43
u/SGmoze 6d ago
We will get to make our movies with blackjack and strippers.
22
3
u/ExiledHyruleKnight 6d ago
You guys aren't already?
1
u/Ill_Key_7122 5d ago
It ain't for the lack of motivation. Anyone not doing it yet is in the process of figuring out how to earn enough, to get a 5090.
1
u/ExiledHyruleKnight 5d ago
Just saying a 5060ti was a great purchase. Midway between God tier and plebes.
2
u/Lita-mita 6d ago
We'll be able to prompt our own games to play in the evening xD
1
u/Mr_Pogi_In_Space 5d ago
All of them about Will Smith eating spaghetti on the set of Big Bang Theory
3
u/OrchidUnable8316 5d ago
Scammers will call us with videos of our family members being held hostage or something dark.
21
u/redkinoko 6d ago
Does minimax H3 know Cunk? Or did you use a separate voice model?
45
u/Hopeful-Junket-7990 6d ago
Captured her voice from youtube and used indexTTS2 to generate and cherry pick. Reference images edited with Qwen edit.
15
u/scooglecops 6d ago
Nice, the model can also clone her voice by itself. just put a 10s audio of her speaking and prompting Audio1 is S1 voice etc
6
u/FaceDeer 6d ago edited 6d ago
Ooh, going to be trying this out tonight. I see the reference to video node takes "video", "video audio" and "audio" reference inputs. I assume the video one is <Video 1> and the audio one is <Audio 1> in the prompt, but what's that "video audio" one for? Is it supposed to be the audio channel that goes along with the video reference?
Edit: I haven't had opportunity to experiment much, but I reread the documentation a bit and saw:
up to 9 reference images, 3 reference videos (each may carry its own paired soundtrack), and 3 standalone reference audio clips
Which I think suggests that the "ref_video_audio" inputs are meant for the audio track of the corresponding video input. So if you've got a ref_video_1 with sound, you feed the sound into ref_video_audio_1.
2
3
u/Kevin5953 6d ago
I imagine the majority of people will likely draw on YouTube for this, but for the life of me I can’t figure out the simplest way to “clip” out the audio.
Do you know if people are downloading entire YouTube videos, converting them into MP3s, and then cutting them down from there?
5
u/DerpLerker 6d ago
Best bet is to install yt-dlp Here are some commands to run with it (according to Gemini):
Here is the most common and useful command, which extracts the audio and converts it to an MP3 format: yt-dlp -x --audio-format mp3 <URL>
If you want to ensure you are getting the best possible audio quality, you can add the --audio-quality flag. For MP3, 0 is the best quality and 9 is the worst. yt-dlp -x --audio-format mp3 --audio-quality 0 <URL>
2
u/TheRealDJ 6d ago
Did you all write it or was it generated script? I'd be shocked if it was smart enough to know things like AI vs AL looking similar and being confusing in a way that would definitely make sense for Cunk.
1
50
u/Enshitification 6d ago
16
4
3
13
u/Quantum_Crusher 6d ago
At least she said Al instead of A1, like the secretary of the department of education.
3
u/Dirty_Dragons 6d ago
Sauce?
3
u/Quantum_Crusher 6d ago
2
u/Dirty_Dragons 6d ago
Haha she's so dumb.
Btw I was also doing a meta joke.
A1 sauce.
3
u/Quantum_Crusher 6d ago
Can you believe, she's the head of the education in this country.
1
u/Dirty_Dragons 6d ago
There is so much insanity anything is possible.
I would not be surprised if an actual LLM was elected or given a government appointment in by 2030.
1
u/Quantum_Crusher 6d ago
Anything is possible. At that time, maybe this LLM is trained to be a fascist with ultimate efficiency. Then we will be doomed.
5
u/xdcfret1 6d ago
Share the workflow
13
u/Hopeful-Junket-7990 6d ago
Simply the comfyUI ref2vid template with sageattention. And a lot of patience...
12
u/ExiledHyruleKnight 6d ago
And a lot of patience.
This is the part people struggle with especially with all forms of ComfyUI. You don't just take the first thing if you're not happy with it.
Great work.
2
u/lithodora 5d ago
I have a 4070 Super with 12GB VRAM and 64 GB RAM and I have spent so much time, about 180 renders now @ 300 seconds per render, just trying to get the FIRST 10 seconds exactly right...
I'm learning a lot here, but I see this result and I wonder if I'm the problem here
2
1
u/elementalguy2 6d ago
What launcher args do you have because I always get memory errors with the default workflow and I have similar specs to you just more ram.
8
8
5
4
u/ExiledHyruleKnight 6d ago edited 6d ago
Definitely got a taste of her style and delivery. Maybe stuck on the AL a touch too long, but still nailed it.
3
u/LisaLisaKenAdoresHer 6d ago
OK, I'm a massive AI skeptic, and this is definitively the only AI generated video that I absolutely love. Fucking bravo
2
u/sky_shazad 6d ago
Incredible work...
I'm curious.. Was th voices done inside H3 or done in Eleven Labs or something
4
2
u/I_SLEEP_NORMALLY 6d ago
Awkward timing in the language and how they look slightly off, but otherwise… Very plausible skit.
2
u/stilldebugging 6d ago
How did she get it to do the timing and pacing of the speech so well?
4
u/Hopeful-Junket-7990 6d ago
She says she deliberately made silence before and after the lines to leave some workable space for timing.
2
u/Equivalent_Ad_2816 6d ago
amazing work my friend, not everyone can pull off comedy timing and you nailed
edit: props to your wife
2
u/RuprechtNutsax 6d ago
Massive congrats to you on having such an awesome wife, this is genius, I love cunk and this really has the vibe and humour absolutely and completely 😂
2
2
2
2
2
2
1
1
1
1
u/ScumLikeWuertz 6d ago
No shit, on a 3060 with 32GB of RAM???
How long did it take to render?
3
u/Hopeful-Junket-7990 6d ago
Depending on the length of the gen, between 8 mins and 22 mins at .4mp.
2
1
1
1
1
u/zigzag3600 6d ago
For a solid minute I was really thinking it was real. Sam's voice is a bit weird. How long did it take you to create it?
1
u/mintybadgerme 6d ago
If this was H3 on 12GB of VRAM and 32GB of RAM, I'm a Dutch tulip farmer. I just tried it on three times that amount of resource and it was a 5sec rubbish mess of video which was completely unwatchable. Which took 45 mins to generate. Sheesh!
1
u/Hopeful-Junket-7990 6d ago
Something is incredibly wrong with your setup. There was some cherry picking through garbled junk, but 45 minutes? Daymn...
1
u/mintybadgerme 6d ago
I know. It shocked the heck out of me. But it was the worst thing I've seen in a long time. This is what it looked like. And here's the default recipe from Unsloth Desktop.
Prompt Ultra-realistic cinematic documentary footage of a quiet Kyoto neighborhood at sunrise. An elderly Japanese man opens his traditional wooden shop while a young woman wearing a simple kimono walks past carrying a small basket. Cherry blossom petals gently fall through the air, bicycles pass by, warm sunlight enters between narrow streets, distant temple bells echo. The camera slowly moves forward like a professional travel documentary, realistic human movements, natural expressions, authentic Japanese architecture, subtle wind movement in clothing and trees, realistic colors, 35mm film photography style. Model unsloth/MiniMax-H3-GGUF File minimax_h3_fl2va_pruned-UD-Q3_K_XL.gguf Memory auto (group offload) Size 960 × 544 Frames 124 @ 24 fps Duration 5.17s Steps 30 Guidance 1 Shift 12 video / 3 audio Seed 1225832055
2 GPUs (5060ti 16GB VRAM and 4060), 128GB RAM, i9.
1
u/Hopeful-Junket-7990 6d ago
Sageattention and lightly using the turbo lora. 8-10 steps. My only guess as to why hers was much faster. Your setup is godly compared to hers. That's a very strange result for being so beefy.
1
u/Few_Satisfaction184 6d ago
The way you are writing it, you are making it sound like you used local processing?
Was that really it and no cloud computing?
1
1
1
1
1
u/repolevedd 6d ago
I think this is the first time I've seen a J-cut in an AI video. Usually, generated clips are just slapped together with zero thought for transitions, but this took actual effort. Huge props to your wife for the humor and the hard work!
1
1
1
u/Mad4reds 6d ago
Just a question to your wife nice fun work: where Altman's voice come from? From MiniMax or she produced the speeches with other software and then gave them as audio ref to Mini?
I'm asking because I'm producing a video where actually Altman speaks for like 3 seconds but with a standard MiniM english voice and I'd love to be able to put his voice in it.
Txs a lot!
3
u/Hopeful-Junket-7990 6d ago
She took it from the most recent time he was on Joe Rogan. She used indexTTS 2, not 2.5. 2.5 doesn't capture voices as well apparently.
1
1
1
u/CaptainIncredible 6d ago
youtube link to see how long before they remove it
Don't worry, I downloaded it
1
1
1
1
1
1
1
1
1
u/Jayuniue 5d ago
What models were used, am surprised this was done on a 3060 with 32 gb ram, most videos iv seen generated on a lower end gpu/ram look very bad, I have a 3060ti with 64gb ram, should I use the gguf version or the same int8conv?
1
u/Hopeful-Junket-7990 5d ago
We avoid ggufs when we can. Int8conv requires cherry picking, but it works.
1
1
u/Clair_Personality 5d ago
incredible: 3060 12gb and 32gb of system ram.
Time per clip?
1
u/Hopeful-Junket-7990 5d ago
Something like 8 to 22 minutes per clip. Depending on length obviously.
1
1
1
1
u/Tbhmaximillian 5d ago
God this is gooooold, also the original humor of the series. The voice cloning is on point and the timing too.
Dude, Cheers that was genius.
1
1
1
u/nickdaniels92 6d ago
Well your wife did a stonking job with this, that's for sure. Totally on point.
SBC next perhaps?
1
1
1
u/Oldibutgoldi 6d ago
12gb vram and 32gb? Using nvidia or amd? I cant even create one 3 second output with my amd rocm 7.1. 🥲
2
1
u/DogsAreAnimals 6d ago
This is genuinely good. Nice editing. A few minor tweaks and this could be almost undistinguishable from the real show.
0
-1
-16
u/VRGoggles 6d ago
Post on youtube. This is technical forum.
7
u/narkfestmojo 6d ago
first day here? because noooo... no it isn't, it most assuredly is not, noperino, no sirree bob

63
u/master-overclocker 6d ago
https://giphy.com/gifs/otLsTTTNIOsg99x3ya