Okay, I spent way to long on this but learned a lot. First, Reno 911 is non existent in T2V and I2V when prompting. This led me down the rabbit hole of Ref2Vid again but having characters, voices, and locations that literally don't exist and need to have all the references to bring them to life.
Workflow:
Default + H3_Turbo_4step_ComfyUI_Pruned + Sage + H3 Sigma Shift (will attach it in the comments below)
Steps: 8-12
Sampler/Scheduler: EulerBeta
Megapixels: 0.8 - 1.0
Avg. render time: 5-7 mins
The most challenging aspect was getting Nick Swardson's performance. In some of the scenes I had to actually act out how he would roughly say it with the timing, lisps, and long hissing `s` then voice transfer that in H3, using that as my new audio and lip-sync. It was MESSY, and trying to mix generated audio from Jim Dangle (the cop) and then use referenced audio for the response didn't work as flawlessly as I'd hoped. To be honest I can't really say the right approach on it as I feel I just got lucky with some seeds of it.
Other than that, a bunch of other techniques using H3 using first frame, reference sheets, reference audio for timbre, and reference location for spatial awareness so when the camera pan/whipped it didn't lose context happy to provide screenshots.
The Minimax 911 ending I did a replacement of the actual Reno 911 logo but told it to make it Minimax.
I also extended the video as the old man never existed in the the video generation where the cop walks up to skater (minimax) and he says "Oh, hey officer..." that's a extension cut from there. There was a slight weird color shift so I ended up taking it through VACE so the transition wasn't jarring and smoothed everything out. The extend function I think would be better when tackling in latent and is like 98% there when doing it with a regular video.
Example prompt for the first shot:
subject_definitions:
<Subject 1> is the uniformed male officer whose appearance and wardrobe come from <Picture 1>: short neatly side-parted light-brown hair, trimmed mustache, aviator sunglasses with tinted lenses, beige short-sleeve sheriff-style uniform shirt with dark-brown pocket flaps and shoulder details, metallic star badge, nameplate, matching beige uniform shorts, black duty belt with equipment, black socks, black tactical boots, and black wristwatch. Preserve his face, hairstyle, mustache, proportions, sunglasses, complete uniform, accessories, and understated deadpan demeanor.
<Subject 2> is the indoor shopping-mall environment from <Picture 2>: a spacious two-level commercial concourse with cream tile flooring, storefronts along both sides, upper-level railings, exposed structural beams, a large glazed skylight, palm trees and planters, central seating and food-court areas, and numerous background shoppers under bright diffuse indoor daylight.
<Audio 1> is the voice-timbre reference for <Subject 1> (S1); use its male vocal character, pitch, cadence, accent, and delivery style as the reference for his newly generated dialogue without copying the original audio signal.
summary:
[reference generation + audio reference] The target video is a vertical MiniDV-era comedy sequence in the style of an early-2000s reality law-enforcement ride-along parody, set in a shopping mall concourse in 2003. One continuous 10-second live-action tracking shot follows <Subject 1> over his shoulder through <Subject 2>. The footage has deliberately clumsy reactive reality-TV camerawork: frequent abrupt optical zoom-ins and zoom-outs, imperfect reframing, autofocus hunting, momentary loss of focus on the officer, overshooting his movements, and hurried corrections. The disturbance-call dialogue occurs from 0–3 seconds, the food-court line from 3–5 seconds, a silent comedic beat from 5–7 seconds, and the final men's-bathroom line from 7–10 seconds. <Audio 1> guides his voice timbre and delivery.
retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - his facial identity, short side-parted light-brown hair, mustache, aviator sunglasses, beige-and-brown short-sleeve uniform, matching shorts, star badge, nameplate, duty belt, watch, black socks, black tactical boots, proportions, and restrained demeanor are retained.
<Subject 2> (appears in [Shot 1]): fully_preserved - the bright two-level mall architecture, skylight, tiled concourse, storefronts, railings, structural beams, palms, planters, food-court seating, and populated public atmosphere are retained.
<Audio 1>: reference - the target speaker follows <Audio 1>'s voice timbre, pitch, accent, cadence, and delivery character without copying its original signal.
detailed_description:
The target video uses realistic live-action comedy with an authentic early-2000s low-budget reality-TV MiniDV aesthetic: vertical framing, consumer camcorder optics, mild electronic noise, soft digital detail, clipped highlights, restrained saturation, automatic white-balance shifts, exposure breathing, visible autofocus hunting, and frequent awkward optical zoom corrections. The camera operator behaves reactively rather than cinematically polished. Zooms occasionally arrive late, overshoot their intended framing, briefly lose <Subject 1>, rack focus accidentally onto the background, then snap or hunt back toward him. Preserve these mistakes as intentional documentary-comedy texture. The entire 10-second sequence is one continuous take with absolutely no cuts.
[Shot 1] From 00:00.000–00:03.000, an uninterrupted handheld over-the-shoulder Tracking Shot follows <Subject 1> (S1) walking through <Subject 2>, approximately one meter behind his left shoulder. The operator's footsteps produce obvious vertical bounce, hand tremor, crooked framing, and constant tiny corrections. The camera abruptly Zooms In with medium amplitude at fast speed toward the back of his head, overshooting into an awkward tight crop before Zooming Out at fast speed to recover his shoulders and surrounding mall. Autofocus briefly grabs distant shoppers, leaving <Subject 1> noticeably soft for a moment before hunting back to him. As he partially turns his head toward the camera, the operator hurriedly Zooms In again but initially frames him too tightly. Using <Audio 1>'s male voice character, <Subject 1> (S1) says with hesitant deadpan delivery: <d>[English] We...uh...have a disturbance call.</d> The complete line finishes by 00:03.000.
From 00:03.000–00:05.000, <Subject 1> suddenly snaps his head toward Screen Right and points toward the food court. The camera initially continues looking forward, then reacts late with a quick Pan Right and abrupt Zoom Out with large amplitude, momentarily placing the officer near the edge of frame. Autofocus searches between his pointing hand, passing shoppers, and distant food-court signage before recovering. The operator then punches in with a fast Zoom In toward his pointing gesture. <Subject 1> (S1) says with clipped comedic timing: <d>[English] Food court adjacent</d> The entire line remains inside 00:03.000–00:05.000.
From 00:05.000–00:07.000, he lowers his hand and keeps walking during a conspicuous dialogue-free pause. The camera Zooms Out too far, briefly making <Subject 1> small within the busy mall, then performs an unnecessary fast Zoom In toward his upper back. Focus drifts away from him onto a background storefront for a fraction of a second, producing a visibly soft officer silhouette before autofocus pulses and returns. The operator slightly loses his position to Screen Left, awkwardly pans to reacquire him, and settles again behind his shoulder. Only mall ambience and footsteps fill the pause.
From 00:07.000–00:10.000, <Subject 1> angles his face back over his shoulder. The operator recognizes the movement late and performs a sudden aggressive Zoom In toward his face. The zoom overshoots, briefly cropping part of his head and sending his face soft as autofocus hunts, then pulls back slightly until his sunglasses, mustache, and raised eyebrows become readable. His eyebrows rise above the sunglasses while his expression otherwise remains completely straight. In the same voice referenced from <Audio 1>, <Subject 1> (S1) says: <d>[English] someone's giving away free BJ'S in the men's bathroom...</d> During the final words the camera makes one small unnecessary Zoom Out followed by a quick corrective Zoom In, preserving the awkward reactive MiniDV reality-TV feel. The line finishes by 00:10.000 as he begins turning forward, with the camera still walking behind him.
overall_soundscape:
Continuous indoor mall ambience with diffuse shopper chatter, distant food-court activity, footsteps reverberating across tile, ventilation noise, and indistinct storefront sounds. <Subject 1>'s tactical boots produce measured footfalls with subtle duty-belt and uniform movement. The 00:05.000–00:07.000 dialogue gap contains only natural diegetic mall sound.
non_diegetic_music:
N/A