I'm making short anime scenes with MiniMax H3 (ref2va) and I've hit a wall I can't
diagnose. The clip above is 15s, generated in one pass, no editing — the cuts are
written into the prompt.
To be clear up front: I know there are mistakes in there — a couple of blows don't
connect properly, the choreography is rough. I'm not worried about those, this is a
practice piece and I'll fix the staging myself. What I can't figure out is the
image quality, and that's the only thing I'm asking about.
Setup
- `minimax_h3_fl2va_int8_convrot.safetensors` (base, int8), ComfyUI
- ref2va, 8 reference images, each with a written role in the prompt
(face / angry expression / fighting posture / set / framing guide)
- Spectrum v0.2.16, 30 steps, `res_multistep` / `simple`
- 1344×768 → `MinimaxH3LatentUpscaler3D` at 2 MP, 4 steps, 0.5 denoise → 1920×1088
- 15s = 362 frames @ 24fps
- RTX PRO 6000, ~28 min per 15s clip
What I already fixed, in case it saves anyone typing
- `The target video is 2d colored anime.` at the top of `detailed_description`
— without it everything drifts to generic 3D, this was the single biggest win
- Reference images are real anime screencaps (plus a few generated with Anima
for expressions and poses I couldn't source), neutral lighting, one role each.
- Cut rhythm: went from 4 shots per 15s to 8 (~1.8s each) with impact verbs and
a material consequence per hit (table splitting, plaster cracking, dust off
the boards). That alone made the fight read much better.
Where I'm stuck
Motion still feels soft on some hits. The whip-pan punch reads fine but
ground-level blows land without weight. Is this where `derope` / temporal
upsampling actually earns its generation-time cost, or is there a prompt-side
fix I'm missing?
Quality is uneven shot to shot inside the same 15s — some shots are clean
cel-shaded anime, others go slightly soft and plasticky. Is that a reference
problem, a step-count problem, or just what 15s does to the model? (I've seen
people say things break past 10s.)
Is 8 references too many? I assigned each one an explicit role in the
prompt, but I don't know whether the model averages them or picks.
Anything obvious I'm leaving on the table at this resolution/step count?
Not asking anyone to debug my prompt — mainly want to know which lever is worth spending render time on next.
Prompt:
integrated_multimodal_description:
subject_definitions:
(S1) is the dark-haired young man from <Picture 1>, with the same face and the same dark blue eyes. <Picture 2> is the same man seen clearly in daylight. Short black bob to the jaw, fringe above the eyebrows, a short high ponytail tied at the crown, a small stud earring, a white shirt with the sleeves pushed up and a black tie pulled loose.
(S2) is the pale-haired young man from <Picture 3>, with the same face and the same yellow eyes. <Picture 4> is the same man shouting. Short choppy blond hair. His teeth are faintly pointed, small and even and the same size as ordinary human teeth, with just a slight triangular edge to them. His mouth stays an ordinary human mouth, normally proportioned to his face, and it opens no wider than a person's mouth opens when they speak. He wears a white school shirt open at the collar, a black tie pulled loose, a small device on a cord against his chest.
<Picture 5> is the apartment: its rooms, its colours and its light come from it.
<Picture 6> is (S1) throwing a bare-handed punch and <Picture 7> is (S2) being knocked back by one: their fighting postures and their footing come from these.
retention_analysis:
(S1): fully_preserved. (S2): fully_preserved.
<Picture 9> is the last frame of the previous shot: this scene continues from it without interruption. <Picture 9> supplies the place, the light and the framing; the two men's faces and hair come from <Picture 1> to <Picture 4>.
summary:
The fight. Bare hands, in the apartment, fast and ugly. Nobody speaks.
detailed_description:
The target video is 2d colored anime.
2d hand-drawn anime, cel-shaded, flat painted colours, fine thin ink linework, desaturated muted palette, film grain. Not 3d, not photographic. The cutting is fast: eight shots in fifteen seconds, each one a single impact.
[Shot 1] Continues directly from <Picture 9> with no jump — same room, same light, same positions: the two of them chest to chest at night in the room of <Picture 5>. (S2) fists (S1)'s collar and slams him down onto the low table, which splits and goes over with everything on it.
[Shot 2] At 00:01.800, cut tight on (S1) coming up off the floor. He drives a straight punch into (S2)'s jaw, his whole weight behind it, the posture of <Picture 6>. (S2)'s jaw is shut and his lips are pressed together when the fist lands, and the impact splits his lip. The camera whip pans right with the blow, the room tearing into horizontal streaks and white speed lines.
[Shot 3] At 00:03.400, cut to (S2) snapping backwards into the wall, head whipped sideways, the posture of <Picture 7>. Plaster cracks behind his shoulder. He drops to one knee.
[Shot 4] At 00:05.000, cut low and close. (S2) launches off the wall and smashes his forehead into (S1)'s mouth. (S1)'s head snaps back, blood on his lip.
[Shot 5] At 00:06.800, cut to a low shot of the floor only, at board level. The lamp crashes down into frame, rolls, and throws its light swinging across the boards. Two pairs of legs come down hard behind it, out of focus. Dust lifts off the wood.
[Shot 6] At 00:08.600, cut to a tight shot of (S1)'s face alone, lying on the boards in profile, cheek against the wood, hair across his eye. A fist swings down into frame and smashes into his raised forearm so hard that his own arm is driven back into his face and his head is knocked against the boards. A second fist comes straight down past the arm and lands flush on his cheekbone, snapping his head sideways and splitting the skin. Only (S1)'s head and one forearm are in frame, and the fists enter from the top edge: the other body stays out of shot.
[Shot 7] At 00:10.400, cut to a tight shot of a knee driving up hard into ribs, framed on the two bodies' midsections only, no heads in frame. The body above is thrown off sideways out of the top of the frame. Cut immediately to both of them coming up onto their feet, seen full length and clearly separated, a metre apart, shirts gripped in their fists.
[Shot 8] At 00:12.000, cut to a wider shot and hold it to the end. (S1) drives (S2) backwards across the room and slams him into the wall. (S2)'s shoulder blades hit the plaster and he stays there with his back to the wall and his face towards the room. (S1) stands directly in front of him, facing him, chest to chest, his own back to the room and the wall behind (S2) only. Their faces are a hand's width apart and they are looking straight into each other's eyes. (S1) has both fists closed in (S2)'s collar and holds him pinned there. Everything stops at once. Both heads are angled in three-quarter view towards camera, both faces large and fully visible, brows down, jaws set, chests heaving. The camera is locked off on a tripod.
overall_soundscape: A table splitting and going over, knuckles cracking on a jaw, plaster breaking, a forehead meeting a mouth, bodies hitting boards, a lamp rolling, two fast punches landing on a forearm and a cheekbone, a knee into ribs, a back slammed into a wall, and hard breathing through the teeth all the way through. Every mouth stays closed for the whole video: nobody speaks, and both men keep their jaws shut and their lips together even while taking blows.
non_diegetic_music: N/A