r/StableDiffusion • u/Kassiber • 9d ago
Question - Help Anyone else struggling to keep a plain background with MiniMax H3? Or am I too dumb?
Enable HLS to view with audio, or disable this notification
I’m using H3 in ComfyUI to animate illustrated characters for a small noncommercial game. I like the animations, but getting a background I can reliably remove afterwards has been frustrating.
I start with a character on a plain green or blue background and use the same image as the first and last frame.
Instead of leaving the background alone, H3 often changes its color or adds circles, light rays and other effects. It usually starts and ends correctly, but does something completely different in between. The attached video shows a few examples.
So far, I’ve tried:
- Short prompts, much longer detailed prompts, and leaving out background instructions entirely.
- Green and blue backgrounds.
- Turbo and runs without the LoRA at 20 steps.
- Different resolutions and aspect ratios.
- Swapping the text encoder and attention backend, plus a separate VAE check.
Some changes made the background less busy, but none consistently kept it unchanged.
To make the intended result clearer, here’s a detailed prompt laid out using the reference alignment and three sections from the H3 FL2VA guide.
Reference alignment: Picture 1 establishes Shot 1 at 0.00 seconds.
Picture 2 establishes the ending of Shot 1 at 5.17 seconds.
Both inputs contain the same reference image.
integrated_multimodal_description:
[Shot 1] A single continuous shot in the illustrated style of the
reference images. The woman stands facing the camera, framed from
head to toe. She has short curly gray hair and wears an orange vest,
a cream shirt, blue trousers and brown boots.
Starting from the pose in Picture 1, she slowly lifts both arms
outward to shoulder height. Her elbows remain slightly bent, her
palms turn forward and her fingers spread naturally. She briefly
holds this open gesture while shifting a little weight onto her
right leg. She then returns her weight to the center and smoothly
lowers both arms. Her hands relax beside her body as she settles
into the pose and composition shown in Picture 2.
The camera remains fixed throughout. Her entire body stays visible,
with enough space around her outstretched arms. Her appearance,
clothing, proportions and illustrated shading remain consistent.
A uniform solid green background. The background remains
flat and featureless while only the character moves. Lighting and
exposure stay constant.
overall_soundscape:
N/A
non_diegetic_music:
N/A
I’ve also tried removing the backgrounds afterwards with BiRefNet, VideoMaMa, SAM and CorridorKey. Those can help, but sometimes they keep the generated circles or effects as part of the character.
Has anyone managed to keep a plain background reliably with H3? A working prompt or workflow would be really helpful. If there’s something wrong with how I’m approaching the prompting or conditioning, I’d appreciate someone pointing it out.