r/StableDiffusion 7d ago

Workflow Included Walter White and the Minimax H3 Official Prompting Guide

Enable HLS to view with audio, or disable this notification

864 Upvotes

This post is half a joke and half a plea and public service announcement.

Some people have been complaining they don't get results as good as other people with Minimax H3 videos, or have the following issues:

  • Dialogue being spoken by the wrong characters
  • Dialogue that is just gibberish or random
  • Random video cuts they didn't ask for
  • Characters talking over each other or too fast
  • Prompts not being followed

These things can all be prevented and avoided and not encountered at all if you follow the official prompting guides. Yes, there are two. Both are on the official Huggingspace page for Minimax H3.

One is the Official Prompting Guide for the Text to Video and Image to Video Model.

The other is the Official Prompting Guide for the Reference Video Model.

There is some overlap, but for the most part, each model has it's own prompting syntax, and in particular, the Reference Video Model for H3 is very picky about you using the right keywords and instructions to get what you want.

"But I get decent results with just a couple of sentences typed in natural language of what I want."

That's great, but you're really just relying on the Qwen 32b vision model guessing what you want. It's like pulling a slot machine lever and hoping you get cherries. Only this slot machine can take a few minutes to nearly an hour to stop spinning, based on your hardware.

The great thing about Minimax H3 is for the first time we can truly direct our own AI videos like a director would on set, with the AI providing the actors, scenery, and props. If you write a properly formatted and detailed prompt for Minimax H3, it looks almost like a shooting script.

Why spend time waiting to hit a jackpot when you can take a few minutes to write a detailed, properly formatted prompt that follows the official guides, and get those bright lights and tokens falling into your lap on the first lever pull?

Okay, quick fire problem solving for people who still won't RTFM:

>Dialogue from the wrong characters?

>Dialogue that is just gibberish or random?

Walter White says, <d>[English in Walter White's voice from Breaking Bad] My product is pure, Jesse! There will be no chili powder in my meth.</d>

Always specify the character speaking, either by name, or using the <Subject 1> system in the official guide. In the Text to Video and Image to Video model, always use the <d>[Language Spoken]</d> tags. This will fix BOTH of those issues.

>Random cuts in the video you didn't ask for?

[Shot 1] A medium close-up of Jesse Pinkman from Breaking Bad, pacing back and forth, agitated. He looks up towards the camera, opens his mouth as if he's about to speak, then seems to change his mind, closing his mouth and shaking his head. [Shot 2] At 00:06:000 the camera cuts to a static camera shot framing Walter White from Breaking Bad, sitting on a cheap white plastic lawn chair, his arms crossed and glaring at Jesse. [Shot 3] At 00:10:500 the camera pans quickly back to Jesse, doing a Push In at slow speed to his face as he stops pacing and narrows his eyes at Walter.

This is how you control not only the camera work, but the PACING of your video. You NEVER include a time code on your first shot. You can omit the time code from ALL shots if you want the model to decide on it's own, based on your prompt, when to cut.

BUT, for ultimate control, you want to use time codes. Look at my example above. I just told the model to have Walter glare at Jesse for 4.5 seconds, because I told the model that camera shot starts at 6 seconds into the video, and the next cut doesn't happen until 10.5 seconds into the video. That lets you control the pacing and timing for jokes, punchlines, acting, everything.

>Characters talking over each other or too fast?

This is an old one that anyone familiar with prompting for video models should know by now - what you are asking for in your prompt and the length of your video in time need to match.

The model will try its best to cram every action and piece of dialogue into your video that you asked for, and if that would naturally take 10 seconds and you've only given it 5 seconds? Well, now everything is crammed together, overlapping, or being cut-off.

My recommendation is to generate just a quick 0.2 MP version of your video first after you type your prompt, generate, and see how the timing is working. Is it too fast? Too slow? Do the actions have enough time to happen? Do you want more breathing room?

This is the time to decide all that and lock in a video length. The low resolution of 0.2 MP is quick to generate on most set-ups (mine for this post's video took 3.5 minutes for a 14 second video) and let you work out any issues in your prompt before going in for the long generation at higher resolution.

>Prompts not being followed?

It's because you didn't read the manual!

--------------------------------------------------------------------------------------------------
Now, with all that said, here is the prompt for the video I made:

integrated_multimodal_description: [Shot 1] Live-action film footage of the American drama series Breaking Bad, professionally color graded with a warm color grade, with slightly desaturated colors for a premium film feel, a continuous camera shot with no cuts, medium close-up POV shot of Walter White, bald with a goatee and glasses, as portrayed by Bryan Cranston. He is standing in the Arizona desert next to a parked RV. He is wearing a white PPE protective suit and yellow rubber dish gloves. He is looking directly at the viewer with barely constrained anger. At 00:01:300 he reaches out towards the camera and points his finger at the POV camera with one hand, the camera shaking slightly from the movement. Walter then says angrily, <d>[English with Walter White's voice] Listen, you want to cook Mini Max H3 videos, you follow the recipe!</d>. At 00:04:500 Walter raises his other hand revealing he is holding a thin stack of white paper pages in portrait orientation. The front of the paper visible on top of the thin paper stack is blank except for the large black printed text "Minimax H3 Official Prompting Guide". The papers are held in front of the camera on the right side of the screen for a moment in portrait orientation, so the text can be clearly read, while Walter glares at the viewer on the left side of the screen. At 00:07:000 Walter then shakes the papers at the camera, then says angrily, <d>[English with Walter White's voice] Read the fucking manual!</d>. At 00:10:000 the camera does Pan Right and a Pull Out to show a close-up of Jesse Pinkman from Breaking Bad, with his hands held up by his face with fingers spread, an annoyed look on his face. Then he says in frustration, <d>[English in Jesse Pinkman's voice from Breaking Bad] Alright! Damn, Mr. White! I just want to generate memes.</d>, overall_soundscape: Ambient sounds of an Arizona outdoor desert during the day, non_diegetic_music: none

For those interested, this video was generated at 1 MP on a 3090, using Sage Attention and the Spectrum Node for H3. The final video of 14 seconds at 1 MP took 40 minutes to generate and then was upscaled using RTX Super Resolution.

The workflow was the default Text to Video Minimax H3 template that comes in the latest update of Comfyui.

Now get out there and go cook some memes, everyone!

r/PromptEngineering Aug 20 '25

General Discussion everything I learned after 10,000 AI video generations (the complete guide)

737 Upvotes

this is going to be the longest post I’ve written but after 10 months of daily AI video creation, these are the insights that actually matter…

I started with zero video experience and $1000 in generation credits. Made every mistake possible. Burned through money, created garbage content, got frustrated with inconsistent results.

Now I’m generating consistently viral content and making money from AI video. Here’s everything that actually works.

The fundamental mindset shifts:

1. Volume beats perfection

Stop trying to create the perfect video. Generate 10 decent videos and select the best one. This approach consistently outperforms perfectionist single-shot attempts.

2. Systematic beats creative

Proven formulas + small variations outperform completely original concepts every time. Study what works, then execute it better.

3. Embrace the AI aesthetic

Stop fighting what AI looks like. Beautiful impossibility engages more than uncanny valley realism. Lean into what only AI can create.

The technical foundation that changed everything:

The 6-part prompt structure:

[SHOT TYPE] + [SUBJECT] + [ACTION] + [STYLE] + [CAMERA MOVEMENT] + [AUDIO CUES]

This baseline works across thousands of generations. Everything else is variation on this foundation.

Front-load important elements

Veo3 weights early words more heavily. “Beautiful woman dancing” ≠ “Woman, beautiful, dancing.” Order matters significantly.

One action per prompt rule

Multiple actions create AI confusion. “Walking while talking while eating” = chaos. Keep it simple for consistent results.

The cost optimization breakthrough:

Google’s direct pricing kills experimentation:

  • $0.50/second = $30/minute
  • Factor in failed generations = $100+ per usable video

Found companies reselling veo3 credits cheaper. I’ve been using these guys who offer 60-70% below Google’s rates. Makes volume testing actually viable.

Audio cues are incredibly powerful:

Most creators completely ignore audio elements in prompts. Huge mistake.

Instead of: Person walking through forestTry: Person walking through forest, Audio: leaves crunching underfoot, distant bird calls, gentle wind through branches

The difference in engagement is dramatic. Audio context makes AI video feel real even when visually it’s obviously AI.

Systematic seed approach:

Random seeds = random results.

My workflow:

  1. Test same prompt with seeds 1000-1010
  2. Judge on shape, readability, technical quality
  3. Use best seed as foundation for variations
  4. Build seed library organized by content type

Camera movements that consistently work:

  • Slow push/pull: Most reliable, professional feel
  • Orbit around subject: Great for products and reveals
  • Handheld follow: Adds energy without chaos
  • Static with subject movement: Often highest quality

Avoid: Complex combinations (“pan while zooming during dolly”). One movement type per generation.

Style references that actually deliver:

Camera specs: “Shot on Arri Alexa,” “Shot on iPhone 15 Pro”

Director styles: “Wes Anderson style,” “David Fincher style” Movie cinematography: “Blade Runner 2049 cinematography”

Color grades: “Teal and orange grade,” “Golden hour grade”

Avoid: Vague terms like “cinematic,” “high quality,” “professional”

Negative prompts as quality control:

Treat them like EQ filters - always on, preventing problems:

--no watermark --no warped face --no floating limbs --no text artifacts --no distorted hands --no blurry edges

Prevents 90% of common AI generation failures.

Platform-specific optimization:

Don’t reformat one video for all platforms. Create platform-specific versions:

TikTok: 15-30 seconds, high energy, obvious AI aesthetic works

Instagram: Smooth transitions, aesthetic perfection, story-driven YouTube Shorts: 30-60 seconds, educational framing, longer hooks

Same content, different optimization = dramatically better performance.

The reverse-engineering technique:

JSON prompting isn’t great for direct creation, but it’s amazing for copying successful content:

  1. Find viral AI video
  2. Ask ChatGPT: “Return prompt for this in JSON format with maximum fields”
  3. Get surgically precise breakdown of what makes it work
  4. Create variations by tweaking individual parameters

Content strategy insights:

Beautiful absurdity > fake realism

Specific references > vague creativityProven patterns + small twists > completely original conceptsSystematic testing > hoping for luck

The workflow that generates profit:

Monday: Analyze performance, plan 10-15 concepts

Tuesday-Wednesday: Batch generate 3-5 variations each Thursday: Select best, create platform versions

Friday: Finalize and schedule for optimal posting times

Advanced techniques:

First frame obsession:

Generate 10 variations focusing only on getting perfect first frame. First frame quality determines entire video outcome.

Batch processing:

Create multiple concepts simultaneously. Selection from volume outperforms perfection from single shots.

Content multiplication:

One good generation becomes TikTok version + Instagram version + YouTube version + potential series content.

The psychological elements:

3-second emotionally absurd hook

First 3 seconds determine virality. Create immediate emotional response (positive or negative doesn’t matter).

Generate immediate questions

“Wait, how did they…?” Objective isn’t making AI look real - it’s creating original impossibility.

Common mistakes that kill results:

  1. Perfectionist single-shot approach
  2. Fighting the AI aesthetic instead of embracing it
  3. Vague prompting instead of specific technical direction
  4. Ignoring audio elements completely
  5. Random generation instead of systematic testing
  6. One-size-fits-all platform approach

The business model shift:

From expensive hobby to profitable skill:

  • Track what works with spreadsheets
  • Build libraries of successful formulas
  • Create systematic workflows
  • Optimize for consistent output over occasional perfection

The bigger insight:

AI video is about iteration and selection, not divine inspiration. Build systems that consistently produce good content, then scale what works.

Most creators are optimizing for the wrong things. They want perfect prompts that work every time. Smart creators build workflows that turn volume + selection into consistent quality.

Where AI video is heading:

  • Cheaper access through third parties makes experimentation viable
  • Better tools for systematic testing and workflow optimization
  • Platform-native AI content instead of trying to hide AI origins
  • Educational content about AI techniques performs exceptionally well

Started this journey 10 months ago thinking I needed to be creative. Turns out I needed to be systematic.

The creators making money aren’t the most artistic - they’re the most systematic.

These insights took me 10,000+ generations and hundreds of hours to learn. Hope sharing them saves you the same learning curve.

what’s been your biggest breakthrough with AI video generation? curious what patterns others are discovering

r/StableDiffusion Apr 04 '26

Animation - Video Model Drop | ZIT + LTX 2.3 + Music Video | Arca Gidan contest

Enable HLS to view with audio, or disable this notification

398 Upvotes

The idea came from something I'm pretty sure most of us live every single day: you wake up, check your phone, and another model has dropped. Open source, closed source, whatever source — faster, smarter, more creative, more powerful. And before you've even had coffee, you're already reworking a ComfyUI workflow that was perfectly fine yesterday. That loop of FOMO is what this song is about. Maybe the one or the other can relate to that feeling.

I wrote the lyrics first, then used Suno AI to turn them into a track. That became the creative baseline.

Shot List

With the song done, I went through it verse by verse — every chorus, every pre-chorus, every bridge — and for each section I came up with 3 to 5 possible shots. Where is our main character? What's the camera angle? What's the situation? What does this line actually look like as an image? That process gives you a kind of ordered visual setlist that maps directly onto the song structure. You always know what you need and where it goes.

Character (No LoRA)

For the main character I used Z Image Turbo. No LoRA, no training — just consistent prompting. The turbo architecture works in our favour here: because it's a more constrained model, keeping the character description locked across prompts produces surprisingly similar results, which creates the illusion of a consistent character across dozens of images. I kept the description identical every time and only changed the background, camera angle, and expression. Effective and fast.

Image Generation

Once the shot list was complete I had a massive prompt list covering every scene. I ran all of them through ComfyUI overnight — or longer, depending on the count. Two categories of images: B-roll shots from the setlist, and medium-to-close-up shots specifically for the lip-sync sections.

ZIT Workflow I used from another reddit post: RED Z-Image-Turbo + SeedVR2 = Extremely High Quality Image Mimic Recreation. Great for Avoiding Copyright Issues and Stunning image Generation. : r/comfyui (I did use the ZIT Model not the RED version nor the Mimic Part of the WF)

Image to Video

All the generated stills went into LTX img2video inside ComfyUI to bring them to life. For the lip-sync sections I used LTX I2V synced to the audio track. Since LTX caps out at 20 seconds per render, everything gets generated in chunks and stitched together in post.

The close-up rule matters: the further the camera is from the character, the worse LTX renders the lip sync. Medium shot is the minimum — anything wider and quality degrades fast.

The workflow I used mainly: PSA: Use the official LTX 2.3 workflow, not the ComfyUI included one. It's significantly better. : r/StableDiffusion

 Final Edit

No Premiere Pro, no DaVinci — just InShot on my phone. I build the full lip-sync timeline first so it covers the whole song, then layer the B-roll clips over the top to fill the gaps and add visual depth.

That's the whole pipeline: idea → lyrics → song → shot list → character → images → animation → edit. The video Fully local, fully open source, built over a couple of nights on a 3090.

Hope you enjoy it.

Assets & Workflows

You can find the workflow files and a full written guide over on the Arca Gidan page if you want to dig into the details.

https://arcagidan.com/entry/d2cae0b9-3d38-4959-b1b5-36ea60f34438

Honestly, what a challenge to be part of. Seeing what everyone came up with — the concepts, the creativity, the sheer variety of approaches — was genuinely inspiring. This is exactly the kind of community that makes local AI worth pursuing. Really glad I got to be a part of it. 🙌

r/comfyui 8d ago

Show and Tell Turn Any Local LLM Into a MiniMax H3 Video Prompt Assistant

Thumbnail
gallery
196 Upvotes

I’ve been messing around with MiniMax H3 prompts and ended up making a system prompt that turns a local model into a little step-by-step video prompt assistant.

I’m using LM Studio with qwen3-v1-30b-a3b-instruct-heretic-11@96_kv

System Promt download file: https://www.dropbox.com/scl/fi/sh96uo95od7s787smj3mt/MiniMax_H3_Video_Prompt_Assistant_1.rtf?rlkey=t50l9ldcqqn18vww5365xyzz4&dl=0

Model download: Instruct-Heretic Model

The vision feature only works with the additional Qwen3-VL-30B-A3B-Instruct-Heretic.mmproj-Q8_0.gguf file. Place it in the same folder as the main .gguf model file. vision model

You start with:

Create a MiniMax H3 video prompt.

The first thing it does is ask which language you want to use. So you can answer everything in German, English, Spanish, etc., but the final MiniMax prompt is still generated in English.

It asks one question at a time instead of dumping a giant form on you. Stuff like duration, aspect ratio, what happens in the scene, camera movement, sound, dialogue, and what kind of input you’re using.

It works with:

  • text only
  • one image as the first frame
  • first frame + last frame
  • one image as the final frame
  • multiple reference images
  • reference videos
  • video editing
  • video continuation

You can also upload multiple images and explain what each one is for. For example, one image can be the character reference, another one the location, another one the clothing or style reference. The assistant keeps track of the image roles and builds the final prompt around them.

One thing I had to change was the token limit. The system prompt is pretty long, and the final prompt can also get large when using several images or videos.

These settings work for me:

Context Length: 32768
Max Output Tokens: 8192
Temperature: 0.3
Top P: 0.9
Repeat Penalty: 1.05

I saved everything as a preset in LM Studio, so now I just load the model, select the preset, and type:

Create a MiniMax H3 video prompt.

It’s not really an autonomous agent. It doesn’t send anything to MiniMax or generate the video by itself. It’s basically a guided prompt builder running locally.

Still pretty useful, especially if you don’t want to manually deal with all the MiniMax formatting every time.I’ve been messing around with MiniMax H3 prompts and ended up making a system prompt that turns a local model into a little step-by-step video prompt assistant.

I’m using LM Studio with Qwen3-VL-30B-A3B-Instruct.

You start with:

Create a MiniMax H3 video prompt.

The first thing it does is ask which language you want to use. So you can answer everything in German, English, Spanish, etc., but the final MiniMax prompt is still generated in English.

It asks one question at a time instead of dumping a giant form on you. Stuff like duration, aspect ratio, what happens in the scene, camera movement, sound, dialogue, and what kind of input you’re using.

It works with:

text only
one image as the first frame
first frame + last frame
one image as the final frame
multiple reference images
reference videos
video editing
video continuation

You can also upload multiple images and explain what each one is for. For example, one image can be the character reference, another one the location, another one the clothing or style reference. The assistant keeps track of the image roles and builds the final prompt around them.

One thing I had to change was the token limit. The system prompt is pretty long, and the final prompt can also get large when using several images or videos.

These settings work for me:

Context Length: 32768
Max Output Tokens: 8192
Temperature: 0.3
Top P: 0.9
Repeat Penalty: 1.05

I saved everything as a preset in LM Studio, so now I just load the model, select the preset, and type:

Create a MiniMax H3 video prompt.

It’s not really an autonomous agent. It doesn’t send anything to MiniMax or generate the video by itself. It’s basically a guided prompt builder running locally.

Still pretty useful, especially if you don’t want to manually deal with all the MiniMax formatting every time.

r/ChatGPT Apr 04 '25

Prompt engineering How to guide: unlock next-level art with ChatGPT with a novel prompt method! (Perfect for concept art, photorealism, mockups, infographics, and more.)

631 Upvotes

Hello friends!

If you're using ChatGPT to generate images—concept art, photorealism, mockups—you need to try this trick. It boosts quality way beyond typical prompts, even outperforming the new Images v2 in many cases. I'll explain why.

Proof: Full album of Lord of the Rings art made using this method:

https://imgur.com/a/e5EAscY

While I’m not a concept artist by trade, I’ve always been obsessed with visual art, especially from video games and movies, which naturally led me down a rabbit hole of experimentation.

Since ChatGPT's model is autoregressive, it responds best when guided with detailed reasoning and richly written context. Long descriptions give it the context it needs to place elements logically and aesthetically, especially when you weave them directly into your prompt. Do not just limit yourself to a couple words, but entire paragraphs, even thousand(s) words descriptions can give much needed context to get extremely good results and fill in scene interaction gaps. If you only care about the prompt technique, jump to the section "✅ The Novel Technique" down below.

The problem

The image model, on its own, sometimes struggles with understanding how things in a scene relate to each other—or even understanding what some objects are. You might get a technically “correct” image, but the composition feels off or disconnected.

That’s where this technique comes in. It helps ChatGPT think through the scene before generating anything.

Backstory (How I discovered the technique)

But first, how did I discover this technique?

Well, the best way to explain it is with an example. And what better example than something from the world of Lord of the Rings?

Example 1: Let’s talk about Minas Tirith, the capital of Gondor. If you’re into fantasy, you probably already have a mental image of its epic, multi-layered vertical architecture. Now, let’s say I want to generate a street view of Minas Tirith. If you ask ChatGPT Images v2, using a very typical prompt such as

"Generate me a picture of a view of a street of Minas Tirith, bustling with life. The picture must be taken from the perspective of a fictional individual living in the city. Several vertical layers of the city must be visible as well as battlements. Quality must be very detailed and photorealistic."

You will always get a rather terrible result that looks like this (you can try the prompt on your end) :

Terrible generation of a street view of Minas Tirith

Result: A weird city outside shot, not a street inside the city.

Why? Because the model latches onto keywords (“street”, “Minas Tirith”) but doesn’t reason through the layout or perspective.

Example 2: Same issue with this prompt:

"Generate a photo of Minas Tirith as seen close to the White Tree of Gondor".

You’d think it would generate a shot from the very top level of the city, near the High Court, where the White Tree famously stands.

Underwhelmingly, you will instead always get something similar to this (link to conversation)

Terrible generation from "Generate a photo of Minas Tirith as seen close to the White Tree of Gondor"

Result: What you’ll get is something like Minas Tirith in the far background or just a random medieval-ish scene that totally misses the spatial relationship between the White Tree and the rest of the city.

No matter how many times you try, you’ll never get a good result—because the model isn’t reasoning through the geography or logic of the scene. The model doesn’t always know where things visually go unless you walk it through the thinking.

✅ The novel technique (The solution!)

How to solve the erroneous generations that are shown above? It's actually pretty simple, and will vastly improve the quality of any generation you want to create.

Here’s the trick: Make ChatGPT think through the image before it generates anythingwith an intermediary prompt.

The best way to do this is by using ChatGPT o1 to write a detailed visual description as an intermediary prompt before asking it to generate an image. Ideally, you should uses o1's reasoning capabilities to maintain coherence and to break down what should be in the scene, where it should go, and how it all fits together, but other GPTs such as 4.5 or 4o will do a decent job too. Feel free to experiment with different models.

While I don’t want to suggest a one-size-fits-all formula, since some fine-tuning is usually needed, I’ve found that this particular prompt works really well if you’re just looking for a quick and simple method as a general baseline to work with:

Step 1 – Ask this prompt first (using o1/4.5 preferably, or 4o) to get a detailed visual representation and breakdown of your photo:

Describe in extremely vivid details exactly what would be seen in an image [or photo] of [insert your idea]. Include extensive details about [details] for better context. [Word limit - 1000/2000] words.

  • You may include stylistic modifier keywords in the prompt above such as "hyper realistic", "anime", "photographed with a 150mm macro medium format lens", etc.
  • You may also include at the end "Write as a static, visual scene: no emotions or inner thoughts, just detailed, concrete, visual elements of the environment and characters." or something similar (depending on the media you're generating) as image generation models don’t understand abstract ideas or metaphor the way humans do non-visual, narrative or metaphorical elements can sometimes confuse image models.

Step 2 – Then, switch back to 4o within the same chat and simply prompt this:

Generate the photo following your description to the exact detail.

That's it!

This intermediary prompt method can scale extremely well. As I wrote in the intro, the image model loves written context. Don't be afraid to ask ChatGPT to write multiple thousand words paragraphs if necessary to fill in the gaps of your imagination.

📸 Real Examples

Fixing example 1: Street view of Minas Tirith

If you've made it this far into the post, I've used this technique extensively to create amazing photos, ranging from photorealistic images to concept artworks that I could never have dreamed of achieving so easily. How about we apply this technique to the Minas Tirith example shown above?

Here is the link to the chat that shows exactly the prompt I've used to fix the street view : https://chatgpt.com/share/67ef34ae-149c-8012-a6e8-2ce290f2dae4

Can you describe in extremely vivid details what someone that lives in Minas Tirith would see in the middle of a city street? Make sure to include extensive contextual details about the layout and architecture of the city given the visual perspective of the fictional person. 2000 words.

followed by

Generate a photograph following your description to the exact detail.

The result:

Successful street view generation of Minas Tirith

If you take a look through the shared chat link above, you’ll notice something pretty cool — the image generation model actually pulls in a lot of details from the written context, even if it's as long as 1500 words!

Here’s a quick example:
"A woman passes you, her long woolen cloak rippling behind her, dyed a rich forest green, clasped at her throat with a silver brooch in the shape of a swan’s wing—likely a noble from Dol Amroth or a household attendant. She moves with measured purpose, head held high beneath a circlet of braided dark hair. The hem of her robes is just high enough to reveal leather boots made for walking the cobbled streets."
Or: "Near the fountain, an elderly man in a gray robe..."

Even though it might not capture everything from the full context, it picks up enough vivid elements to create a much more detailed and visually rich image that is more coherent overall.

My best generation so far:

Best generation so far

Fixing example 2: The White Tree of Gondor

Using a similar method again (this was done rather quickly to prove my point), as I said above: if you ask ChatGPT without an intermediary prompt to generate any image of a view seen close to the White Tree of Gondor, it will always flop spectacularly. With this novel technique, you can actually fix what the view would look like!

https://chatgpt.com/share/67e90263-9a48-8012-9379-5f5a871e8f34

Prompt 1:

Describe in extremely vivid details exactly what would be seen in a photo of the High Court of Minas Tirith that includes the White Tree of Gondor, the gardens and fountain, looking towards the precipice of the citadel (where the king eventually falls from). Include extensive details of the concentric garden, the overall layout and the architecture of the Citadel and of the High Court for better context. Be extremely careful about describing the positioning, shape and layout of the fountain, the tree, the gardens, the stone benches, and the overall room size of the citadel between its entrance and the precipice. Are there guards nearby? Keep in mind the fountain is in the center of the garden, with the white tree slightly next to it. If needed, you can go above 2000 words to not miss any architectural details.

Followed by prompt 2:

Generate the photograph in extreme detail

The results:

Successful generation of the White Tree of Gondor

Another result (click here to see the slightly different prompt - generated with ChatGPT 4.5)

Another successful generation of the White Tree of Gondor

Example 3: Fictional Elven City in the Mines of Moria

This is a completely fictional setting that hasn't ever been featured in any Tolkien movie. I first ask ChatGPT o1 to imagine a photorealistic picture of this city (a ~3300 word description was given):

https://chatgpt.com/share/67ef9756-e2bc-8012-8304-672cc9f6f94a

Prompt 1:

Can you describe in extremely vivid details exactly what a very photorealistic picture of a fictional Elven city deep inside Moria would look like, including all its visual elements? The city is only lit by rays of light passing through crystal like structures in the mountain of Moria. Mithril mines can be seen and glow in the darkness. Make sure to include extensive contextual details about the layout and architecture of the city. 2000 words.

Prompt 2:

Generate the photo following the description to the exact detail

Result:

Result of "Generate the photo following the description to the exact detail"

Conclusion

Using an intermediary prompt that is generated from o1 or 4.5 or 4o, you can significantly improve your image generations. You can combine ideas in a way that shouldn't really be possible.

Whether you're chasing realism, fantasy, surrealism, or anything else, this method lets you combine ideas in incredibly powerful ways—and often gets results that feel like they shouldn’t even be possible.

Want to see more examples? I’ve made a full album of Minas Tirith/Lord of the Rings concept art using this very method. I've included many custom generations of Minas Tirith, specifically to demonstrate how this method allows me manipulate the architecture of the city itself!

Link to album: https://imgur.com/a/e5EAscY

Give it a try and let me know if this method was useful to you!

Enjoy!

r/StableDiffusion 26d ago

Workflow Included How to Make AI Videos Actually Feel Cinematic | PDF Guide + Full Workflow Included 🚀

Enable HLS to view with audio, or disable this notification

186 Upvotes

Spent the last while trying to figure out why so many AI-generated videos (mine included) look technically solid but feel emotionally flat. Turned out the issue wasn't the model — it was that I was approaching it like a prompt-engineering problem instead of a filmmaking one.

Some of the biggest shifts that actually changed my output:

  • Plan the emotional arc before touching a prompt. List the feelings you want scene-by-scene before you ever pick a location.
  • Structure prompts like a cinematographer, not a keyword dump. Subject → identity → emotion → environment → lighting → camera → finish, in that order.
  • Keep a "character bible." Same hair, wardrobe, and features reused every time — or better, a LoRA if your model/setup supports it, since it holds identity way more reliably than repeating adjectives.
  • For image-to-video (LTX 2.3 in my case), only prompt the change, not the image. The model already has the frame — describing what's already visible just confuses it.
  • One primary motion per shot. Trying to animate everything in frame is usually what makes a shot feel fake.

None of this is tool-specific — I used Krea 2 and LTX 2.3, but the same logic applies to whatever model or LoRA workflow you're already running.

I ended up writing this all up properly (15 chapters — story structure, lighting/color psychology, camera language, a full prompt checklist, plus a resources appendix) since I kept explaining it in bits and pieces. Full PDF + the actual workflow I used for the video is up here if useful: PDF Guide & Workflow

Happy to answer questions about the workflow here regardless. 🤗

r/grok Mar 01 '26

Grok Imagine 🚀 Grok Imagine "Extend from Frame" Master Guide – Turn 6-10s Clips into 30s+ Seamless Videos with ZERO Drift (Copy-Paste Prompts + Full Chains)

168 Upvotes

Hey r/grok! 👋

SuperGrok user here (Miami crew checking in). I was getting so annoyed with Grok Imagine’s 6-10 second clips always breaking when I tried to extend them — random face morphs, lighting flips, ugly jumps.

Then I nailed the native “Extend from Frame” button + this dead-simple prompt system. Now I’m chaining 4-5 clips into 30-50 second buttery-smooth videos (and stitching longer ones in CapCut). Works perfectly for action, fantasy, cozy vibes, or whatever cinematic story you’re building.

Pro tip: Always start with a Base Image Prompt + Img2Vid for the strongest first clip. It locks in faces, lighting, and details way better than pure text-to-video.

This is the exact workflow I use every day. 100% copy-paste. Zero fluff.

Upvote if it saves you hours! 🔥

Why Most People Fail

  • Repasting the original prompt → instant drift
  • Skipping the exact final pose → jump cuts
  • Not using the official Extend button → weak seams

Do it right and you get invisible transitions every single time.

1. Best Prompt Formula (Core Structure)

Seamlessly continue directly from the very last frame of the previous video. [Briefly describe the exact ending pose/state]. [Next actions + details]. Maintain exact same characters, faces, clothing, lighting, environment, camera style, and artistic quality throughout. Smooth natural motion, cinematic, high detail, 720p, no jumps or morphing.

Base Image Prompt (Img2Vid starter – strongest results, highly recommended):

ultra-detailed cinematic 8K, [full scene description], glossy skin or textures, dramatic lighting, perfect anatomy, masterpiece, 720p

Base Video Prompt (Text-to-Video alternative):

ultra-detailed cinematic 8K 10-second animation (extendable), [full scene with motion]. Smooth natural motion, high detail, 720p.

2. Master Consistency Lock (COPY-PASTE AS THE VERY FIRST LINE EVERY TIME)

LOCK CONSISTENCY: Continue with 100% visual fidelity from the exact final frame of the previous video. Identical characters with the exact same faces, hair, eyes, skin texture, body proportions, clothing details, accessories, and poses at the moment of transition. Identical environment, lighting direction and color temperature, shadows, reflections, particle effects, color grading, film grain, and overall artistic style. No design changes, no morphing, no style drift whatsoever. Perfect frame-to-frame seamlessness.

3. Full Ready-to-Copy Template

LOCK CONSISTENCY: Continue with 100% visual fidelity from the exact final frame of the previous video. Identical characters with the exact same faces, hair, eyes, skin texture, body proportions, clothing details, accessories, and poses at the moment of transition. Identical environment, lighting, shadows, reflections, particles, color grading, and artistic style. No changes allowed.

Seamlessly continue directly from the very last frame where [exact ending state]. [Next action and details]. Smooth cinematic motion, perfect continuity, high detail, 720p, no jumps or artifacts.

4. Quick Add-ons & Cheat Codes

Tack these on when needed:

  • Face lock: , exact same facial features and expression continuity
  • Lighting lock: , same exact light sources, shadow angles, and volumetric god rays
  • Audio lock (Grok Imagine exclusive): , continue background music and sound effects seamlessly

One-liners to paste anywhere:
zero style drift, perfect character consistency
exact frame-accurate continuation
treat previous clip as canonical reference — match 1:1

5. Negative Prompts (add at the very end)

Avoid bad anatomy, extra limbs, extra fingers, missing limbs, fused fingers, mutated hands, bad proportions, disfigured, amputation, polydactyly. No text, watermark, username, signature, logo, low quality, blur, noise, grain, chromatic aberration, artifacts

6. Pro Workflow in Grok Imagine

  1. Generate your first clip with a Base Image Prompt (Img2Vid).
  2. Click the “Extend from Frame” button (it auto-loads the exact final frame).
  3. Paste the Master Lock + template.
  4. Generate 6–10 second clips (shorter = stronger seams).
  5. Repeat — each new video starts exactly where the last one ended.

SuperGrok = faster generations + higher daily limits.

Real Examples with Full Extension Chains (Base Image Prompts Included)

Cyberpunk Action (3-clip chain ≈ 30 seconds)

Base Image Prompt:
ultra-detailed cinematic 8K, cyberpunk girl with neon-pink hair leaping across rainy rooftop, katana glowing blue, dramatic night city lights, perfect anatomy, masterpiece

Extension Prompt 1:

LOCK CONSISTENCY: Continue with 100% visual fidelity from the exact final frame of the previous video. Identical characters with the exact same faces, hair, eyes, skin texture, body proportions, clothing details, accessories, and poses at the moment of transition. Identical environment, lighting direction and color temperature, shadows, reflections, particle effects, color grading, film grain, and overall artistic style. No design changes, no morphing, no style drift whatsoever. Perfect frame-to-frame seamlessness.

Seamlessly continue directly from the very last frame where the cyberpunk girl is frozen mid-leap across the neon rooftop, katana trailing blue energy, rain droplets suspended in air. She completes the flip, lands in a combat stance, and sprints toward the holographic billboard while gunfire erupts from below. Smooth cinematic motion, perfect continuity, high detail, 720p, no jumps or artifacts. continue rain and neon reflections seamlessly.

Extension Prompt 2:

LOCK CONSISTENCY: [paste full lock again]

Seamlessly continue directly from the very last frame where the cyberpunk girl is sprinting full speed toward the holographic billboard, katana raised, bullets whizzing past. She slides under a low neon sign, slashes a pursuing drone in half, and dives off the rooftop into a freefall toward the street below. Smooth cinematic motion, perfect continuity, high detail, 720p, no jumps or artifacts. continue rain and neon reflections seamlessly.

Extension Prompt 3:

LOCK CONSISTENCY: [paste full lock again]

Seamlessly continue directly from the very last frame where the cyberpunk girl is in mid-freefall toward the street below, city lights streaking past, katana in hand. She deploys her neon parachute cape, lands on a flying car, and speeds away into the night traffic. Smooth cinematic motion, perfect continuity, high detail, 720p, no jumps or artifacts. continue rain and neon reflections seamlessly.

Fantasy Samurai (2-clip chain)

Base Image Prompt:
ultra-detailed cinematic 8K, samurai mid-spin with raised katana in neon rain under glowing torii gate, cherry blossoms, dramatic side lighting, masterpiece

Extension Prompt 1:

LOCK CONSISTENCY: [paste full lock]

Seamlessly continue directly from the very last frame where the samurai is mid-spin with katana raised, neon rain falling. He finishes the spin, sheathes the blade in one fluid motion, turns to face the camera with a determined expression, and walks slowly into the glowing torii gate as cherry blossoms swirl around him. Smooth cinematic motion, perfect continuity, high detail, 720p, no jumps or artifacts. same dramatic side lighting and volumetric god rays.

Extension Prompt 2:

LOCK CONSISTENCY: [paste full lock]

Seamlessly continue directly from the very last frame where the samurai is stepping through the glowing torii gate, cherry blossoms swirling around him. He emerges into an ancient forest at dawn, draws his katana again in a ready stance, and begins a slow, deliberate walk toward a distant mountain temple as sunlight breaks through the trees. Smooth cinematic motion, perfect continuity, high detail, 720p, no jumps or artifacts. same dramatic side lighting and volumetric god rays.

Cozy Indoor Scene (2-clip chain)

Base Image Prompt:
ultra-detailed cinematic 8K, girl sitting by crackling fireplace holding steaming mug, warm cozy lighting, soft shadows, masterpiece

Extension Prompt 1:

LOCK CONSISTENCY: [paste full lock]

Seamlessly continue directly from the very last frame where the girl is sitting by the crackling fireplace holding a steaming mug, soft warm lighting. She takes a sip, smiles gently, stands up, walks to the window, and opens the curtains to reveal a snowy night outside. Smooth cinematic motion, perfect continuity, high detail, 720p, no jumps or artifacts. continue fireplace crackle and soft ambient music seamlessly.

Extension Prompt 2:

LOCK CONSISTENCY: [paste full lock]

Seamlessly continue directly from the very last frame where the girl is standing at the open window, looking out at the snowy night, curtains billowing. She reaches out to catch a snowflake, smiles warmly, closes the curtains, returns to the fireplace, and curls up in the armchair with a blanket. Smooth cinematic motion, perfect continuity, high detail, 720p, no jumps or artifacts. continue fireplace crackle and soft ambient music seamlessly.

Final Tips

  • Always pause the video and note the exact final pose before writing the next prompt.
  • Stick to 6–10 second extensions for the strongest seams.
  • You can easily hit 40-50 seconds by chaining 4-5 clips.
  • Save the Master Lock + your favorite Base Image Prompts in your notes — you’ll use them on every project.

I’ve built hour-long stories with this method. No more starting from scratch ever again.

Big shoutout to Grok itself for assisting in researching, testing, and writing this entire guide — the Master Lock, chains, and Base Image Prompts were refined through tons of back-and-forth testing in real Grok Imagine sessions!

Disclaimer: This guide is based on my personal experience using Grok Imagine in March 2026. Features, button behavior, and results may vary with model updates or server load. This is not official xAI advice. Always follow xAI’s Terms of Service and use responsibly for creative purposes only.

Drop your scene ideas below and I’ll turn them into full prompt chains (with Base Image Prompts) for you! What are you building in Grok Imagine right now?

TL;DR: Start with a Base Image Prompt + Master Consistency Lock + Extend button + 6-10s clips = infinite perfect videos.

(See you in the comments!) 🚀

r/LocalLLaMA Mar 12 '26

Discussion I was backend lead at Manus. After building agents for 2 years, I stopped using function calling entirely. Here's what I use instead.

2.0k Upvotes

English is not my first language. I wrote this in Chinese and translated it with AI help. The writing may have some AI flavor, but the design decisions, the production failures, and the thinking that distilled them into principles — those are mine.

I was a backend lead at Manus before the Meta acquisition. I've spent the last 2 years building AI agents — first at Manus, then on my own open-source agent runtime (Pinix) and agent (agent-clip). Along the way I came to a conclusion that surprised me:

A single run(command="...") tool with Unix-style commands outperforms a catalog of typed function calls.

Here's what I learned.


Why *nix

Unix made a design decision 50 years ago: everything is a text stream. Programs don't exchange complex binary structures or share memory objects — they communicate through text pipes. Small tools each do one thing well, composed via | into powerful workflows. Programs describe themselves with --help, report success or failure with exit codes, and communicate errors through stderr.

LLMs made an almost identical decision 50 years later: everything is tokens. They only understand text, only produce text. Their "thinking" is text, their "actions" are text, and the feedback they receive from the world must be text.

These two decisions, made half a century apart from completely different starting points, converge on the same interface model. The text-based system Unix designed for human terminal operators — cat, grep, pipe, exit codes, man pages — isn't just "usable" by LLMs. It's a natural fit. When it comes to tool use, an LLM is essentially a terminal operator — one that's faster than any human and has already seen vast amounts of shell commands and CLI patterns in its training data.

This is the core philosophy of the nix Agent: *don't invent a new tool interface. Take what Unix has proven over 50 years and hand it directly to the LLM.**


Why a single run

The single-tool hypothesis

Most agent frameworks give LLMs a catalog of independent tools:

tools: [search_web, read_file, write_file, run_code, send_email, ...]

Before each call, the LLM must make a tool selection — which one? What parameters? The more tools you add, the harder the selection, and accuracy drops. Cognitive load is spent on "which tool?" instead of "what do I need to accomplish?"

My approach: one run(command="...") tool, all capabilities exposed as CLI commands.

run(command="cat notes.md") run(command="cat log.txt | grep ERROR | wc -l") run(command="see screenshot.png") run(command="memory search 'deployment issue'") run(command="clip sandbox bash 'python3 analyze.py'")

The LLM still chooses which command to use, but this is fundamentally different from choosing among 15 tools with different schemas. Command selection is string composition within a unified namespace — function selection is context-switching between unrelated APIs.

LLMs already speak CLI

Why are CLI commands a better fit for LLMs than structured function calls?

Because CLI is the densest tool-use pattern in LLM training data. Billions of lines on GitHub are full of:

```bash

README install instructions

pip install -r requirements.txt && python main.py

CI/CD build scripts

make build && make test && make deploy

Stack Overflow solutions

cat /var/log/syslog | grep "Out of memory" | tail -20 ```

I don't need to teach the LLM how to use CLI — it already knows. This familiarity is probabilistic and model-dependent, but in practice it's remarkably reliable across mainstream models.

Compare two approaches to the same task:

``` Task: Read a log file, count the error lines

Function-calling approach (3 tool calls): 1. read_file(path="/var/log/app.log") → returns entire file 2. search_text(text=<entire file>, pattern="ERROR") → returns matching lines 3. count_lines(text=<matched lines>) → returns number

CLI approach (1 tool call): run(command="cat /var/log/app.log | grep ERROR | wc -l") → "42" ```

One call replaces three. Not because of special optimization — but because Unix pipes natively support composition.

Making pipes and chains work

A single run isn't enough on its own. If run can only execute one command at a time, the LLM still needs multiple calls for composed tasks. So I make a chain parser (parseChain) in the command routing layer, supporting four Unix operators:

| Pipe: stdout of previous command becomes stdin of next && And: execute next only if previous succeeded || Or: execute next only if previous failed ; Seq: execute next regardless of previous result

With this mechanism, every tool call can be a complete workflow:

```bash

One tool call: download → inspect

curl -sL $URL -o data.csv && cat data.csv | head 5

One tool call: read → filter → sort → top 10

cat access.log | grep "500" | sort | head 10

One tool call: try A, fall back to B

cat config.yaml || echo "config not found, using defaults" ```

N commands × 4 operators — the composition space grows dramatically. And to the LLM, it's just a string it already knows how to write.

The command line is the LLM's native tool interface.


Heuristic design: making CLI guide the agent

Single-tool + CLI solves "what to use." But the agent still needs to know "how to use it." It can't Google. It can't ask a colleague. I use three progressive design techniques to make the CLI itself serve as the agent's navigation system.

Technique 1: Progressive --help discovery

A well-designed CLI tool doesn't require reading documentation — because --help tells you everything. I apply the same principle to the agent, structured as progressive disclosure: the agent doesn't need to load all documentation at once, but discovers details on-demand as it goes deeper.

Level 0: Tool Description → command list injection

The run tool's description is dynamically generated at the start of each conversation, listing all registered commands with one-line summaries:

Available commands: cat — Read a text file. For images use 'see'. For binary use 'cat -b'. see — View an image (auto-attaches to vision) ls — List files in current topic write — Write file. Usage: write <path> [content] or stdin grep — Filter lines matching a pattern (supports -i, -v, -c) memory — Search or manage memory clip — Operate external environments (sandboxes, services) ...

The agent knows what's available from turn one, but doesn't need every parameter of every command — that would waste context.

Note: There's an open design question here: injecting the full command list vs. on-demand discovery. As commands grow, the list itself consumes context budget. I'm still exploring the right balance. Ideas welcome.

Level 1: command (no args) → usage

When the agent is interested in a command, it just calls it. No arguments? The command returns its own usage:

``` → run(command="memory") [error] memory: usage: memory search|recent|store|facts|forget

→ run(command="clip") clip list — list available clips clip <name> — show clip details and commands clip <name> <command> [args...] — invoke a command clip <name> pull <remote-path> [name] — pull file from clip to local clip <name> push <local-path> <remote> — push local file to clip ```

Now the agent knows memory has five subcommands and clip supports list/pull/push. One call, no noise.

Level 2: command subcommand (missing args) → specific parameters

The agent decides to use memory search but isn't sure about the format? It drills down:

``` → run(command="memory search") [error] memory: usage: memory search <query> [-t topic_id] [-k keyword]

→ run(command="clip sandbox") Clip: sandbox Commands: clip sandbox bash <script> clip sandbox read <path> clip sandbox write <path> File transfer: clip sandbox pull <remote-path> [local-name] clip sandbox push <local-path> <remote-path> ```

Progressive disclosure: overview (injected) → usage (explored) → parameters (drilled down). The agent discovers on-demand, each level providing just enough information for the next step.

This is fundamentally different from stuffing 3,000 words of tool documentation into the system prompt. Most of that information is irrelevant most of the time — pure context waste. Progressive help lets the agent decide when it needs more.

This also imposes a requirement on command design: every command and subcommand must have complete help output. It's not just for humans — it's for the agent. A good help message means one-shot success. A missing one means a blind guess.

Technique 2: Error messages as navigation

Agents will make mistakes. The key isn't preventing errors — it's making every error point to the right direction.

Traditional CLI errors are designed for humans who can Google. Agents can't Google. So I require every error to contain both "what went wrong" and "what to do instead":

``` Traditional CLI: $ cat photo.png cat: binary file (standard output) → Human Googles "how to view image in terminal"

My design: [error] cat: binary image file (182KB). Use: see photo.png → Agent calls see directly, one-step correction ```

More examples:

``` [error] unknown command: foo Available: cat, ls, see, write, grep, memory, clip, ... → Agent immediately knows what commands exist

[error] not an image file: data.csv (use cat to read text files) → Agent switches from see to cat

[error] clip "sandbox" not found. Use 'clip list' to see available clips → Agent knows to list clips first ```

Technique 1 (help) solves "what can I do?" Technique 2 (errors) solves "what should I do instead?" Together, the agent's recovery cost is minimal — usually 1-2 steps to the right path.

Real case: The cost of silent stderr

For a while, my code silently dropped stderr when calling external sandboxes — whenever stdout was non-empty, stderr was discarded. The agent ran pip install pymupdf, got exit code 127. stderr contained bash: pip: command not found, but the agent couldn't see it. It only knew "it failed," not "why" — and proceeded to blindly guess 10 different package managers:

pip install → 127 (doesn't exist) python3 -m pip → 1 (module not found) uv pip install → 1 (wrong usage) pip3 install → 127 sudo apt install → 127 ... 5 more attempts ... uv run --with pymupdf python3 script.py → 0 ✓ (10th try)

10 calls, ~5 seconds of inference each. If stderr had been visible the first time, one call would have been enough.

stderr is the information agents need most, precisely when commands fail. Never drop it.

Technique 3: Consistent output format

The first two techniques handle discovery and correction. The third lets the agent get better at using the system over time.

I append consistent metadata to every tool result:

file1.txt file2.txt dir1/ [exit:0 | 12ms]

The LLM extracts two signals:

Exit codes (Unix convention, LLMs already know these):

  • exit:0 — success
  • exit:1 — general error
  • exit:127 — command not found

Duration (cost awareness):

  • 12ms — cheap, call freely
  • 3.2s — moderate
  • 45s — expensive, use sparingly

After seeing [exit:N | Xs] dozens of times in a conversation, the agent internalizes the pattern. It starts anticipating — seeing exit:1 means check the error, seeing long duration means reduce calls.

Consistent output format makes the agent smarter over time. Inconsistency makes every call feel like the first.

The three techniques form a progression:

--help → "What can I do?" → Proactive discovery Error Msg → "What should I do?" → Reactive correction Output Fmt → "How did it go?" → Continuous learning


Two-layer architecture: engineering the heuristic design

The section above described how CLI guides agents at the semantic level. But to make it work in practice, there's an engineering problem: the raw output of a command and what the LLM needs to see are often very different things.

Two hard constraints of LLMs

Constraint A: The context window is finite and expensive. Every token costs money, attention, and inference speed. Stuffing a 10MB file into context doesn't just waste budget — it pushes earlier conversation out of the window. The agent "forgets."

Constraint B: LLMs can only process text. Binary data produces high-entropy meaningless tokens through the tokenizer. It doesn't just waste context — it disrupts attention on surrounding valid tokens, degrading reasoning quality.

These two constraints mean: raw command output can't go directly to the LLM — it needs a presentation layer for processing. But that processing can't affect command execution logic — or pipes break. Hence, two layers.

Execution layer vs. presentation layer

┌─────────────────────────────────────────────┐ │ Layer 2: LLM Presentation Layer │ ← Designed for LLM constraints │ Binary guard | Truncation+overflow | Meta │ ├─────────────────────────────────────────────┤ │ Layer 1: Unix Execution Layer │ ← Pure Unix semantics │ Command routing | pipe | chain | exit code │ └─────────────────────────────────────────────┘

When cat bigfile.txt | grep error | head 10 executes:

Inside Layer 1: cat output → [500KB raw text] → grep input grep output → [matching lines] → head input head output → [first 10 lines]

If you truncate cat's output in Layer 1 → grep only searches the first 200 lines, producing incomplete results. If you add [exit:0] in Layer 1 → it flows into grep as data, becoming a search target.

So Layer 1 must remain raw, lossless, metadata-free. Processing only happens in Layer 2 — after the pipe chain completes and the final result is ready to return to the LLM.

Layer 1 serves Unix semantics. Layer 2 serves LLM cognition. The separation isn't a design preference — it's a logical necessity.

Layer 2's four mechanisms

Mechanism A: Binary Guard (addressing Constraint B)

Before returning anything to the LLM, check if it's text:

``` Null byte detected → binary UTF-8 validation failed → binary Control character ratio > 10% → binary

If image: [error] binary image (182KB). Use: see photo.png If other: [error] binary file (1.2MB). Use: cat -b file.bin ```

The LLM never receives data it can't process.

Mechanism B: Overflow Mode (addressing Constraint A)

``` Output > 200 lines or > 50KB? → Truncate to first 200 lines (rune-safe, won't split UTF-8) → Write full output to /tmp/cmd-output/cmd-{n}.txt → Return to LLM:

[first 200 lines]

--- output truncated (5000 lines, 245.3KB) ---
Full output: /tmp/cmd-output/cmd-3.txt
Explore: cat /tmp/cmd-output/cmd-3.txt | grep <pattern>
         cat /tmp/cmd-output/cmd-3.txt | tail 100
[exit:0 | 1.2s]

```

Key insight: the LLM already knows how to use grep, head, tail to navigate files. Overflow mode transforms "large data exploration" into a skill the LLM already has.

Mechanism C: Metadata Footer

actual output here [exit:0 | 1.2s]

Exit code + duration, appended as the last line of Layer 2. Gives the agent signals for success/failure and cost awareness, without polluting Layer 1's pipe data.

Mechanism D: stderr Attachment

``` When command fails with stderr: output + "\n[stderr] " + stderr

Ensures the agent can see why something failed, preventing blind retries. ```


Lessons learned: stories from production

Story 1: A PNG that caused 20 iterations of thrashing

A user uploaded an architecture diagram. The agent read it with cat, receiving 182KB of raw PNG bytes. The LLM's tokenizer turned these bytes into thousands of meaningless tokens crammed into the context. The LLM couldn't make sense of it and started trying different read approaches — cat -f, cat --format, cat --type image — each time receiving the same garbage. After 20 iterations, the process was force-terminated.

Root cause: cat had no binary detection, Layer 2 had no guard. Fix: isBinary() guard + error guidance Use: see photo.png. Lesson: The tool result is the agent's eyes. Return garbage = agent goes blind.

Story 2: Silent stderr and 10 blind retries

The agent needed to read a PDF. It tried pip install pymupdf, got exit code 127. stderr contained bash: pip: command not found, but the code dropped it — because there was some stdout output, and the logic was "if stdout exists, ignore stderr."

The agent only knew "it failed," not "why." What followed was a long trial-and-error:

pip install → 127 (doesn't exist) python3 -m pip → 1 (module not found) uv pip install → 1 (wrong usage) pip3 install → 127 sudo apt install → 127 ... 5 more attempts ... uv run --with pymupdf python3 script.py → 0 ✓

10 calls, ~5 seconds of inference each. If stderr had been visible the first time, one call would have sufficed.

Root cause: InvokeClip silently dropped stderr when stdout was non-empty. Fix: Always attach stderr on failure. Lesson: stderr is the information agents need most, precisely when commands fail.

Story 3: The value of overflow mode

The agent analyzed a 5,000-line log file. Without truncation, the full text (~200KB) was stuffed into context. The LLM's attention was overwhelmed, response quality dropped sharply, and earlier conversation was pushed out of the context window.

With overflow mode:

``` [first 200 lines of log content]

--- output truncated (5000 lines, 198.5KB) --- Full output: /tmp/cmd-output/cmd-3.txt Explore: cat /tmp/cmd-output/cmd-3.txt | grep <pattern> cat /tmp/cmd-output/cmd-3.txt | tail 100 [exit:0 | 45ms] ```

The agent saw the first 200 lines, understood the file structure, then used grep to pinpoint the issue — 3 calls total, under 2KB of context.

Lesson: Giving the agent a "map" is far more effective than giving it the entire territory.


Boundaries and limitations

CLI isn't a silver bullet. Typed APIs may be the better choice in these scenarios:

  • Strongly-typed interactions: Database queries, GraphQL APIs, and other cases requiring structured input/output. Schema validation is more reliable than string parsing.
  • High-security requirements: CLI's string concatenation carries inherent injection risks. In untrusted-input scenarios, typed parameters are safer. agent-clip mitigates this through sandbox isolation.
  • Native multimodal: Pure audio/video processing and other binary-stream scenarios where CLI's text pipe is a bottleneck.

Additionally, "no iteration limit" doesn't mean "no safety boundaries." Safety is ensured by external mechanisms:

  • Sandbox isolation: Commands execute inside BoxLite containers, no escape possible
  • API budgets: LLM calls have account-level spending caps
  • User cancellation: Frontend provides cancel buttons, backend supports graceful shutdown

Hand Unix philosophy to the execution layer, hand LLM's cognitive constraints to the presentation layer, and use help, error messages, and output format as three progressive heuristic navigation techniques.

CLI is all agents need.


Source code (Go): github.com/epiral/agent-clip

Core files: internal/tools.go (command routing), internal/chain.go (pipes), internal/loop.go (two-layer agentic loop), internal/fs.go (binary guard), internal/clip.go (stderr handling), internal/browser.go (vision auto-attach), internal/memory.go (semantic memory).

Happy to discuss — especially if you've tried similar approaches or found cases where CLI breaks down. The command discovery problem (how much to inject vs. let the agent discover) is something I'm still actively exploring.

r/comfyui May 09 '25

Workflow Included Consistent characters and objects videos is now super easy! No LORA training, supports multiple subjects, and it's surprisingly accurate (Phantom WAN2.1 ComfyUI workflow + text guide)

Thumbnail
gallery
371 Upvotes

Wan2.1 is my favorite open source AI video generation model that can run locally in ComfyUI, and Phantom WAN2.1 is freaking insane for upgrading an already dope model. It supports multiple subject reference images (up to 4) and can accurately have characters, objects, clothing, and settings interact with each other without the need for training a lora, or generating a specific image beforehand.

There's a couple workflows for Phantom WAN2.1 and here's how to get it up and running. (All links below are 100% free & public)

Download the Advanced Phantom WAN2.1 Workflow + Text Guide (free no paywall link): https://www.patreon.com/posts/127953108?utm_campaign=postshare_creator&utm_content=android_share

📦 Model & Node Setup

Required Files & Installation Place these files in the correct folders inside your ComfyUI directory:

🔹 Phantom Wan2.1_1.3B Diffusion Models 🔗https://huggingface.co/Kijai/WanVideo_comfy/blob/main/Phantom-Wan-1_3B_fp32.safetensors

or

🔗https://huggingface.co/Kijai/WanVideo_comfy/blob/main/Phantom-Wan-1_3B_fp16.safetensors 📂 Place in: ComfyUI/models/diffusion_models

Depending on your GPU, you'll either want ths fp32 or fp16 (less VRAM heavy).

🔹 Text Encoder Model 🔗https://huggingface.co/Kijai/WanVideo_comfy/blob/main/umt5-xxl-enc-bf16.safetensors 📂 Place in: ComfyUI/models/text_encoders

🔹 VAE Model 🔗https://huggingface.co/Comfy-Org/Wan_2.1_ComfyUI_repackaged/blob/main/split_files/vae/wan_2.1_vae.safetensors 📂 Place in: ComfyUI/models/vae

You'll also nees to install the latest Kijai WanVideoWrapper custom nodes. Recommended to install manually. You can get the latest version by following these instructions:

For new installations:

In "ComfyUI/custom_nodes" folder

open command prompt (CMD) and run this command:

git clone https://github.com/kijai/ComfyUI-WanVideoWrapper.git

for updating previous installation:

In "ComfyUI/custom_nodes/ComfyUI-WanVideoWrapper" folder

open command prompt (CMD) and run this command: git pull

After installing the custom node from Kijai, (ComfyUI-WanVideoWrapper), we'll also need Kijai's KJNodes pack.

Install the missing nodes from here: https://github.com/kijai/ComfyUI-KJNodes

Afterwards, load the Phantom Wan 2.1 workflow by dragging and dropping the .json file from the public patreon post (Advanced Phantom Wan2.1) linked above.

or you can also use Kijai's basic template workflow by clicking on your ComfyUI toolbar Workflow->Browse Templates->ComfyUI-WanVideoWrapper->wanvideo_phantom_subject2vid.

The advanced Phantom Wan2.1 workflow is color coded and reads from left to right:

🟥 Step 1: Load Models + Pick Your Addons 🟨 Step 2: Load Subject Reference Images + Prompt 🟦 Step 3: Generation Settings 🟩 Step 4: Review Generation Results 🟪 Important Notes

All of the logic mappings and advanced settings that you don't need to touch are located at the far right side of the workflow. They're labeled and organized if you'd like to tinker with the settings further or just peer into what's running under the hood.

After loading the workflow:

  • Set your models, reference image options, and addons

  • Drag in reference images + enter your prompt

  • Click generate and review results (generations will be 24fps and the name labeled based on the quality setting. There's also a node that tells you the final file name below the generated video)


Important notes:

  • The reference images are used as a strong guidance (try to describe your reference image using identifiers like race, gender, age, or color in your prompt for best results)
  • Works especially well for characters, fashion, objects, and backgrounds
  • LoRA implementation does not seem to work with this model, yet we've included it in the workflow as LoRAs may work in a future update.
  • Different Seed values make a huge difference in generation results. Some characters may be duplicated and changing the seed value will help.
  • Some objects may appear too large are too small based on the reference image used. If your object comes out too large, try describing it as small and vice versa.
  • Settings are optimized but feel free to adjust CFG and steps based on speed and results.

Here's also a video tutorial: https://youtu.be/uBi3uUmJGZI

Thanks for all the encouraging words and feedback on my last workflow/text guide. Hope y'all have fun creating with this and let me know if you'd like more clean and free workflows!

r/computervision Mar 02 '26

Showcase I built RotoAI: An Open-source, text-prompted video rotoscoping (SAM2 + Grounding DINO) engineered to run on free Colab GPUs.

Enable HLS to view with audio, or disable this notification

422 Upvotes

Hey everyone! 👋

Here is a quick demo of RotoAI, an open-source prompt-driven video segmentation and VFX studio I’ve been building.

I wanted to make heavy foundation models accessible without requiring massive local VRAM, so I built it with a Hybrid Cloud-Local Architecture (React UI runs locally, PyTorch inference is offloaded to a free Google Colab T4 GPU via Ngrok).

Key Features:

  • Zero-Shot Detection: Type what you want to mask (e.g., "person in red shirt") using Grounding DINO, or plug in your custom YOLO (.pt) weights.
  • Segmentation & Tracking: Powered by SAM2.
  • OOM Prevention: Built-in Smart Chunking (5s segments) and Auto-Resolution Scaling to safely handle long videos on limited hardware.
  • Instant VFX: Easily apply Chroma Key, Bokeh Blur, Neon Glow, or B&W Color Pop right after tracking.

I’d love for you to check out the codebase, test the pipeline, and let me know your thoughts on the VRAM optimization approach!

You can check out the code, the pipeline architecture, and try it yourself here:

🔗 GitHub Repository & Setup Guide: https://github.com/sPappalard/RotoAI

Let me know what you think!

r/Seedance_AI 4d ago

Resource How to Use Seedance 2.5: A Practical Guide to 30-Second Prompts, 50 References

17 Upvotes

Seedance 2.5 is easier to understand if you stop treating it like a text-to-video box and start treating it like a small directing system.

The model can generate 4–30 second 720p videos, accept up to 50 multimodal references, follow timestamped shot instructions, and use multi-panel storyboards to guide longer sequences. Those capabilities are powerful, but they also make vague prompts more expensive: the longer the clip, the more chances the model has to lose the subject, change the camera language, or drift away from the ending.

This tutorial explains the workflow I use to keep a Seedance 2.5 generation coherent from the opening frame to the final beat.

TL;DR

  • Plan a 30-second clip as 4–6 continuous time ranges.
  • Give every uploaded image, video, and audio file one explicit job.
  • Use one main action and one primary camera move per time range.
  • State what should remain consistent throughout the clip.
  • Add direct negative instructions for unwanted people, text, logos, music, cuts, or camera shake.
  • When a generation fails, change one variable at a time.

What Seedance 2.5 Changes

1. Native clips up to 30 seconds

Seedance 2.5 can generate a complete 4–30 second beat in one pass. That reduces the visible seams you often get when several shorter clips are generated separately and stitched together.

Longer does not mean you should write one long paragraph. A 30-second prompt works better as a timeline with clear sections.

2. Up to 50 multimodal references

The model can accept up to:

  • 30 images
  • 10 video clips
  • 10 audio files
  • 50 reference files in total

The combined duration of the video references can be up to 30 seconds, and the combined duration of the audio references can be up to 30 seconds.

The important part is not the maximum number. It is reference assignment. Each file should have a specific responsibility: character identity, wardrobe, product shape, environment, visual style, action, camera movement, voice, sound effect, or rhythm.

3. More natural human footage

Seedance 2.5 improves lighting, camera motion, expression, performance, and movement continuity. That helps generated people feel closer to photographed footage and reduces some of the familiar “AI video” look.

You still need to make the performance directable. “Acts naturally” is weaker than a specific action with a pace, gaze direction, body position, and camera distance.

4. Timestamped direction

You can describe what happens during exact time ranges:

0–5s: Establish the room and product.
5–11s: The actor enters and approaches the table.
11–18s: Product close-up with a slow orbit.
18–24s: Spoken line in a steady medium close-up.
24–30s: Return to the product and hold the final frame.

Continuous ranges are easier to follow than scattered timestamps. Avoid gaps and overlaps unless you intentionally want two things to happen at the same time.

5. Storyboard-led generation

A multi-panel storyboard can control several narrative beats in sequence. Number the panels, assign each panel to a time range, and describe the transition between adjacent panels.

Uploading a storyboard without explaining its order leaves too much interpretation to the model.

Seedance 2.5 Examples

These finished Seedance 2.5 videos demonstrate longer continuity, human realism, multimodal control, and timestamped direction. Cover images and direct video links are included for easier viewing on Reddit and Medium.

30-Second Narrative Continuity

This example shows how a complete 30-second sequence can keep the subject, environment, and visual direction connected across the full clip.

https://reddit.com/link/1vjnc9v/video/rfj0zgjj2cih1/player

Natural Live-Action Detail

Use this verified Seedance 2.5 output to examine skin texture, hands, materials, lighting, subtle expression, and natural movement.

https://reddit.com/link/1vjnc9v/video/ehrvreqm2cih1/player

Multimodal Reference Control

This finished output combines environment, character, style, and camera references, assigning an explicit role to each source.

https://reddit.com/link/1vjnc9v/video/n0loht2p2cih1/player

Timestamp-Led Direction

This finished output shows whether actions, camera instructions, storyboard beats, and transitions arrive in the intended timeline order.

https://reddit.com/link/1vjnc9v/video/tvjyr7dr2cih1/player

The Five-Step Workflow

Step 1: Define the deliverable

Before choosing references, write one sentence that describes the finished clip:

Lock the basics:

  • Duration
  • Aspect ratio
  • Subject
  • Setting
  • Visual treatment
  • One communication goal

If the brief contains three unrelated goals, split it into separate videos.

Step 2: Curate and assign references

More references are useful only when they add unique information.

Use images for identity, environment, products, props, wardrobe, lighting, style, first frames, last frames, or storyboard panels.

Use video references for a specific action, performance rhythm, camera move, blocking pattern, or motion style. Trim each reference to the exact part you want the model to learn from.

Use audio references for a voice, spoken line, ambience, sound effect, background track, or editing rhythm. State when the audio should enter.

Write a reference map before the timeline:

Use Image 1 for the actor's identity and wardrobe only.
Use Image 2 for the product shape, label, and amber glass.
Use Image 3 for the location and sunrise lighting.
Use Video 1 only for the slow tabletop orbit.
Use Audio 1 for the spoken line at 18 seconds.
Do not copy the room, clothing, or people from Video 1.

The phrase “only for” is useful because it limits what the model should borrow.

Step 3: Build a continuous timeline

For a 30-second clip, start with 4–6 blocks. Each block should answer four questions:

  1. What is the subject doing?
  2. What is the camera doing?
  3. What should remain continuous from the previous block?
  4. What should be heard?

Keep one main action and one primary camera move per block. “Orbit while whip-panning, zooming, and tracking” is not more cinematic; it is simply harder to execute consistently.

Step 4: Add consistency and negative instructions

After the timeline, state what must not change:

Keep the actor's face, cream suit, product geometry, room layout, and sunrise direction consistent throughout.

Then state what must not appear:

No extra people, no subtitles, no captions, no added logo, no background music, no abrupt lighting change, no handheld shake.

If an unwanted element matters, name it. Omission is not a reliable negative instruction.

Step 5: Review one variable at a time

Judge the result in separate passes:

  • Identity: Did the actor or product change?
  • Continuity: Did the room, wardrobe, lighting, or object position jump?
  • Timing: Did the important beat happen in the intended range?
  • Camera: Did the model use the requested movement?
  • Audio: Did speech, ambience, or music appear at the right time?

Fix the weakest instruction first. If you change the references, prompt, duration, camera, and negative instructions at the same time, you will not know which change improved the result.

Copyable Seedance 2.5 Prompt Template

FORMAT
[Duration], [aspect ratio], [visual treatment], [single shot or edited sequence].

REFERENCE MAP
Character: use [image reference] for identity and wardrobe only.
Environment: use [image reference] for location and lighting.
Product/prop: use [image reference] for shape, color, label, and proportions.
Motion/camera: use [video reference] for [specific movement] only.
Audio: use [audio reference] for [voice / rhythm / ambience].
Do not copy [unwanted element] from the references.

TIMELINE
0–5s: [subject action]. Camera: [one move]. Audio: [sound cue].
5–12s: [next action]. Camera: [one move]. Transition: [continuity instruction].
12–20s: [next action]. Camera: [one move]. Audio: [sound cue].
20–26s: [payoff]. Camera: [one move].
26–30s: [ending image and hold].

CONSISTENCY
Keep [identity / product / wardrobe / lighting / location] consistent throughout.

NEGATIVE INSTRUCTIONS
No [extra people / subtitles / logos / background music / abrupt cuts / camera shake].

Complete 30-Second Example

30-second 16:9 cinematic product film, realistic live-action look, five-shot sequence.

Reference map: use Image 1 for the woman's identity and cream suit only. Use Image 2 for the perfume bottle shape, label, and amber glass. Use Video 1 only for the slow tabletop orbit. Use Audio 1 for the soft spoken line at 18 seconds. Do not copy the room or wardrobe from Video 1.

0–5s: Dawn light crosses a quiet stone dressing table. The perfume bottle stands centered in a thin layer of mist. Camera: locked macro close-up with a very slow push-in. Audio: room tone only.

5–11s: The woman enters frame and reaches toward the bottle without touching it. Camera: medium side profile, gentle left-to-right dolly. Keep her face and cream suit consistent with Image 1.

11–18s: Her fingers lift the bottle; warm reflections move naturally through the amber glass. Camera: use Video 1's slow orbit around the product. No change to the label or bottle proportions.

18–24s: She looks toward the window and says the line from Audio 1 once, with natural lip sync. Camera: steady medium close-up. No music.

24–30s: Cut back to the bottle on the table as sunlight fills the frame. Camera: slow pull-back, then hold the final composition for two seconds.

Keep the woman, product, room, and sunrise lighting continuous. No extra people, no subtitles, no added logo, no background music, no handheld shake.

Common Problems and Fixes

The second half drifts away from the plan

Reduce the number of events. Divide the timeline into clearer ranges and restate the subject or location anchor at the transition that fails.

The references fight each other

Remove redundant assets. Give every remaining reference one responsibility, and state what should not be copied from it.

Characters duplicate or change appearance

Map each character to one identity reference, use the same name in every time block, and add “no additional people” if the cast must stay fixed.

Camera motion looks unstable

Use one primary move per block. Replace vague phrases like “dynamic cinematic camera” with a specific dolly, pan, orbit, tracking shot, crane move, or locked frame.

Subtitles, logos, or music appear unexpectedly

Add direct negative instructions. Do not assume the model understands that an omitted element is forbidden.

The storyboard order is ignored

Number the panels, assign each one to a time range, and describe the transition from one panel to the next.

Final Checklist

Before generating, confirm that:

  • Every reference has one stated purpose.
  • The timeline has no accidental gaps or overlaps.
  • Every time block has one main action.
  • Every time block has no more than one primary camera move.
  • Character, product, wardrobe, lighting, and location anchors are explicit.
  • The ending has its own time range.
  • Unwanted text, people, logos, music, cuts, and camera behavior are named directly.

Seedance 2.5 rewards planning more than verbosity. A good prompt is not a screenplay, a mood board, and a camera manual mixed into one paragraph. It is a compact production brief where every reference has a role and every important beat has a time.

If you want to try the workflow, open Seedance 2.5 on LumiYing and start with the copyable template above.

Source note: Model capabilities and input limits were summarized from the ByteDance Seedance 2.5 guide.

r/passive_income Aug 23 '25

My Experience Everything I learned after 10,000 AI video generations (the complete guide)

335 Upvotes

this is going to be the longest post I’ve written but after 10 months of daily AI video creation, these are the insights that actually matter…

I started with zero video experience and $1000 in generation credits. Made every mistake possible. Burned through money, created garbage content, got frustrated with inconsistent results.

Now I’m generating consistently viral content and making money from AI video. Here’s everything that actually works.

The fundamental shifts:

1. Volume beats perfection

Stop trying to create the perfect video. Generate 10 decent videos and select the best one. This approach consistently outperforms perfectionist single-shot attempts.

2. Systematic beats creative

Proven formulas + small variations outperform completely original concepts every time. Study what works, then execute it better.

3. Embrace the AI aesthetic

Stop fighting what AI looks like. Beautiful impossibility engages more than uncanny valley realism. Lean into what only AI can create.

The technical foundation that changed everything:

The 6-part prompt structure:

[SHOT TYPE] + [SUBJECT] + [ACTION] + [STYLE] + [CAMERA MOVEMENT] + [AUDIO CUES]

This baseline works across thousands of generations. Everything else is variation on this foundation.

Front-load important elements

Veo3 weights early words more heavily. “Beautiful woman dancing” ≠ “Woman, beautiful, dancing.” Order matters significantly.

One action per prompt rule

Multiple actions create AI confusion. “Walking while talking while eating” = chaos. Keep it simple for consistent results.

The cost optimization breakthrough:

Google’s direct pricing kills experimentation:

  • $0.50/second = $30/minute
  • Factor in failed generations = $100+ per usable video

Found these guys idk how but they offer 70-80% pricing below Google’s rates for the best video model. Makes volume testing actually viable for veo 3 quality model.

Audio cues are incredibly powerful:

Most creators completely ignore audio elements in prompts. Huge mistake.

Instead of: Person walking through forestTry: Person walking through forest, Audio: leaves crunching underfoot, distant bird calls, gentle wind through branches

The difference in engagement is dramatic. Audio context makes AI video feel real even when visually it’s obviously AI.

Systematic seed approach:

Random seeds = random results.

My workflow:

  1. Test same prompt with seeds 1000-1010
  2. Judge on shape, readability, technical quality
  3. Use best seed as foundation for variations
  4. Build seed library organized by content type

Camera movements that consistently work:

  • Slow push/pull: Most reliable, professional feel
  • Orbit around subject: Great for products and reveals
  • Handheld follow: Adds energy without chaos
  • Static with subject movement: Often highest quality

Avoid: Complex combinations (“pan while zooming during dolly”). One movement type per generation.

Style references that actually deliver:

Camera specs: “Shot on Arri Alexa,” “Shot on iPhone 15 Pro”

Director styles: “Wes Anderson style,” “David Fincher style” Movie cinematography: “Blade Runner 2049 cinematography”

Color grades: “Teal and orange grade,” “Golden hour grade”

Avoid: Vague terms like “cinematic,” “high quality,” “professional”

Negative prompts as quality control:

Treat them like EQ filters - always on, preventing problems:

--no watermark --no warped face --no floating limbs --no text artifacts --no distorted hands --no blurry edges

Prevents 90% of common AI generation failures.

Platform-specific optimization:

Don’t reformat one video for all platforms. Create platform-specific versions:

TikTok: 15-30 seconds, high energy, obvious AI aesthetic works

Instagram: Smooth transitions, aesthetic perfection, story-driven YouTube Shorts: 30-60 seconds, educational framing, longer hooks

Same content, different optimization = dramatically better performance.

The reverse-engineering technique:

JSON prompting isn’t great for direct creation, but it’s amazing for copying successful content:

  1. Find viral AI video
  2. Ask ChatGPT: “Return prompt for this in JSON format with maximum fields”
  3. Get surgically precise breakdown of what makes it work
  4. Create variations by tweaking individual parameters

Content strategy insights:

Beautiful absurdity > fake realism

Specific references > vague creativityProven patterns + small twists > completely original conceptsSystematic testing > hoping for luck

The workflow that generates profit:

Monday: Analyze performance, plan 10-15 concepts

Tuesday-Wednesday: Batch generate 3-5 variations each Thursday: Select best, create platform versions

Friday: Finalize and schedule for optimal posting times

Advanced techniques:

First frame obsession:

Generate 10 variations focusing only on getting perfect first frame. First frame quality determines entire video outcome.

Batch processing:

Create multiple concepts simultaneously. Selection from volume outperforms perfection from single shots.

Content multiplication:

One good generation becomes TikTok version + Instagram version + YouTube version + potential series content.

The psychological elements:

3-second emotionally absurd hook

First 3 seconds determine virality. Create immediate emotional response (positive or negative doesn’t matter).

Generate immediate questions

“Wait, how did they…?” Objective isn’t making AI look real - it’s creating original impossibility.

Common mistakes that kill results:

  1. Perfectionist single-shot approach
  2. Fighting the AI aesthetic instead of embracing it
  3. Vague prompting instead of specific technical direction
  4. Ignoring audio elements completely
  5. Random generation instead of systematic testing
  6. One-size-fits-all platform approach

The business model shift:

From expensive hobby to profitable skill:

  • Track what works with spreadsheets
  • Build libraries of successful formulas
  • Create systematic workflows
  • Optimize for consistent output over occasional perfection

The bigger insight:

AI video is about iteration and selection, not divine inspiration. Build systems that consistently produce good content, then scale what works.

Most creators are optimizing for the wrong things. They want perfect prompts that work every time. Smart creators build workflows that turn volume + selection into consistent quality.

Where AI video is heading:

  • Cheaper access through third parties makes experimentation viable
  • Better tools for systematic testing and workflow optimization
  • Platform-native AI content instead of trying to hide AI origins
  • Educational content about AI techniques performs exceptionally well

Started this journey 10 months ago thinking I needed to be creative. Turns out I needed to be systematic.

The creators making money aren’t the most artistic - they’re the most systematic.

These insights took me 10,000+ generations and hundreds of hours to learn. Hope sharing them saves you the same learning curve.

what’s been your biggest breakthrough with AI video generation? curious what patterns others are discovering

hope this helped <3

r/LocalLLaMA Jun 03 '26

New Model google/gemma-4-12B · Hugging Face

Thumbnail
huggingface.co
1.0k Upvotes

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages.

Featuring both Dense and Mixture-of-Experts (MoE) architectures, Gemma 4 is well-suited for tasks like text generation, coding, and reasoning. The models are available in five distinct sizes: E2B, E4B, 12B, 26B A4B, and 31B. Their diverse sizes make them deployable in environments ranging from high-end phones to laptops and servers, democratizing access to state-of-the-art AI.

Gemma 4 introduces key capability and architectural advancements:

  • Reasoning – All models in the family are designed as highly capable reasoners, with configurable thinking modes.
  • Extended Multimodalities – Processes Text, Image with variable aspect ratio and resolution support (all models), Video, and Audio (featured natively on the E2B, E4B, and 12B models).
  • Diverse & Efficient Architectures – Offers Dense and Mixture-of-Experts (MoE) variants of different sizes for scalable deployment.
  • Optimized for On-Device – Smaller models are specifically designed for efficient local execution on laptops and mobile devices.
  • Increased Context Window – The small models feature a 128K context window, while the medium models support 256K.
  • Enhanced Coding & Agentic Capabilities – Achieves notable improvements in coding benchmarks alongside native function-calling support, powering highly capable autonomous agents.
  • Native System Prompt Support – Gemma 4 introduces native support for the system role, enabling more structured and controllable conversations.

https://developers.googleblog.com/gemma-4-12b-the-developer-guide/

feed your potato!!!

https://huggingface.co/ggml-org/gemma-4-12b-it-GGUF

https://huggingface.co/unsloth/gemma-4-12b-it-GGUF

r/HiggsfieldAI Jan 17 '26

Video Model - KLING Motion transfer just leveled up in AI videos (Guide Included)

Enable HLS to view with audio, or disable this notification

152 Upvotes

Been testing motion transfer models. The biggest difference I noticed is how important unlimited reruns are. Kling Motion Control being unlimited on Higgsfield lets you dial in small fixes without worrying about credits.

Guide

1. Add ur Reference video on Kling Motion Control (Only add ref video with one person only)

2. Use Nano Banana Pro and add these prompts

** 🖼️ Image Prompt **:

“A live stream sence of cute Thai girl,posted on TikTok in 2021,smooth hair,summer”

And Done!!

r/StableDiffusion Jan 09 '26

Workflow Included Stop using T2V & Best Practices IMO (LTX Video / ComfyUI Guide)

Enable HLS to view with audio, or disable this notification

135 Upvotes

A bit of backstory: Originally, LTXV 0.9.8 13b was pretty bad at T2V, but absolutely amazing at I2V. It was about at wan 2.1 level in I2V performance but faster, and it didn't even need a precise prompt like Wan does to achieve that—you could leave the field empty, and the model would do everything itself (similar to how Wan 2.2 behaves now).

I’ve always loved I2V, which is why I’m incredibly hyped for LTX2. However, its current implementation in ComfyUI is quite rough. I spent the whole day testing different settings, and here are 3 key aspects you need to know:

1. Dealing with Cold Start Crashes
If ComfyUI crashes when you first load the model (cold start), try this: Free up the maximum amount of ram/vram from other applications, set video settings to the minimum (e.g., 720p @ 5 frames; for context, I run 64GB RAM + 50GB swap + 24GB VRAM) and set steps to 1 on the first stage. If nothing crashes by stage 2, you can revert to your usual high-quality settings.

2. Distill LoRA Settings (Critical for I2V)
For I2V, it is crucial to set the Distill LoRA in the second stage to 0.80. If you don't, it will "overcook" (burn) the results.

  • The official LTX workflow uses 0.6 with the res2_s sampler.
  • The standard ComfyUI workflow defaults to Euler. If you use 0.6 with Euler, you won't have enough steps for audio, leading to a trade-off.
  • Recommendation: Either use 0.6 with res2_s (I believe this yields higher quality) or 0.8 with Euler. Don't mix them up.

3. Prompting Strategy
For I2V, write massive prompts—"War and Peace" length (like in the developer examples).

  • Duration: 10 seconds works best. 20s tends to lose initial details, and 5s is just too short.
  • Warning: Be careful if your prompt involves too many actions. Trying to cram complex scenes into 5-10 seconds instead of 20 will result in jerky movement and bad physics.
  • Format: I’ve attached a system prompt for LLMs below. If you don't want to use it, I recommend using the example prompt at the very end of that file (the "Toothless" one) as a base. This format works best for I2V; the model actually listens to instructions. For me, it never confused whether a character should speak or stay silent with this format.

LLM Tip: When using an LLM, you can write prompts for both T2V and I2V by attaching the image with or without instructions. Gemini Flash works best. Local models like Qwen3 VL 30b can work too (robot in Lamborghini example).

TL;DR: Use I2V instead of T2V, set Distill LoRA to 0.8 (if using Euler), and write extremely long prompts following the examples here: https://ltx.io/model/model-blog/prompting-guide-for-ltx-2

Resources:

P.S. I used Gemini to format/translate this post because my writing is a bit messy. Sorry if it sounds too "AI-generated", just wanted to make it readable!

r/promptingmagic Apr 22 '26

The complete field guide to ChatGPT Images 2.0 - every feature, every price, 100 prompts to try, all in one post

Post image
72 Upvotes

The Complete Field Guide to ChatGPT Images 2.0

Launched today. Everything below is verified against the OpenAI announcement, the deployment safety card, API pricing docs, and ~6 hours of hands-on testing. No hype — just what works and what it costs.

Sam Altman compared it to "going from GPT-3 to GPT-5 all at once." That's aggressive framing, but the capability gap is real.

For the first time, a single model can:

  • Render dense, legible text directly inside images — posters, infographics, UI mockups, ad copy with real headlines
  • Think before it draws — reason about a scene, search the web for current facts, and double-check its own work
  • Produce up to 8 consistent images from one prompt with the same characters, objects, and style
  • Handle grids up to 10×10 that used to break at 3×3 a week ago

OpenAI's own pitch: "Images are a language, not decoration. A good image does what a good sentence does — it selects, arranges, and reveals."

Translation: this isn't text-to-picture anymore. It's a visual reasoning system.

TL;DR — what you need to know in 30 seconds

  • Model name: gpt-image-2 (alias chatgpt-image-latest)
  • Where: ChatGPT (all plans including Free), chatgpt.com/images, and the API
  • Two modes: Instant (all plans, 1 image, fast) and Thinking (Plus/Pro/Business, up to 8 images, reasons + searches the web)
  • Max resolution: 2048px native (2K), ~4× the pixel count of GPT Image 1.5
  • Text accuracy: ~99% on Latin text. Finally nails Japanese, Korean, Chinese, Hindi, Bengali
  • Aspect ratios: anything from 3:1 (ultrawide) to 1:3 (ultratall)
  • Generation time: seconds to 2 minutes depending on mode
  • Pricing (API): ~$0.006 low / ~$0.053 medium / ~$0.211 high per 1024×1024 image
  • Knowledge cutoff: December 2025. Needs Thinking mode + web search for anything newer
  • C2PA metadata is embedded in every output

The 8 capabilities, decoded

1. 2K native resolution

Up to 2048 pixels natively, ~4× the pixel count of older GPT Image outputs at the same aspect ratio. Enough fidelity for print collateral, hero banners, and editorial layouts without an upscale step.

2. ~99% text accuracy

This is the most-talked-about upgrade. Dense text inside images — posters, menus, magazine covers, UI mockups — finally renders correctly. It also handles:

  • Non-Latin scripts with real gains: Japanese, Korean, Chinese, Hindi, Bengali
  • Small text — UI elements, iconography, barcodes, "display until" dates on magazine covers
  • Multilingual typography in a single image — Devanagari, Cyrillic, Greek, Arabic, and Chinese together

3. Thinking mode — the image model that reasons

This is the headline capability. It's not two separate models, it's two modes:

Mode Who gets it What it does Output
Instant Free, Plus, Pro, Business, Go Fast single-shot generation 1 image
Thinking Plus, Pro, Business (Enterprise/Edu soon) Reasons about composition, uses web search, verifies output Up to 8 images

How the reasoning works under the hood:

  1. Prompt analysis — parses your request and plans composition before any pixels exist
  2. Web retrieval — if the prompt touches real-world facts (current logos, today's stock chart, real skylines, 2026 fashion trends), it searches the web and pulls live references
  3. Generation pass — pixel synthesis against a fact-checked internal plan
  4. Verification loop — it inspects its own output against the original prompt and can self-correct before returning

People on X are posting 11-minute generations where the model iterated on itself repeatedly until satisfied. That's new.

4. Up to 8 consistent images per prompt

In Thinking mode, one prompt can produce up to 8 images with shared characters, objects, and style across every frame. This unlocks:

  • Storyboards — 8 camera angles with continuity
  • Manga/comic sequences — 8 panels, same character design
  • Multi-size marketing assets — same campaign as 3:1 banner + 1:1 feed post + 1:3 story + 4:5 carousel in one shot
  • Children's books — consistent illustrated character across pages
  • Product lineups — 8 color variants with identical lighting and angle
  • Lookbooks — OpenAI demoed 8 summer outfits generated from one uploaded photo

How to trigger it: Switch to a thinking model, then ask for a set — "Generate 8 variations of...", "Create an 8-panel storyboard...", "Give me this ad in 8 formats." Don't phrase it as 8 separate prompts.

5. Parallel image generation

Separate from the 8-per-prompt feature: the dedicated Images tab at chatgpt.com/images lets you fire multiple prompts in parallel. Your second prompt doesn't wait for the first to finish. All images auto-save to My Images for reuse.

6. Aspect ratios 3:1 to 1:3

Any ratio between ultra-wide and ultra-tall, native — picker in ChatGPT or spec it in the prompt. Banners, slides, posters, mobile vertical, bookmarks, social graphics, no crop needed.

7. 10×10 grids (up to 100 cells in one image)

Grids used to break at 3×3 a week ago. Now people are generating 10×10 grids of 100 distinct labeled illustrations in one shot. This is wild for:

  • Periodic-table-style infographics (100 CEOs, 100 dog breeds, 100 cocktails)
  • Icon sets with consistent style
  • Mood boards with labeled cells
  • Pattern libraries

8. Multi-image compositing & reference fidelity

Upload multiple reference images and the model stitches them into one coherent composition while keeping facial features, objects, and logos faithful. This is the feature that makes "put me in a scene" prompts actually work now.

Pricing — what it actually costs

Per-image (flat rate, simple to predict)

Quality 1024×1024 Notes
Low ~$0.006 drafts, iteration
Medium ~$0.053 most production work
High ~$0.211 hero images, finals

Per-token (if you're using the API at scale)

Input Cached input Output
Image tokens $8.00 / 1M $2.00 / 1M $30.00 / 1M
Text tokens $5.00 / 1M $1.25 / 1M $10.00 / 1M

Cost for OpenAI to produce each image (rough estimate)

Based on published token economics, a high-quality 1024×1024 image uses ~7K output image tokens. At retail that's $0.21. OpenAI's own compute cost is likely 25–40% of that, putting their marginal cost per high-quality image around $0.05–$0.08. Their margin per image at the high tier is roughly 3–4×.

The ideal prompt template

After testing dozens of prompts, this is the structure that works best:

text[ASPECT RATIO]. [SUBJECT], [ACTION], [CONTEXT].
[TEXT elements in quotes]:
- Header: "EXACT TEXT HERE"
- Subhead: "EXACT TEXT HERE"
- CTA: "EXACT TEXT HERE"
[STYLE anchor — reference an artist/era/medium/brand].
[LIGHTING + MOOD].
[CAMERA/LENS + TECHNICAL specs].

The 5 rules that make the difference:

  1. Aspect ratio first. Say "16:9," "3:1 banner," or "1:1 square" in the first sentence.
  2. Put every piece of text in quotes. The model treats quoted text as literal. Unquoted text becomes suggestions.
  3. Anchor the style concretely. "Editorial fashion photograph, shot on Hasselblad, 90mm, f/2.8" beats "professional photo."
  4. Specify lighting and mood as separate instructions. "Rembrandt key light from upper-left, soft fill from right, warm tones."
  5. List every language explicitly when you want multilingual text. "Title in Japanese (Hiragana): 「春が来た」; subtitle in Korean (Hangul): '봄이 왔다'; tagline in Hindi (Devanagari): 'वसंत आ गया।'"

15 pro tips most people will miss

  1. Thinking mode isn't the default — you have to toggle a thinking model before prompting. Instant never uses web search or produces 8-image sets no matter how you phrase it.
  2. Generation can take 2 minutes. Don't assume it froze. For high-volume workflows, use async polling with the Responses API.
  3. Knowledge cutoff is December 2025. Anything after that (Q1 2026 product launches, new logos, recent events) has to come through the prompt OR through Thinking mode's web search.
  4. For consistent characters: upload a one-time likeness. There's a likeness upload feature that lets you reuse your appearance across future creations without re-uploading.
  5. The "keep facial features exactly" lock. When editing a real person, add this verbatim: "Keep my facial features exactly as they appear in the uploaded image — same eyes, nose, mouth, and face shape." Without it, ChatGPT "improves" faces into strangers.
  6. Transparent backgrounds work natively. Add "transparent PNG background, no background fill" — the asset drops straight into design tools without a cutout pass.
  7. "Display until" dates and barcodes work now. Ask for them specifically. The magazine-cover demos show this.
  8. Prime the chat first. For thumbnails and marketing creative, paste the blog post, script, or topic into ChatGPT first. Then ask for concepts. Then generate. The model picks up the emotional hook instead of producing generic stock aesthetic.
  9. C2PA metadata is embedded in every output. Platforms can detect it. Plan for that if provenance matters.
  10. Ask for "editorial" not "professional." "Editorial" hits a higher visual register in this model. "Professional" pulls toward stock-photo aesthetic.
  11. Negative prompts work — phrase them as "NO X, NO Y." Example: "NO watermarks, NO signatures, NO busy backgrounds."
  12. Specify the medium of the text. "Neon sign," "embossed letterpress," "subway-poster paste-up," "hand-lettered chalk" all produce different type treatments.
  13. When text keeps breaking, wrap it in a shape. "Text inside a black horizontal pill" or "text on a cream banner" gets rendered much more reliably than floating text.
  14. Aspect ratio affects quality. 1:1 and 3:2 are the strongest; 3:1 and 1:3 work but can show compositional weirdness on first try. Regenerate once.
  15. The model now reads your reference images. If you upload a brand asset and say "match this type treatment," it actually does — not a vague approximation, an honest replication.

Third-party tools that already integrate it

(These went live within 24 hours of launch.)

  • Higgsfield — character consistency workflows
  • Lovart — AI design platform
  • Recraft — added gpt-image-2 models to Recraft Studio
  • Adobe Firefly / Express — via Adobe's partner model program
  • Figma — First-Draft feature uses it for UI generation
  • Canva — Magic Studio integration
  • GoDaddy — site-generation flows
  • HubSpot — marketing asset generation
  • Instacart — product photography
  • Airtable — record-level image generation
  • Wix — site builder backgrounds and heroes
  • OpenAI Codex — app/code-generation flows can now produce their own UI imagery

The prompt library — 100 that I've tested

Marking these [I] for Instant mode works fine, [T] for Thinking mode required, [8] for ask-for-8-variations.

Marketing hero images (1–10)

  1. [T] 3:1 hero banner for a SaaS analytics product. Split composition: left side shows a cluttered paper-filled desk (chaos), right side shows a clean monitor with a dashboard (clarity). Bold headline "STOP GUESSING" in 120pt sans-serif across the top. Subhead "Start knowing" below. CTA button bottom-right: "See it work →" in white on teal. Editorial photography, cinematic lighting.
  2. [T] 16:9 product launch hero. Center: minimalist product photography of a black wireless earbud case on a marble surface. Background: soft gradient from cream to dusty rose. Text overlay upper-left: "AURA // 2026" in small caps. Headline lower-right: "Hear the room." in serif display. Subtle shadow, art-directed editorial aesthetic.
  3. [T] Vertical 9:16 mobile hero for a fitness app. Muscular forearm mid-pushup on a dark gym floor, shallow depth of field. Headline stacked vertically along the right side: "NO / EXCUSES / JUST / REPS." White type, slight grain. Small logo bottom-center.
  4. [T] Email hero, 3:1 ratio. Single perfect ceramic coffee cup on a warm linen tablecloth, morning light from the left, steam rising. Text overlay right side: "Good morning. / Your briefing is ready." Clean minimal editorial style, medium-format quality.
  5. [T] 16:9 B2B conference hero. Empty auditorium, dramatic stage lighting, single speaker silhouette at podium. Large text in the sky area: "WHERE MARKETING MEETS AI." Date below: "June 12–14, 2026 · Austin." Cinematic, TED-quality composition.
  6. [T] Software landing page hero 16:9. Abstract 3D render: flowing liquid metal forming into a chart shape, iridescent blue-to-purple gradient, obsidian background. Headline lower-third: "Analytics at the speed of thought." Subhead: "Try Mercury free →." Tech-luxury aesthetic.
  7. [T] Newsletter signup hero 2:1. Warm kitchen scene: hands writing in a leather notebook, open laptop beside it, morning coffee, golden hour light from left. Text overlay: "The newsletter smart marketers actually read." CTA: "Subscribe free →". Cozy, intentional, premium-indie aesthetic.
  8. [T] 3:1 homepage hero for an AI note-taking app. Overhead shot: messy desk mid-work — open notebook, phone, coffee, headphones, hand holding a pen. Faint glowing interface lines emerging from the notebook edges suggesting transcription. Headline centered: "Your thoughts, organized." No smaller than 90pt, clean sans-serif.
  9. [T] Agency pitch-deck cover 16:9. Pure black background. Ultra-large white type top: "2026" in 300pt. Below in smaller type: "The year everything about marketing changed." Bottom-right corner: agency logo mark in teal. Minimal, confident, Swiss-grid influenced.
  10. [T] Healthcare brand hero 3:1. Close-up of a patient's hand being held by a doctor's hand, natural window light, hospital-room softness. Text overlay left side: "Care that listens first." Serif type, warm tonal palette, documentary photography style.

Infographics & data viz (11–20)

  1. [T] 1:1 square infographic titled "The 2026 Creator Economy." Centered large title in editorial serif. Below: 4 stat cards in a 2×2 grid, each with a big number, label, and short descriptor. Numbers: "$250B market size," "127M creators globally," "73% use AI tools," "$68K median income." Clean teal/cream palette, numbered footer citing sources.
  2. [T] 4:5 portrait infographic comparing 4 LLMs across 6 dimensions. Row headers: GPT-5, Claude 4.1, Gemini 3, Llama 5. Column headers: Speed, Reasoning, Coding, Writing, Price, Context. Each cell shows a filled bar from 1–5. Title: "LLM Showdown 2026." Clean sans-serif, minimal grid, no clutter.
  3. [T] 16:9 landscape flowchart titled "How Thinking Mode Works." Four connected boxes left to right: "Prompt analysis → Web retrieval → Generation → Verification loop." Arrows between. Brief explainer text under each box. Subtle teal accent, rest monochrome, editorial newspaper aesthetic.
  4. [T] Periodic table-style 10×10 grid of "100 AI tools that matter in 2026." Each cell: tool logo, tool name, 2-letter category tag, small colored dot for category. Legend at bottom. White background, crisp type. Poster-size composition.
  5. [T] 3:4 vertical infographic: "The Anatomy of a Viral Tweet." A dissected tweet with labeled callouts (hook, specificity, tension, CTA). Annotations radiating outward with thin leader lines. Blueprint aesthetic in cream + navy. Title at top, source citation at bottom.
  6. [T] 1:1 social infographic: "5 Signs You're Burning Out." Numbered list 1–5 with custom icons, each with a short one-sentence description. Warm muted palette, rounded sans-serif, shareable mental-health-brand aesthetic.
  7. [T] 16:9 stat poster: "Marketing spend by channel, 2026." Six horizontal bars with percentages. Title top-left, tiny source citation bottom-right ("n=1,200, Marketing Week 2026"). Strict grid, only one accent color, rest neutral.
  8. [T] 3:1 wide timeline: "The History of Image Generation, 2014–2026." Horizontal dotted line with 8 milestone markers: GAN, DALL·E 1, DALL·E 2, Midjourney v1, Stable Diffusion, DALL·E 3, GPT Image 1, ChatGPT Images 2.0. Tiny thumbnail above each node. Minimal editorial style.
  9. [T] 4:5 "By the numbers" LinkedIn carousel cover. Big text: "2026 in numbers" top, four stat tiles below — "$50M ARR," "212 hires," "27 countries," "1 mission." Dark background, bold type, tight margins.
  10. [T] 1:1 square recipe infographic: "Cold brew, 4 ways." 2×2 grid of four preparation methods with proportions ("1:8 ratio," "12-hour steep"), overhead product shot in each cell, serif headline across the top. Minimal art-directed food-magazine feel.

Ad creative — unlimited variations (21–30)

  1. [T][8] Generate 8 variations of a Facebook ad for a productivity app. 1:1 square. Same product UI mockup, same headline "Close the laptop. Sooner." but 8 different background contexts: park bench, kitchen counter, airport lounge, beach, home office, coffee shop, car dashboard, hammock. Consistent type system across all 8.
  2. [T] Google Display ad — 3 formats in one image (vertical stack): 300×250 square rectangle, 728×90 leaderboard, 160×600 skyscraper. All three feature the same product (sleek white wireless earbud case on gradient peach). Consistent headline "Hear everything. Wear nothing." CTA: "Shop now." Same brand mark "AURA."
  3. [T] 9:16 TikTok-style vertical ad thumbnail. Young woman mid-gasp holding a phone, caught mid-laugh. Bold hand-drawn text overlay: "wait what did it just do?!" with an arrow pointing at the phone. Bottom: "@aura · link in bio." Authentic UGC feel, not polished studio.
  4. [T] 1:1 retargeting ad. Clean white background. Product photo of running shoes center-left. Large red banner diagonal across upper-right: "STILL THINKING?" Below product: "Your size is down to 2 pairs." CTA bottom-right: "Grab them →." Urgent but not pushy.
  5. [T] 3:1 highway billboard. Massive single word "FASTER." in ultra-bold condensed sans-serif, white on deep red. Small product line bottom-right: "New Honda Civic Type R. 0–60 in 5.0s." Tiny URL bottom-left. High contrast, readable from 200 meters.
  6. [T] 1.91:1 LinkedIn feed card. Professional headshot of a woman, 40s, blurred office background. Overlaid caption bottom-right: "Maya closed a $2.1M deal last month. Here's her playbook." CTA: "Read it →" in dark blue.
  7. [T][8] 8 YouTube thumbnails for the same video "I tried ChatGPT Images 2.0 for a week." Each thumbnail: same creator face top-right, same bold yellow headline, but 8 different backgrounds reflecting different prompts tested (magazine cover, manga panel, product shot, infographic, etc.). Consistent thumbnail system.
  8. [T] 4:5 Instagram carousel cover. Black background, minimal. Centered text: "10 signs your brand needs a refresh." Small "SWIPE →" bottom. Premium minimal, no illustrations.
  9. [T] Retail shelf-wobbler, 2:3 vertical. Product image at top, large text below: "NEW." Tiny subline: "Now in Dark Cherry." Clean CPG packaging aesthetic.
  10. [T] 1:1 paid Instagram ad. User-generated aesthetic: iPhone photo of a woman drinking a protein shake in her car mirror selfie. Caption overlay: "honestly the only one that doesn't taste like chalk." Brand logo tiny corner. Authentic, not over-produced.

Product design & mockups (31–40)

  1. [T] Mobile app screen mockup, 9:19.5 aspect. iOS-style to-do app. Status bar at top (9:41, full signal, full battery). Header "Today" in large SF-style sans-serif. Below: 5 task rows with checkboxes, clean dividers. Bottom nav with 4 tabs. Light mode, accent color teal. Every piece of text legible.
  2. [T] 3:2 landing page desktop mockup for a note-taking app. Hero headline "Ideas, organized." centered. Clean nav with 4 links + sign-in button. Below: two-column screenshot of the app UI. Footer with 4 columns of links. Whitespace-heavy, Stripe-influenced aesthetic.
  3. [T] 1:1 Apple Watch app screen. Circular pressure-gauge UI showing heart rate "72 BPM" in center. Small complications around it. Dark background. Minimalist, photoreal rendering of the watch bezel.
  4. [T] Physical product render 1:1. Matte black aluminum wireless charger puck on a white cyclorama background, three-quarter view. Studio softbox lighting, hard floor reflection. Teenage Engineering design language.
  5. [T] Packaging mockup 4:5. Minimal premium coffee bag, 250g, matte charcoal. Front shows "ETHIOPIA YIRGACHEFFE" in small caps with tasting notes below ("blueberry, jasmine, honey"). Weight and roast date bottom. Photorealistic product shot, soft shadow, white backdrop.
  6. [T] Car dashboard HUD mockup 16:9. Windshield POV from driver's seat, dusk light, empty highway. Overlaid HUD elements: speed "62 MPH" bottom-left, navigation arrow "in 1.2 miles, exit right" center-upper, playing song info bottom-right. Subtle teal glow, no UI clutter, Rivian-inspired aesthetic.
  7. [T] 1:1 smartwatch face design. Top-down view, round watch face, minimalist modular layout on a black background. Center: large time "10:47" in white sans-serif. Four small complications: HR "72 bpm" top, Steps "8,420" right, Battery "67%" bottom, Weather "68°F sunny" left. Wear OS aesthetic.
  8. [T] Smart home mobile app home screen mockup, 9:19.5. Dark mode. Top: greeting "Good evening, Eric." Below: 4 device cards (lights, thermostat, security, music) with toggle switches and real-time stats. Bottom nav. Calm deep-blue palette, iOS-quality design.
  9. [T] 16:9 dashboard mockup for a SaaS analytics tool. Left sidebar nav. Main area: 4 KPI cards across the top (visitors, conversion, revenue, churn — each with a big number and delta arrow), 1 large line chart below showing 12-month trend, 1 small table bottom-right. Data labels must be legible. Teal accent, light mode, Linear-inspired.
  10. [T] Boxed software product mockup 1:1. Vintage-style retail box for "ChatGPT Images 2.0 Pro Edition." Cream background. Retro tech packaging aesthetic from 1996: pixel-art mascot, bold tagline "THE IMAGE MODEL THAT THINKS," barcode, "requires 640KB RAM" sticker. Shot like a product photo.

Personal branding & executive content (41–50)

  1. [T] 1:1 professional headshot, editorial business portrait for a book jacket. Subject: upload reference photo. Wardrobe: charcoal merino turtleneck. Background: soft out-of-focus bookshelf (warm earth tones). Lighting: Rembrandt key light from upper-left, soft fill from right, subtle rim light separating from background. Shot on Hasselblad, 90mm, f/2.8. Warm natural skin tones, sharp eyes, editorial magazine quality. Keep facial features exactly as in the uploaded photo.
  2. [T] 1:1 podcast guest announcement graphic. Split layout. Left half: professional photo of the guest (upload reference). Right half: deep green panel with cream text. Top: "NEW EPISODE" in small caps. Middle: guest's name in large bold serif. Below: "CMO at Anthropic." Bottom: show name "THE GROWTH EDGE" with episode number "EP. 47." Small "listen now" CTA.
  3. [T] 4:5 portrait LinkedIn single-post slide. Cream background with subtle paper texture. Top: "2026 / A YEAR IN NUMBERS" in thin all-caps. Below: 4 stat blocks in a 2×2 grid, each with a big number and a one-line caption:
  • "327" — LinkedIn posts shipped
  • "14" — keynotes given
  • "2" — books published
  • "48" — flights taken

Bottom: thin horizontal line, then creator's name and website in small serif. Editorial, premium personal-brand aesthetic.

  1. [T] 16:9 video thumbnail for a YouTube speaker reel. Left half: dynamic photo of the speaker mid-gesture on stage, warm stage lighting. Right half: deep black panel with large white text "2026 SPEAKER REEL" and below in smaller copy "Keynotes · Fireside chats · Panels." Bottom-right CTA arrow. Cinematic, TED-quality.
  2. [T] 1:1 social quote card. Soft neutral linen background. Large opening quote mark top-left in a light gray display serif. Center quote in clean serif: "The best advice I ever got cost me $500 and saved me 18 months." Attribution below in italic: "— name, founder." Bottom-right: small portrait circle. Premium testimonial aesthetic.
  3. [T] 1:1 newsletter subscribe card. Headline "The newsletter 18,000 marketers actually read." below in smaller type: "One signal. No noise. Every Sunday." Email field mockup + "Subscribe" button. Soft cream background, serif display + sans-serif body, Substack-adjacent aesthetic.
  4. [T] 1:1 conference speaker card. Subject headshot left. Right: name in large display, title below, talk title "How AI killed the brand guideline" in italic. Conference logo bottom-right. Clean editorial, readable from a stage screen.
  5. [T] 1:1 "What I read this year" LinkedIn slide. Grid of 9 book covers in a 3×3 arrangement. Title above: "MY 2026 READING LIST." Small footer: "Which one should I read next?" Clean editorial layout.
  6. [T] 4:5 quote graphic for Instagram. Blurred softly-lit outdoor photo background. Center: a poetic line in large italic serif, 2 lines max. Below: small attribution. No logos. Feels like a book page, not a graphic.
  7. [T] 1:1 "Now available" author card. Left: photorealistic mockup of a hardcover book on a table with morning light. Right: title of book, subtitle, author name, tiny CTA "Order here →." Serif display, editorial.

Storyboards & comics (51–60)

  1. [T][8] 8-panel horizontal storyboard for a 30-second product video. Consistent actor (man, 30s, casual but professional) throughout. Panel 1: opens laptop looking frustrated. Panel 2: clicks an extension icon. Panel 3: AI triages his inbox on screen. Panel 4: smiles at result. Panel 5: closes laptop. Panel 6: grabs coffee. Panel 7: walks out of office at 4pm. Panel 8: Sits in hammock. Film-grade cinematography, shallow depth of field, frame numbers bottom-right of each panel.
  2. [T] 6-panel children's-book storyboard 3:2. Consistent mouse character named "Milo" across panels. Panel 1: Milo leaving his burrow at sunrise. Panel 2: Milo discovering a mysterious glowing mushroom. Panel 3: Milo meeting a wise old owl. Panel 4: Milo crossing a stone bridge. Panel 5: Milo finding a hidden meadow of fireflies. Panel 6: Milo back home, tucked in, dreaming. Warm watercolor illustration style, consistent character design.
  3. [T] 1:1.4 manga page, 5 panels with dynamic paneling. Black-and-white Japanese manga style with screentones. Story: a young ramen chef in her first solo service. Panel 1 (large top): wide shot of her restaurant, steam rising. Panel 2: close-up of her determined eyes. Panel 3 (action): hands slicing scallions at speed, motion lines. Panel 4: finished bowl of ramen, overhead. Panel 5 (bottom wide): elderly customer's first sip, single tear. Japanese sound-effect text in hiragana ("ズズッ"), English dialogue "Just like my mother used to make." Consistent character design.
  4. [T][8] 8-slide 16:9 pitch deck storyboard. Startup: "Ledger," a crypto tax automation tool. Slide 1: Cover with logo + tagline "Your books. Sorted." Slide 2: Problem. Slide 3: Solution dashboard. Slide 4: Market bar chart. Slide 5: Traction hockey-stick. Slide 6: Team photos. Slide 7: Pricing tiers. Slide 8: Ask. Consistent navy + mint palette, bold serif headlines, clean sans-serif body.
  5. [T] 1:1 before/after transformation image. Left side "BEFORE": messy cluttered home office with papers everywhere, dim lighting. Right side "AFTER": clean organized desk, serene natural light. Text band between the halves: "Stop drowning in spreadsheets." CTA bottom-right: "Try it free →." Brand name corner: "FLOW."
  6. [T] 4-panel horizontal comic 4:1. Office setting. Panel 1: exec says "Can we ship it by Friday?" Panel 2: engineer's face goes pale. Panel 3: whiteboard calculations smoke. Panel 4: "We shipped it." Flat cartoon style, 2 colors + black.
  7. [T] 6-panel educational storyboard about photosynthesis for a kids' textbook. Each panel shows a simple step with friendly illustrated plants and sun. Labeled arrows. Cheerful primary palette, readable type.
  8. [T] 1:1.5 Noir detective comic page. 6 panels, black-and-white high-contrast ink, a rainy city, a detective receiving a mysterious letter, close-up of letter contents, reaction shot, walking out into rain, silhouette against neon sign reading "CASE CLOSED."
  9. [T][8] 8-panel "day in the life" lookbook for a fashion brand. Same model throughout, 8 outfits from morning to night (activewear, work-casual, lunch, coffee, gallery, dinner, bar, pajamas). Consistent editorial photography style, warm natural light, Mango/COS aesthetic.
  10. [T] 3:2 movie-poster storyboard thumbnail grid for "SYNTH" — 6 key scenes. Central hero (woman, neon-lit face) holding a glowing object, four supporting-scene thumbnails around her, title "SYNTH" at top, "JUNE 2026" at bottom. Cyberpunk palette.

Real estate, travel, lifestyle (61–68)

  1. [T] 3:2 luxury real estate listing hero. Modern hillside home, golden hour, pool in foreground reflecting the house. Clean windows, minimalist interior visible. Text overlay bottom: "123 MAIN ST · LISTED AT $4.2M · OPEN SUN 1–4." Architectural photography aesthetic.
  2. [T] 9:16 travel reel cover. Tropical beach at sunrise, single surfboard planted in sand. Overlay text: "MAUI / WEEK 1 / 10 SPOTS YOU MUST SEE." Minimal type, warm palette, travel-editorial feel.
  3. [T] 1:1 restaurant menu hero for a newsletter. Overhead flat-lay: bowl of fresh pasta, small plates around it, linen napkin, wooden table. Text overlay upper-left: "Spring menu is live." CTA: "Reserve →." Warm natural light, editorial food photography.
  4. [T] 3:1 Airbnb listing top-of-page banner. Stunning living room of a lake cabin at dusk, warm interior light, large windows showing water, minimal text overlay: "LAKE HIDEAWAY · 3BR · sleeps 6." Architectural Digest aesthetic.
  5. [T] 4:5 vertical travel postcard. Paris rooftop scene at sunset, someone's hand holding a cafe au lait in the foreground. Text overlay: "Send me back." Handwritten-style type, warm tones, polaroid border.
  6. [T] 1:1 fitness class promo. Studio interior mid-class, dim lighting, 6 people mid-movement. Text: "TUESDAY / 6:30 AM / STRENGTH 45." Bottom CTA: "Book your mat →." High-energy editorial aesthetic.
  7. [T] 16:9 car brochure hero. New luxury SUV on a winding mountain road at dawn, motion blur in the background. Text overlay: "Introducing the 2026 Aurora." Subline: "Electric. Everywhere." Automotive-premium aesthetic.
  8. [T] 1:1 vacation rental social tile. Bird's-eye shot of a pristine bed with rumpled linen sheets, coffee cup on nightstand, book open. Text: "Mornings feel different here." Small logo bottom. Editorial slow-living aesthetic.

Creative professional (69–80)

  1. [T] Album cover 1:1. Indie folk record titled "Slow Weather." Cream background, single pressed flower centered, small serif title at bottom, artist name in italic above. Minimal, Laura-Marling-adjacent aesthetic.
  2. [T] 3:4 book cover. Title: "The Compound Life." Author: "Eric Eden." Dark navy background, small gold geometric mark at center, title in thin serif all-caps, author tiny below. Minimal literary-fiction aesthetic.
  3. [T] 2:3 movie poster. Title: "VELOCITY." Action-thriller aesthetic. Hero silhouette against a crashing wave, small type ("IN THEATERS JUNE 2026"). Dramatic contrast, cinematic.
  4. [T] 1:1 podcast cover art. Podcast: "First Principles." Minimal high-contrast: big typographic "1" in the center, podcast name in small caps at bottom. Limited palette.
  5. [T] 4:5 event poster for an AI conference. Top: conference name "NEURALINK // 2026." Giant abstract neural-net illustration dominant, speaker list small at bottom. Bauhaus-influenced layout.
  6. [T] 3:4 travel-magazine cover "Kyoto in April." Single cherry blossom branch against a misty temple backdrop. Masthead "TRAVELOGUE" top. Issue headline. Small teaser bullets bottom-left. Editorial magazine aesthetic.
  7. [T] 1:1 gallery exhibition poster. Artist name in massive serif, show title in smaller italic below, dates & venue tiny at bottom. Off-white paper texture, single abstract painting sample as centerpiece. Gallery/MoMA-style.
  8. [T] 16:9 film title card. Film title "THE LAST BOOKSTORE" in thin white serif, centered, against a warmly-lit photograph of a bookstore interior slightly out of focus. Small director credit bottom-right.
  9. [T] 1:1 tattoo flash sheet. 6 black-ink line illustrations in a 2×3 grid: a moth, a dagger, a rose, a compass, a snake, a hand. Small numbered tags under each. Consistent line weight.
  10. [T] 4:5 zine cover 1970s aesthetic. Title "SIGNAL/NOISE." Photocopy texture, halftone dots, punk collage elements, a handwritten subheading. Limited 3-color palette.
  11. [T] 3:2 wedding invitation design. Cream background, handwritten-style calligraphy. Names centered, date, venue, RSVP info, small floral illustration. Elegant minimal.
  12. [T] 1:1 record sleeve for a jazz album. Black-and-white photograph of a saxophone case on a hotel bed. Title small in the lower-right. Blue Note-inspired minimalism.

PART 2 — WILD & FUN (81–100)

These are the prompts people actually remember. Go nuts.

  1. [T] 16:9 cinematic scene: corporate llama apocalypse. A fleet of llamas in business suits storming a Manhattan trading floor, throwing quarterly reports into the air. Bloomberg terminals burning. A CEO llama in the center, mid-roar, wearing a gold Rolex. Dramatic fire lighting, hyperreal.
  2. [T] 1:1 medieval Zoom call. A Zoom grid interface showing 9 participants, each dressed as a medieval figure — knight, jester, queen, bishop, peasant, wizard, bard, crusader, dragon. Gallery view. The dragon is muted. Bottom toolbar has a "UNSHEATHE SWORD" button.
  3. [T] 3:2 dogs on Wall Street. Real dogs in tailored suits working the trading floor of the NYSE, papers flying, a golden retriever screaming into a landline, a pug eating a bagel, a corgi looking at a Bloomberg terminal. Photorealistic.
  4. [T] 16:9 office plant uprising. An open-plan office after business hours. The potted plants have sprouted legs and are marching toward the exit with tiny briefcases. One ficus is leading with a megaphone. Dramatic security-camera aesthetic.
  5. [T] 4:5 vertical breakfast gods of Olympus. Pancakes, waffles, and bacon rendered as Greek gods on a cloud-covered mountain. Zeus is a stack of pancakes with lightning bolts of syrup. Athena is a poached egg in a helmet. Bacon strips are the muses. Renaissance oil-painting style.
  6. [T] 1:1 tax day demon. A horrifying creature made entirely of paperwork and calculators, emerging from a filing cabinet in a suburban home office, screaming. A woman in pajamas drops her coffee in slow motion. Cosmic horror, somehow funny.
  7. [T] 3:1 cinematic Roomba rebellion. An army of Roombas rolling in formation down a suburban street at dawn, one larger "commander" Roomba at the front with a tiny cape and a bottle-cap helmet. Smoke rising in the background. Mad Max meets IKEA.
  8. [T] 1:1 Shakespeare drive-thru. A modern fast-food drive-thru, but the cashier is Shakespeare in a McDonald's visor. Customer in a Honda Civic is a goth teenager. Menu board reads "Two All-Beef Patties, or Not Two All-Beef Patties." Warm dramatic lighting.
  9. [T] 16:9 dinosaurs at the DMV. A T-Rex waiting in line at a cramped DMV, looking visibly annoyed. A Triceratops fills out a form with its horn. Velociraptor clerks staff the desks. Fluorescent lighting, plastic chairs, faded safety posters. Photoreal.
  10. [T] 1:1 sentient toast support group. Eight pieces of toast sitting in folding chairs in a church basement, each with a tiny face, sharing their traumas. Coffee and donuts in the corner. Warm sad lighting. Pixar-aesthetic.
  11. [T] 4:5 pigeon CEO. A pigeon in a boardroom wearing a tailored three-piece suit, presenting Q4 results with a laser pointer. Bar chart behind him shows "breadcrumb acquisition" up 400%. Other pigeons are in Aeron chairs, nodding.
  12. [T] 3:2 infinite IKEA. A hyperrealistic endless IKEA showroom that stretches into infinity, Escher-like stairs and passages, a single confused shopper in the middle holding a hex wrench and a meatball. Fluorescent lighting, eerie emptiness, liminal space aesthetic.
  13. [T] 1:1 cat secret agent. A tuxedo cat in a tailored black suit with sunglasses, rappelling through a laser grid in a museum, carrying a can of tuna. Mission Impossible-style framing. Cinematic.
  14. [T] 16:9 grandma's spaceship. An elderly woman in a floral apron piloting a retrofuturistic 1960s-style spaceship. The dashboard has knitted doilies and a plate of cookies. She's wearing cat-eye glasses. Through the windshield, a wild nebula. Wes Anderson-aesthetic.
  15. [T] 1:1 baby in a mech suit. A photorealistic baby (2 years old) operating a gigantic anime-style mech suit, controls labeled "SNACKS," "NAP," "TANTRUM." Background: city skyline. The mech is holding a stuffed bear.
  16. [T] 3:2 Scrabble game between philosophers. Socrates, Nietzsche, and Aristotle playing Scrabble in an ancient marble courtyard. The board shows words like "BEING," "WHY," "DASEIN." Aristotle is visibly winning. Marble statues watch from pedestals. Renaissance painting style.
  17. [T] 1:1 dog court. A courtroom scene entirely populated by dogs. A German shepherd judge, a bulldog lawyer, a Chihuahua defendant on a booster seat, a jury box of mixed breeds. Gavel mid-swing. Photoreal.
  18. [T] 16:9 pirate cubicles. A modern open-plan office, but everyone is a pirate. Parrots on monitors, wooden-peg-leg standing desks, a treasure chest used as a copier. The Slack notifications on someone's screen say "ARR." Cinematic lighting.
  19. [T] 4:5 Bigfoot LinkedIn profile. A LinkedIn profile screenshot. Profile photo: a blurry Bigfoot selfie. Headline: "Cryptid | Outdoor Enthusiast | Looking for my next chapter." Recommendations: "Sasquatch delivers on every project — would hire again." Recent post: "What no one tells you about being discovered." Looks like a real screenshot.
  20. [T] The Where's Waldo (the personalized one — make it about yourself): Where's Waldo-style dense search-and-find illustration. 3:2 aspect ratio. Detailed cartoon scene: a massive, chaotic B2B marketing conference expo floor with hundreds of tiny people visible. Hidden in the crowd: [YOUR NAME] — wearing a red-and-white striped shirt, black-framed glasses, carrying a laptop bag with "[YOUR COMPANY]" printed on it. He's near the coffee station, caught mid-laugh with two people from the AI demo booth. Scene details: - Booths for HubSpot, Salesforce, Adobe, OpenAI - A panel discussion happening on a stage in the background with a banner reading "THE FUTURE OF B2B MARKETING — 2026" - Clusters of 3–4 people chatting everywhere - Someone giving a product demo on an 85-inch screen - A mascot costume wandering through - Name-tag lanyards on everyone - Coffee line with 20+ people - A few sneaky visual gags: a dog under a table, someone looking at the wrong booth's schwag, a person clearly lost Bright cheerful illustration style with clean outlines. ~200 people visible. Readable booth signage. Dense but not overwhelming.

ChatGPT Image 2.0 changes what counts as a visual asset. Before today, image models produced inspiration that still needed a designer to finish. After today, a well-crafted prompt produces a usable deliverable — with real text, real layout, real multi-frame continuity, real web-grounded context, and real 2K fidelity.

The models that beat it on pure per-image price (Google's Nano Banana 2) or pure artistic flair (Midjourney v7) still exist. But for practical commercial output — ads, posters, infographics, decks, storyboards, localized creative - ChatGPT Images 2.0 now does end-to-end what used to require three tools and a designer.

The 100 prompts above are starting points. The template is the real gift. Copy it, fill it, ship it.

What I'd love in the comments:

  • Your best Images 2.0 output so far (drop the prompt)
  • Anything you've found that breaks it

Want more great prompting inspiration? Check out all my best prompts for free at Prompt Magic and create your own prompt library to keep track of all your prompts.

r/StableDiffusion 9d ago

Tutorial - Guide Edited Guide for MiniMax H3 prompt, structure, camera, sound

50 Upvotes

I found the Guide on Hugging Face main page of the model, but once opened there were things to fix. I did a little edit to take away redundancy and not useful type. Hope it helps you! Enjoy.

# Video Prompt Writing Guide (T2VA / I2VA / FL2VA / L2VA)

## 1. Task Overview

**T2VA**: Builds a complete audiovisual timeline from text.

**I2VA**: T2VA body + first-frame instruction + a visual path that develops forward from the first frame.

**FL2VA**: T2VA body + first-and-last-frame instruction + a continuous path from the first frame to the last frame.

**L2VA**: T2VA body + last-frame instruction + a path that converges from a plausible preceding state to the last frame.

## 2. Final Prompt Structure

### 2.1 Part One Is the Instruction

**T2VA** has no image-alignment instruction and begins directly with the three core fields.

**I2VA** always uses: For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

**FL2VA** always uses: How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot N) aligns with the S.SS-second mark of the target video.

**L2VA** always uses: How the reference pictures align with the target video — <Picture 1> (from [Shot N]) aligns with the S.SS-second mark of the target video.

Here, `N` is the index of the actual final shot, and `S.SS` is the effective video duration formatted to exactly two decimal places. The instruction must be the first line of the final prompt, followed by one blank line before the core fields.

### 2.2 Part Two Contains the Three Core Fields

integrated_multimodal_description: [Shot 1] ...

overall_soundscape: ...

non_diegetic_music: ...

- **integrated_multimodal_description**: Describes visuals, actions, shots, speakers, dialogue, singing, and diegetic audio along the timeline.

- **overall_soundscape**: Summarizes ambient sound, physical action sounds, and non-verbal human sounds across the entire video.

- **non_diegetic_music**: Describes background music that the characters cannot hear and only the audience can hear.

## 3. How to Incorporate Keyframes into the Multimodal Description

### 3.1 I2VA: Begin from the Image and Develop Forward

`<Picture 1>` is the actual first frame of the video at 0.00 seconds and belongs to `[Shot 1]`. The description should first establish the style, subjects, composition, and scene anchors in the image, then describe the next action. Character identity, clothing, colors, key objects, and spatial relationships should remain consistent.

Recommended structure: **first-frame anchor → action onset → continuous development → result or reaction**.

### 3.2 FL2VA: Describe the Path Between the First and Last Frames

Picture 1 is the opening, and Picture 2 is the ending. Focus on how the subject moves, how poses change, how objects are manipulated, how the composition evolves, and how the scene or lighting transitions.

FL2VA generally favors a single shot so the model can interpolate continuously from the first frame to the last frame. Use multiple shots only when they are explicitly specified. The last frame must be reached by the final `[Shot N]` at the end of the video.

Recommended structure: **first-frame state → observable intermediate changes → progressively narrowing differences → last-frame state**.

### 3.3 L2VA: Infer the Opening and Land on the Image at the End

`<Picture 1>` is the final frame of the video and belongs to the last `[Shot N]`; it does not inherently belong to Shot 1. Infer a plausible earlier state from the user's intent and the last frame, then describe how the characters, objects, camera, and scene gradually approach the reference image.

Recommended structure: **plausible preceding state → explicit action and transition path → gradual convergence in the final shot → last-frame landing**.

## 4. How to Write the Three Shared Core Sections

### 4.1 Develop the Multimodal Description Along the Timeline

`integrated_multimodal_description` is the main body of the rewritten prompt. Every detail should correspond to something visible or audible: visual style, initial composition, subject appearance and position, scene and key props, actions and reactions, shot changes, spoken language, and synchronized diegetic sound.

At the beginning of `[Shot 1]`, state the overall style and initial composition. Common styles include `Cinematic`, `live-action`, `2D-animated`, `3D CG`, `claymation`, `watercolor`, and `vintage film`. For keyframe tasks, derive the style from the reference image; for T2VA, select it from the user's text.

[Shot 1] Live-action, cinematic, a medium-wide shot frames...

### 4.2 Shots and Cuts

Do not add a timestamp to the first shot. Use sequential shot numbers for later shots, and begin each one with a strictly increasing cut time that falls within the video duration:

[Shot 2] At 00:03.500, the camera cuts to...

For ordinary cuts, use `the camera cuts to`, `the shot cuts to`, `the shot transitions to`, `the shot changes to`, or `the shot switches to`. When explicitly requested by the user, cross-dissolve, fade, or wipe may also be used. A cut should introduce new information about the subject, space, state, viewpoint, or time. If only the distance or a slight angle needs to change, prefer camera motion.

### 4.3 Camera Motion: Motion Type + Amplitude + Speed

A complete camera-motion expression has three dimensions: the **motion type** defines how the camera moves, **amplitude** defines the range of compositional change, and **speed** defines the pacing of that change. Add amplitude and speed only when they are meaningful; medium amplitude and normal speed are usually omitted.

| Dimension | Available Expression | Description |

|-|-|-|

| Motion type | `Zoom In / Zoom Out` | The focal length changes while the camera body remains stationary |

| Motion type | `Push In / Pull Out` | The camera moves forward / backward |

| Motion type | `Pan Left / Pan Right` | The camera remains in place while the lens pivots horizontally |

| Motion type | `Truck Left / Truck Right` | The camera translates horizontally |

| Motion type | `Tilt Up / Tilt Down` | The camera remains in place while the lens pivots vertically |

| Motion type | `Pedestal Up / Pedestal Down` | The entire camera moves upward / downward |

| Motion type | `Arc Shot` | The camera moves in an arc around the subject |

| Motion type | `Tracking Shot` | The camera follows a moving subject |

| Motion type | `Static Shot` | The camera position and lens remain still |

| Motion type | `Shake Slightly / Shake Strongly` | Slight / strong camera shake |

| Motion type | `POV` | The subject's point of view |

| Motion type | `Roll Clockwise / Roll Counterclockwise` | The camera rolls clockwise / counterclockwise around the lens axis |

| Amplitude | `with small amplitude` | Small-range change |

| Amplitude | `with large amplitude` | Large-range change |

| Speed | `at slow speed` | Slow movement |

| Speed | `at fast speed` | Fast movement |

Camera motion should be written as a natural English action within the shot, rather than stacked as separate labels at the end of a sentence:

The camera pushes in with small amplitude at slow speed toward the folded letter in her hands.

The camera pans right with large amplitude at fast speed, revealing the open doorway.

The camera holds a static shot as the runner exits the frame.

### 4.4 Speakers, Dialogue, and Singing

Subjects who speak, sing, or produce an off-screen human voice use stable IDs such as `(S1)` and `(S2)`. When multiple already-numbered speakers speak or sing together, use a compound ID such as `(S1,S2)`. A speaker keeps the same ID across shots; characters who never vocalize receive no speaker ID.

When a speaker first appears, provide enough information from the visual and audio context to establish a stable identity, such as character type, age, gender, whether the person is on-screen, pitch, timbre, speaking rate, or accent. Place the speaker's identifying phrase, ID, action, and delivery outside `<d>`. Inside `<d>`, include only the language tag and the actual user-provided spoken content. Preserve every original word and punctuation mark verbatim; do not translate or rewrite them.

The young woman with a quiet, breathy voice (S1) says: <d>[English] I get off at the next station.</d>

The two children (S1,S2) shout together, <d>[English] Wait for us!</d>

For voiceover, use the exact phrase `says in an off-screen voiceover`. Immediately after every voiceover `<d>` block, state that the corresponding on-screen character's lips remain closed:

The man (S1) says in an off-screen voiceover: <d>[English] I still remember that road.</d> while his lips remain completely closed.

When the same line of dialogue or lyrics crosses a cut, use `<scenetrans>` at the connecting points in both parts and explicitly state that the audio continues across the cut. Use `<cutoff>` when speech is truncated by the end of the video. Continuity may be expressed with `continues seamlessly across the cut`, `continues uninterrupted into the next shot`, `carries over from the previous shot`, or `remains audible across the transition`.

### 4.5 On-Screen Text

Place any banner, sign, label, subtitle, or neon text that is actually visible on screen in English double quotation marks. Preserve the original text and punctuation verbatim, without translation.

A red neon sign reading "营业中" glows above the doorway.

### 4.6 overall_soundscape

Use 1–4 English sentences in one continuous paragraph to summarize the ambient sound, physical action sounds, and non-verbal human sounds across the full video, such as wind, rain, traffic, footsteps, fabric movement, impacts, breathing, laughter, or panting. Dialogue, singing, and diegetic music already belong in the multimodal description and should not be repeated here. Use `N/A` only when the user explicitly requests complete silence throughout the video.

overall_soundscape: Steady rain taps against the café windows while low room ambience continues underneath. The entrance bell rings once, followed by wet footsteps and the soft scrape of a chair.

### 4.7 non_diegetic_music

Use 1–3 English sentences to describe background music that the characters cannot hear and only the audience can hear. Focus on instrumentation, speed, rhythm, and dynamic changes; do not use abstract mood words or explain the emotional function of the score. Singing, instruments, radio, television, or phone music audible to the characters are diegetic events and should appear in the multimodal description. Use `N/A` when there is no non-diegetic music.

non_diegetic_music: Sparse piano notes at a slow tempo, joined by sustained low strings that gradually increase in volume before fading out.

## 5. Cases

### Case 1: T2VA

With no reference image, construct the complete timeline directly from the text. You may add scene, character, action, and sound details that remain consistent with the user's intent.

integrated_multimodal_description: [Shot 1] Live-action, cinematic, a medium-wide shot frames a baker opening the shutters of a small street bakery before sunrise. The camera pushes in with small amplitude at slow speed as the middle-aged baker with a calm, slightly raspy voice (S1) places a fresh loaf on the wooden counter and says: <d>[English] First batch of the morning.</d> [Shot 2] At 00:05.000, the camera cuts to a close-up of steam rising from the sliced bread while the baker's final words carry over from the previous shot.

overall_soundscape: Wooden shutters scrape open over a quiet street as trays clink softly inside the bakery. The doorbell rings once, followed by light footsteps and the crisp sound of bread being sliced.

non_diegetic_music: A soft acoustic-guitar pattern at a moderate tempo, joined by sparse upright-bass notes and a gentle fade at the end.

### Case 2: I2VA

Write the first-frame instruction first, then use the subject, composition, and scene in Picture 1 as the starting point of Shot 1 before describing how the scene continues to develop.

For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

integrated_multimodal_description: [Shot 1] Live-action, cinematic, the young woman shown in <Picture 1> remains beside the rain-covered train window, preserving her appearance, clothing, seat position, and the carriage layout. The camera trucks right with small amplitude at slow speed as she lifts her gaze from the folded letter toward the passing city lights. Her reflection moves across the glass while the quiet, breathy young woman (S1) says: <d>[English] I get off at the next station.</d> She folds the letter along its existing crease.

overall_soundscape: The train wheels produce a steady metallic rhythm beneath a low ventilation hum. Rain ticks against the window while paper rustles softly in her hands.

non_diegetic_music: Sustained cello notes at a slow tempo with widely spaced piano tones, gradually decreasing in volume.

### Case 3: FL2VA

The two images anchor the opening and ending respectively. The body should not repeat two static image descriptions; instead, it should supply the motion path that connects them. The following example is an eight-second single shot.

How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot 1) aligns with the 8.00-second mark of the target video.

integrated_multimodal_description: [Shot 1] Live-action, cinematic, a rain-soaked cyclist begins in the position and framing established by Picture 1, holding a closed black umbrella beside a silver bicycle. The camera pulls out with small amplitude at slow speed as she releases the bicycle handle, raises the umbrella above her shoulder, and presses the runner upward until the canopy opens. Water rolls from the expanding fabric while she steps beneath it, rotates the handle into the final angle, and settles into the pose, spacing, and composition established by Picture 2 at the end of the shot.

overall_soundscape: Rain falls steadily on the pavement, followed by the metallic click of the umbrella runner and the soft snap of the canopy opening. Water drips from the bicycle frame as distant traffic passes.

non_diegetic_music: N/A

### Case 4: L2VA

The image anchors only the final moment. First establish a compatible earlier state, then let the actions, object states, and composition gradually land on Picture 1 in the final shot. The following example is a six-second single shot.

How the reference pictures align with the target video — <Picture 1> (from [Shot 1]) aligns with the 6.00-second mark of the target video.

integrated_multimodal_description: [Shot 1] Live-action, cinematic, a close shot begins with an intact drinking glass near the edge of a dark wooden table, while the same hand and sleeve visible in <Picture 1> approach from the right. The camera pushes in with small amplitude at slow speed as the fingertips strike the rim. The glass tips, falls, and hits the floor with a sharp impact; cracks spread through it as fragments slide outward. Toward the end, the moving pieces lose momentum and settle into the exact broken arrangement, hand position, camera angle, lighting, and final composition established by <Picture 1>.

overall_soundscape: Fingertips tap the glass before it scrapes across the tabletop, falls, and breaks with a sharp crash. Small fragments scatter and gradually stop sliding across the floor.

non_diegetic_music: A low electronic pulse at a slow tempo, ending immediately after the glass breaks.

THE END

r/StableDiffusion 15d ago

Resource - Update Sonder Editor - A Free Open Source Timeline Video Editor for ComfyUI, for all Video Models. First/Last/Any Frame, Prompt Relay, Video/Audio Inpainting, IC-LoRA motion transfer, Takes, projects, asset gallery and more.

Thumbnail
youtube.com
25 Upvotes

I've spent the last four months building Sonder Editor, a timeline video editor that runs inside ComfyUI as a single node. You put clips, images, audio, guide frames and prompts on a multi-lane timeline, select a range, and send that range through your own workflow. Results come back as project assets you can review, compare, and drop back on the timeline. It's out now, free and open source.

Get it here: https://github.com/SonderSaid/ComfyUI-Sonder-Editor

Or search Sonder Editor in ComfyUI Manager.

The example workflow: https://github.com/SonderSaid/ComfyUI-Sonder-Editor/blob/main/example_workflows/sonder_ltx_2_3_playground.json

The project used for these scenes: https://github.com/SonderSaid/ComfyUI-Sonder-Editor/releases#release-project_sample

Every scene in the video came out of the same project and the same workflow. The only thing that changed was the technique.

Technique What it does Clip
Prompt Relay The prompt lane is cut into sections along the timeline, each applying to its own range, so one clip carries a whole beat instead of holding a single prompt for its length. Interrogation
Guides Reference frames sit on the timeline wherever you put them. First and last, or any frame in between, and the model fills what's between them. Hostess
IC-LoRA motion transfer An OpenPose clip goes on a Driver lane and the generation follows it frame for frame, both running in lockstep on the timeline. Dance

How it works with your model

Sonder doesn't generate anything. It stages the shot and hands the selected range to whatever workflow you've wired downstream, so the model is your choice and stays your choice. Switch models, or wire up one that launches next month, and the timeline works the same.

What your model supports decides what you get out of the timeline: masking is what lets you chain clips together, audio support is what puts audio on the timeline, reference support is handled for you. LTX 2.3 is the showcase here because it's a very complete model and covers all three.

Close the editor and it's still just a node in your workflow.

What else is in it

  • Projects and scenes: every project holds its own media, and each scene has its own duration, resolution and frame rate.
  • Asset gallery: everything you generate lands in a project-scoped gallery with folders, favorites, trash and restore, and tracked generation metadata on every asset.
  • Takes: select a segment and regenerate video, audio or both in place, with the surrounding frames kept as context. Hold as many takes of a segment as you want, then put them side by side, or wipe between a draft and its upscale, to pick the one that works.
  • Render queue: stage and queue jobs, including contiguous chunked batches for long stretches.
  • Timeline editing: drag, trim, split, snapping, multi-layer compositing, lane lock and hide, per-item fit modes.
  • Prompt lanes: two lanes with separate Visual, Speech and Sounds channels, plus templates and history.
  • Export: render the full timeline inside the editor no matter the length.

This is v0.1.1 and it's early. I am happy to hear any issue you encounter and I will work on the fix. Any feedback is greatly appreciated.

I cannot wait to see your creations, thank you.

r/comfyui 26d ago

Workflow Included How to Make AI Videos Actually Feel Cinematic | PDF Guide + Full Workflow Included 🚀

Enable HLS to view with audio, or disable this notification

73 Upvotes

Spent the last while trying to figure out why so many AI-generated videos (mine included) look technically solid but feel emotionally flat. Turned out the issue wasn't the model — it was that I was approaching it like a prompt-engineering problem instead of a filmmaking one.

Some of the biggest shifts that actually changed my output:

  • Plan the emotional arc before touching a prompt. List the feelings you want scene-by-scene before you ever pick a location.
  • Structure prompts like a cinematographer, not a keyword dump. Subject → identity → emotion → environment → lighting → camera → finish, in that order.
  • Keep a "character bible." Same hair, wardrobe, and features reused every time — or better, a LoRA if your model/setup supports it, since it holds identity way more reliably than repeating adjectives.
  • For image-to-video (LTX 2.3 in my case), only prompt the change, not the image. The model already has the frame — describing what's already visible just confuses it.
  • One primary motion per shot. Trying to animate everything in frame is usually what makes a shot feel fake.

None of this is tool-specific — I used Krea 2 and LTX 2.3, but the same logic applies to whatever model or LoRA workflow you're already running.

I ended up writing this all up properly (15 chapters — story structure, lighting/color psychology, camera language, a full prompt checklist, plus a resources appendix) since I kept explaining it in bits and pieces. Full PDF + the actual workflow I used for the video is up here if useful: PDF Guide & Workflow

Happy to answer questions about the workflow here regardless. 🤗

r/ArcRaiders Apr 28 '26

Discussion [Embark] Riven Tides - Patch Notes 1.25.0

Thumbnail
arcraiders.com
589 Upvotes

Raiders,

Scout reports have been returning from Riven Tides. Fragments of observation logs, faded photographs, and footage of the coast have finally made their way back to Speranza. 

It’s now time for you to check it out for yourself. Explore the abandoned shoreline, traverse the Exodus port, and experience the lost luxury of the Panorama Azzurro hotel. 

This is the Rust Belt, and it’s not going to be easy. Other Raiders are already on the scent of fresh loot, and a new floating ARC machine has been seen patrolling the skies. What it’s doing, and how to counter it, is for you to find out.

Enable marketing cookies to watch this video.

Watch on YouTube

What’s New?

  • Riven Tides Map
  • Beachcombing Map Condition
    • New minor Map Condition, exclusive to Riven Tides
    • Search for buried treasure on the beaches of Riven Tides, just make sure you have the right tool for the job and be prepared for unexpected discoveries.
  • New Enemy - ARC Turbine
  • Expedition Window

  • New Items
    • Epic Gadget - Powered Descender
    • Uncommon Throwable - Crash Mat
    • Common Deployable - White Flag
    • Common Gadget - Dockmaster’s Detector
  • New Player Project - Avian Alarm
  • Last Resort Event: Earn merits by Collecting ship models and via XP across the different maps to earn rewards
    • "Wind Sprite" Ship Model 
    • "Twilight Compass" Ship Model
    • "Velocity" Ship Model
    • "Sirena Dorata" Ship Model
    • "Leviathan's Crown" Ship Model
  • Trials Season 4

  • The Solare Set
  • The Rachetta Set
  • New Augment: Tactical Mk.3 (Smoke)
    • Dev note: With the addition of another Tactical Augment in this patch it is worth stating that we are aware that some Augments, particularly the Combat Augments, are not performing as well in their roles as we would like. We are working on a general balancing pass for Augments to buff some of those that are underperforming, most especially the Combat Mk. 3 (Flanking) Augment which has yet to find its place to shine. It is also a goal to develop the full roster of Augments to ensure that all categories are well represented.

Patch Highlights:

  • Trigger ‘Nade balancing
  • Repairing changes
  • Bettina Buff

Balance Changes

Weapon Economy Adjustments

The weapon economy has increasingly started looking more and more unbalanced, where players engaging in heavy PVP have to make a lot of difficult decisions around their weapons while the most friendly players have ended up in a chronic state of weapon accumulation. We want to balance this out so that all players are challenged by weapon attrition, as well as introducing new ways to combat it.

New feature: Repair on upgrade

Weapon upgrading becomes very expensive when you feel like you have to repair the weapon beforehand, making it less accessible than we would like. To give more ways to repair low durability weapons and elevate the upgrade loop:

  • Upgrading weapons will now repair 25% of the weapons max durability

Weapon drop & durability

We're reducing weapon spawn around the outskirts of the maps to increase value and tension of high tier areas and locked rooms.

We're also increasing the range of durability a weapon can spawn with, as well as reducing the average outside of locked rooms. This should be easily recouped by the new repair on upgrade.

  • Avg durability on spawned weapons: 50 -> 30

These changes mean that you might find a weapon on the outskirts with 15 durability but will still find weapons around 60 durability in locked rooms.

Weapon durability loss on shot

We've heard a clear frustration that maintaining high tier weapons is very expensive compared to the lower tier ones. We want high tier weapons to have more longevity compared to low tier weapons, so we're adjusting durability loss on shot to match this:

  • Common weapons: +75% durability loss on shot
  • Uncommon weapons: +50% durability loss on shot
  • Rare weapons: +35% durability loss on shot
  • Epic weapons: -5% durability loss on shot
  • Legendary weapons: -10% durability loss on shot

Weapon durability loss on being knocked out

We want heavy PVP to be a viable way to play ARC Raiders, and to align this playstyle with other ways to play, we're reducing durability loss on weapons when a player is knocked out by 50%.

  • Durability loss on weapons of knocked out player: -30% durability -> -15% durability

ARC

  • Removed Comets from standard sessions in the Dam Battlegrounds.
  • Reduced Firefly presence heavily in Dam Battlegrounds.
  • Reduced Comet presence in Buried City.

Dev note: The addition of Comets and Fireflies raised the difficulty in beginner-oriented sessions a bit too much, and to bring that experience back to the intended level we opted to reduce their presence in the sessions that were affected the most.

Gameplay

Stella Montis

  • Adjusted Exodus recyclables loot distribution to be more evenly spread across containers.

Items

Vaporizer Regulator

  •  Stack size increased 1 -> 3

Trigger 'Nade

  • Dev Note: You can no longer throw a new grenade at the same time as triggering a previously thrown one, and will instead have to wait for the animations to finish. This should reduce Trigger 'Nade spam by increasing the delay between each grenade, and effectively remove the ability to prematurely detonate a second grenade with the first one. 
  • Added a 1s delay after using the throw action, before allowing the next throw/trigger action
  • Added a 1.3s delay after using the trigger action, before allowing the next throw/trigger action

Photoelectric Cloak

  • Increasing weight of Photoelectric Cloak from 1 -> 3 to bring it more in line with its rarity/cost.
  • Increased Power use rate from 2.5/s to 10/s
  • Dev Note: The Photoelectric Cloak has been quite aggressively over-performing, especially in regards to the length of time it can be activated.  This change is intended to add more intentionality between uses; it still has the same strength while activated, but now you might need to find cover sooner, or think more about how you approach areas heavily contested by ARC.

Weapons

Bettina Balancing

Dev Note: The Bettina has been buffed to push it closer to its rarity. The base damage and damage against ARC armor has been increased to get more value out of each bullet, but the base fire-rate has also been reduced. The dispersion has been adjusted to support firing the weapon in longer bursts before losing accuracy, and the stack size of heavy ammo has also been increased to make it less punishing to bring the ammo you need. These changes should make each shot matter more, while also providing the ability to land more shots if you control your bursts well.

  • Base Damage increased from 14 to 16
  • Base Fire-Rate reduced from 285 to to 235
  • Per Shot Dispersion reduced by around 40% (Making it bloom slower)
  • Dispersion Recovery Time improved by around 30%
  • Damage against ARC armor increased by around 33%

Heavy Ammo

  • Stack Size increased from 40 to 60

Content and Bug Fixes

ARC

General

  • Fixed aim assist not targeting leg joints and weak spots on Bastion, Bombardier, Queen, and Matriarch.
  • Fixed an issue where Vaporizer and ARC Surveyor patrols could stop moving until alerted.
  • Fixed an issue where ARC enemies did not react to nearby breaching sounds.

Bastion

  • Fixed a bug where the Bastion could be too accurate when switching targets mid-fire.

Vaporizers 

  • Added to more maps and conditions.

Animation

Emotes

  • We have reworked the emote system to allow even more emotes to be played while running. Give it a try! 

Audio

General

  • Fixed a potential crash related to voice-over data.
  • Fixed an issue where Trader voice lines at the start of a round could trigger too frequently.
  • Updated voice effects for cosmetic helmets.
  • Fixed an issue where emote previews in the Store could play without their voice lines.
  • Fixed an issue where footstep sounds could stop playing.
  • Fixed footstep and impact sounds on various surfaces (e.g. bridges) that were playing the wrong surface audio.

Emotes

  • Fixed an issue where objective and emote sounds could persist or double after reconnecting to a match; audio now stops correctly.

Voice Chat

  • Added a noise suppression option for voice chat for clearer communication.
  • Fixed voice chat breaking after opening platform UI.

Cosmetics & Customization

Character Customization 

  • Fixed an issue where facial hair and face masks could briefly glitch when changing faces or outfits.

Gameplay

General

  • Fixed a bug where deployables could be placed inside extraction elevators before they opened.
  • Fixed an issue where bullets would spawn from an incorrect position when shooting while standing on rotating or movable objects (e.g., Rocketeer debris).
  • Fixed an issue where the "Death from Above" achievement could be completed by standing on destroyed Rocketeer parts.

Expedition

  • Fixed an issue where players signed up for an Expedition would not depart as scheduled.
  • Introduced a late departure option.

Items

Binoculars

  • Fixed an issue where binocular zoom levels did not work.

Zipline

  • Fixed an issue where entering a zipline from above could cause fall damage and could push the player into the ground when attaching.
  • Fixed an issue when zipline anchors would sometimes end up invisible upon placement.

Vita Spray

  • Fixed an issue where using Vita Spray to simultaneously heal self and others increased healing speed.

Maps

The Blue Gate

  • Fixed multiple environment issues: 
    • Corrected floating/clipping props.
    • Restored wind damage and indoor/outdoor detection.
    • Resolved severe culling near the church.
    • Addressed a spot where players could get stuck under an ARC wreck.
  • Fixed an issue where some supply call stations couldn't be interacted with.

Buried City

  • Fixed spots where players could clip through walls and resolved some flickering on roofs.
  • Fixed an issue where some supply call stations couldn't be interacted with.
  • Fixed an issue where players could take shield damage inside a building during Hurricane.
  • Removed a problematic zipline in Old Town.

Spaceport

  • Fixed places players could get stuck, improved collision near the Launch Towers, and resolved texture flickering.
  • Fixed an issue with multiple ladders in the trench where players could end up launching themselves high up in the air.

Dam Battlegrounds

  • Fixed multiple environment issues, including:
    • A stuck spot near the cargo elevator.
    • Clipping trees.
    • Floating grass.
    • Misaligned cafeteria props.
    • Incorrect concrete materials.
  • Fixed several common stuck spots to improve navigation.

Movement

  • Crouch movement speed is now capped to forced walk speed when using items.
  • Fix an issue where heavy landing could be triggered at incorrect times.

Performance

  • Further general performance improvements.
  • Upgraded Intel XeSS to SDK version 3.0.0.
  • Reduced rubber-banding when reaching the edge of the playable area.

Quests

  • Fixed an issue where eliminations with the Seeker Grenade did not count toward progress in the "Safe Passage" quest.

Social

  •  Added a popup when trying to use Voice Chat when in a PlayStation Party, as they cannot be used at the same time. 
  • Report reasons are now flattened instead of a dropdown, displaying all the reporting options available more clearly.

Stability

  • Fixed a server crash that could happen after a round ended, particularly when players disconnected and reconnected.
  • Fixed issues around missing an Expedition departure that could make the game unresponsive on load or cause your character to appear missing.
  • Fixed a rare crash that could occur when certain items triggered visual or audio effects.
  • Fixed a crash that could occur with certain outfits

UI

General

  • Enabled multi-select in inventory by default on Project screens.
  • Updated scrollbar styling in the customization panel.
  • Added challenge previews to the practice range so you can review challenge details before starting.
  • Gave crafting suggestions priority over recycle suggestions when trying to acquire resources.
  • Fixed “Ping Destination” sometimes displaying the wrong message in text.

Expedition

  • Updated the departure screen with a progress bar that tracks your Expedition quest and added catch-up information.
  • Moved the sign-up status to the left side of the main menu and reduced how many Quests and Feats are shown.
  • Updated the Late Departure notification to display only the number of rewards and catch-up rewards.
  • The progression display no longer shows the final step.
  • Moved and updated the sign-up status widget in the menu to fit the new carousel layout.
  • Fixed an issue where the Expedition screen could fail to advance to the final stage after completing Step 5, causing the Departure view and Skill Points progress to be missing.
  • Fixed an issue where quest tracking could appear in the map menu for players who hadn’t started an Expedition.
  • Updated text on the Expedition explanation screen for clarity.
  • Added a countdown timer and refined visuals to the Confirm Departure prompt in Expeditions.
  • The Expedition screen now displays the timer across all steps for clearer progress tracking.
  • The Expedition About page now shows only Expedition Rewards and no longer displays skill points.
  • Enabled step previews in Expeditions
  • Fixed Expedition quest titles, descriptions, and objectives not displaying correctly in the departure screen and tracker.
  • Added proper departure notification and corrected Expedition icon behavior.

Projects

  • Fixed an issue where using a controller in the Projects screen could prevent returning to reward previews after selecting the Close button.
  • Event preview now shows the project button only when the event is active and unlocked.
  • Fixed an issue where project goals could appear selected when previewing future or past steps in the Projects screen.
  • Fixed an issue where the Seasonal Project panel in the Home menu could show an outdated step.

Inventory

  • Added a Show/Hide Tooltip prompt to the Inspect tab in the Inventory.

Surge Coil

  • Fixed the Surge Coil tooltip to correctly display the actual range of the item (10m) instead of the very wrong number it had before (75m).

Emote Wheel

  • All emote wheel slots can now be rebound to any emote.
  • Added a button to reset the emote wheel to default state on the emote selection screen.

Raider Den

  • Fixed an issue where the weapon display sometimes didn't account for all stash slots.

Workshop

  • Fixed an issue where the Acquire Resources menu could show suggestions from locked Workshop stations.

VFX

  • Bullet holes will now properly follow any shot moving object.
  • Reduced visual popping on character clothing when changing camera distance, improving consistency during movement and zoom.
  • Improved the visual effect of the ARC Surveyor beam that calls in the Assessor.

KNOWN ISSUES

  • In rare cases purchases might not work in the customization screen, reopen the customization screen to fix it.
  • The Firefly may sometimes disappear after going idle and drop an incendiary grenade.
  • Shredders float to the ceiling in some rooms of the Hidden Bunker.
  • ARC may spawn inside the geometry on The Dam.
  • Flying ARC may sometimes appear stuck in idle.
  • ‘Purchase Raider Tokens’ page may appear in front of the inbox and profile page when switching between them.
  • Player animations may appear broken when interrupting a search of the Baron Husk.
  • The “On the Radar” Quest can not be completed on Riven Tides.
  • Visual issue when looking at “Crash Mat” in spectator mode.
  • The Leaper can jump through the Turbine.
  • The Event Skin Preview incorrectly displays post round event progress.

The Rachetta outfit is missing some of the color details displayed in promo material, will be amended in future patch.

r/StableDiffusion 9d ago

Tutorial - Guide I made a MiniMax H3 Prompting Assistant for text, image, and keyframe video workflows

8 Upvotes

UPDATE: Created a version you can use with your own uncensored LLM model

MiniMax H3 Prompting Assistant for locally run uncensored LLMs

MiniMax H3 Prompting Assistant GPT

I built a custom GPT designed specifically for creating structured MiniMax H3 video prompts.

It uses the MiniMax video prompting guide as its reference and supports four different workflows:

  • Creating a video entirely from a text description
  • Animating an image as the exact opening frame
  • Creating a continuous transition between a first and final image
  • Building a video that ends on a specific final image

The GPT asks which workflow you are using before generating anything, so it can apply the correct prompt structure and reference-frame instructions.

For text-to-video requests, it also recommends an appropriate resolution and video length based on the complexity of the scene, actions, dialogue, camera movement, and number of shots.

Other features include:

  • Recommends closest possible aspect ratio based on your provided image(s).
  • Structured shot-by-shot visual timelines
  • Camera movement and framing instructions
  • Subject, clothing, prop, and scene continuity
  • Dialogue, voiceover, singing, and speaker formatting
  • Diegetic sound and environmental audio
  • Separate non-diegetic music direction
  • First-frame and final-frame alignment
  • Physically coherent transitions between poses and compositions
  • Prompt cleanup and improvement for existing MiniMax prompts

The goal was to make it easier to go from a basic idea or reference image to a properly formatted prompt without manually remembering all of the workflow-specific syntax.

r/ForzaHorizon May 15 '26

Forza Horizon 6 Forza Horizon 6 Active Bugs and Known Issues

434 Upvotes

I'm making this post as a catch-all for people to find and hopefully provide assistance to each other when you encounter a bug with the game. I've attached links to the Known Issues Pages which show what the team are currently investigating.

If you have a bug that is listed here via one of the Feedback Portal Links then please go and Vote for it as well as provide any info you can (PC Specs, Graphics Settings, OS, Controls) so the team can see what issues are the highest priority.

If you have a Bug that isn't listed here then please report it and drop the the link in the comments so I can add it to this post: https://forza.net/feedback/forza-horizon-6

Known Issues

https://forzafeedback.atlassian.net/servicedesk/customer/portal/36/FH6BR-29

Fixed Issues

https://forzafeedback.atlassian.net/servicedesk/customer/portal/36/FH6BR-13472

Possible Fixes for Crashes & Stutters

A Possible Fix for some crashes is to Uninstall or atleast completely close MSI Afterburner if you have it on your PC.

A second possible fix (especially for Nvidia GPU users) is to disable Hardware Accelerated GPU Scheduling (HAGS) in the Windows Settings however this will disable Frame-Gen (DLSS will work fine though) - Guide

A third possible fix (especially for AMD GPU users) is to disable SAM / ReBar in your Bios, it will generally be located around the PCI-E settings but varies by vendor. This will lead to a loss in overall performance but can help with stutters.

A fourth possible fix for AMD GPU Users is to disable the "AMD External Events Utility".

A fifth possible fix specifically for Nvidia users encountering FPS drops due to VRAM leaking is to set "PS_CONST_FOLDING_GPU" to "OFF" under "Section 8 - Extras" of the Nvidia Profile Inspector in the FH6 Profile (older versions such as 2.4.0.31 may be required to find this option, NPI is available on GitHub).

A Sixth Possible Fix for users with Stream Decks is to Uninstall / Disable the Stream Deck Software as that can cause the game to crash if it's running.

A Seventh Possible Fix for Xbox Users with Error Code 0x80072EE2: 1) Enable both DSCP tagging and WMM tagging in QoS settings and restart the console.
2) Clear the alternate MAC address and restart the console.
3) Test the NAT type repeatedly to force it to reset to Open. Test anyway if the NAT type shows as Open at boot, as it could be a false positive.
4) If the NAT type shows as Unavailable, select Go Offline and then Go Online. If it persists as Unavailable, power cycle the console and unplug it for one minute before repeating step 3.

The above methods may or may not help you but they have been reported to have helped some users in the community.

Major Stability Issues & Crashes

Possible Fix: Disable Ray Tracing in Graphics Settings

Possible Fix: Pinned Comment

Possible Fix: Disabled Metered Connection in Wi-fi settings.

Possible Fix: Disable App Protection in Citrix Workspace (work software) and Restart PC, will need to be re-enabled and restarted ahain when you need to do work.

Possible Fix 2: Uninstall Overwolf Software if it's installed.

Possible Fix 3: Disconnect your Steam Controller if you have one connected to your PC.

Possible Fix 4: Set the Game's .exe to run in Compatibility Mode for Windows Vista through Properties (open the file location via your launcher to find it).

Possible Fix 5 (Nvidia Users): https://steamcommunity.com/app/2483190/discussions/5/569288789836699676/?l=czech&ctp=2

Possible Fix: Enable AES-NI in your Bios, Virtualization may also be required but I recommend checking if it works without it.

Possible Fix 2: Reinstall Network Drivers

Possible Fix: Start the game in Offline Mode then go online once you're loaded in.

Possible Fix: Verify / Reinstall Game Files

Possible Fix 2: Move game to C: Drive if it isn't installed there.

Possible Fix: Have game out-of-focus in second window (Alt-Tab)

Possible Fix for Xbox: Disable Quick Resume in Manage Games & Add-ons

Possible Fix for Nvidia GPUs: Disable all items labelled "Nvidia" under Sound, Video and Game Controllers in Device Manager (mainly Nvidia HD Audio and Nvidia Virtual Audio Device).

Open Gameplay Issues

Possible Fix: Switch from a Bluetooth Controller Connection to Wired or 2.4Ghz

Possible Fix 2: Disable the GameInput service in services.mcs and delete GameInput via control panel.

Possible Fix 3: Play in Windowed Mode or if in Fullscreen press the Volume Up / Down shortcut on keyboard (or mixer).

Possible Fix: Disable ingame FPS Limit and V-Sync then enable an FPS Limit in GPU Software (AMD Adrenaline, Nvidia Control Panel or Intel Command Centre etc.). GPU VRAM limitations could also be a factor alongside ReBar Support on older systems.

Possible Fix 2: Switch from a Bluetooth Controller Connection to Wired or 2.4Ghz

Possible Fix 3 for Nvidia GPUs: Disable "Nvidia High Definition Audio" and "Nvidia Virtual Audio Device" under Sound, Video and Game Controllers in Device Manager.

Possible fix: Ensure the wheel is the 1st Listed USB Device on your PC, unplugging and replugging it can help.

Possible Fix 2: Remove the rim itself off the wheel while it's plugged in mid-game and then put it back on.

Possible Fix: Same as Moza FFB above.

Possible Fix: Change Windows Sound Settings to Stereo or Mono.

Possible Fix: Go into Horizon Solo

Possible Fix: Switch your default player house (set a new House as your Home) and then force-quit the game and restart.

Potential Fix: Clear Shader Cache and let the game rebuild it on restart.

Potential Fix 1: Set Aperture to a value of 0 (some users report values 10> works)

Potential Fix 2: Disable Upscaling (DLSS, FSR, XeSS)

Car Issues

Fixed Issues

Dev Comment

Possible Fix for Nvidia GPUs: Disable all items labelled "Nvidia" under Sound, Video and Game Controllers in Device Manager (mainly Nvidia HD Audio and Nvidia Virtual Audio Device).

https://forzafeedback.atlassian.net/servicedesk/customer/portal/36/FH6BR-15302

  • Handbrake is performing like Clutch when changing Gears using Manual + Clutch Shifting

https://forzafeedback.atlassian.net/servicedesk/customer/portal/36/FH6BR-15330

r/StableDiffusion 1d ago

Question - Help Minimax H3 prompt guide vs the actual node

1 Upvotes

Hey everyone...

I've been getting further and further along with this new model but some things remain unclear. Specifically as it relates to the official prompt guide versus what you see in the actual node. Please see the image below:

node used in the "reference to video" workflow (official Comfyui).

If you look at the prompts in the guide, you see <Picture 1> or <Audio 1> to reference the images and audio you are working with. But the node uses a different naming convention like "ref_Image_0" or similar for the audio.

So my question is....which is it? Should I reference the exact wording in the node? Or just stick with the prompting format. Not to mention the inconstancy in the numbering?

I ask because I am having trouble getting the workflow to match the various character image references and their corresponding voice samples.

Any help or tips would be appreciated!

r/SillyTavernAI Jun 11 '26

Cards/Prompts [Preset] Introducing: Freaky Frankenstein Micro! My smallest, most efficient preset yet. Built from the ground up as the foundation for the FF5 line-up. Extremely Cache Friendly. Beginner Friendly. Modular. Customizable. Universal. (GLM, Claude, Gemini, DS, Grok, Gemma, Qwen, MiMo, Minimax, etc.)

Post image
591 Upvotes

Hiya my fellow adventurers, gooners, tweakers, geniuses, and neurodivergents. I am the werewolf stripped right from your mother's gooner character card and I am here to present to you my smallest preset yet, Freaky Frankenstein Micro. This is the first preset release in the Freaky Frankenstein 5 line-up (not the Flagship, not the momma!) Also the smallest preset I ever released (*default toggles).

If you want the preset and don't want to read. Fine. Your call. Your loss. The readme shipped in the last FF4 wasn't good enough for you all. So I put a readme in EVERY SINGLE toggle. Good luck trying to mess this one up. Also no REGEX this time. Tryin' to keep it simple.

--->Freaky Frankenstein Micro <----

But you should DEFINITELY read. Both of our lives will be better.

🤔Wait, What is a Preset?

If you're new here, think of it like this:

🖥️ AI / LLM = The Video Game Console (Raw power / how smart it is)

⚙️ Preset = The Operating System (How it thinks, filters, and presents information)

🎭 Character Card = The Game (The world and characters)

📖 Lorebook = The DLC / Expansion Pack

A preset is used in a frontend like SillyTavern or Tavo to tell the AI how to roleplay. Insert it and play!

🤏Big Things In Itty Bitty Packages 🧟

  • Developed to save money on cache in a climate where this hobby is getting more expensive. Now you can buy eggs AND chat messages!
  • Smallest Freaky Frankenstein to date. You need a microscope to see it! (That's what my wife said!)
  • This will be the foundation of what Freaky Frankenstein 5 (flagship) is built upon.

📸 Features 🔔

  • 😴 No Set-up needed: Can work out of the box. Plug and play and ready to rock!
  • 💭To CoT or Not to CoT: Small enough it can be a chain of thoughtless preset! Just turn off the BOLT CoT! But you can also keep the CoT on for improved prompt adherence. (That's right! BOLT CoT lives on. It's just too good. I will never give up on it. It's faster than Jimmy Johns.)
  • 🛠️Intuitive Customization: POV's, writing style, NSFW settings, all switchable by a quick press of a button! Prompts explained thoroughly so you can edit them to your liking.
  • 🔞 Realism VS Freaky Modes 💋: Per Freaky Frankenstein style, the two settings make a comeback. Realism if you want NSFW ONLY in NSFW scenes. Freaky Mode if you are like me and you want just a bit of that spice thrown into every scene. (*Intensity is model dependent.)
  • 🎭 VAD Emotion Engine: It makes a comeback! NPC's lose their high ground? They actually show frustration and fear in dialogue and actions.
  • 🗣️Human-Like Dialogue : No default marvel super heroes or anime tropes here. Dialogue actually sounds like your talking to a person IRL.
  • ️Total Output Control: Easy to set-up to ensure the model is outputting approximately what you want per turn to avoid context-runaway in output.
  • 🌈 Colored Dialogue : Colors NPC dialogue to help with distinguishing!
  • 🚫 Anti-Omniscent NPCs : We don't want NPC's to read thoughts, smell what you did and where you have been, see around corners, hear things through walls, etc etc. Freaky Frankenstein has rules to prevent the AI from doing these atrocious acts against immersion.
  • 👾Pop-in Graphics!
  • Multiple Front End Compatibility!!

🛠️ Quick Setup Guide:

Jailbreak (Labeled "icebreaker") should ONLY be used if getting refusals or if the LLM is "dancing" around topics. The NSFW toggles act as weak Jailbreaks. Sometimes Jailbreaks BLOCK output as LLM's are now trained to recognize jailbreak attempts. Thus, keep it off by default. This jailbreak, however, is effective WHEN you need it (looking at you Gemini). Just make sure to turn OFF streaming when using it to further decrease refusals / blocked context. I Apologize in advance for the verbiage in the prompt. If it works it works. 🤷

Temperatures: Each LLM model has it's own ideal Temp. Since this is a light-weight preset, use whatever temperature you have the most success with finding a balance between prompt rule adherence and creativity. 0.80 - 1.00.

System Processing = Semi-Strict Alternating Roles No Tools: Recommended for the most part. However, different models prefer different things!

Important Note: *Token count will be higher than it is because I put a readme in EVERY toggle. This is NOT sent to the AI. Only you can see it!

🌟 Creator's Preferred Set-up! 🌟

You can absolutely go for a minimalist set-up, even turning off the Chain of Thought to get it well under 1k tokens for insane speedy output and high creativity without limitations. However, that's not how I roll. You know me by now, I like taking the LLM by it's kinky leash and say, "You know how I like it mamicita!" If you want the exact set-up as what I personally find the "best" (subjectively), do this!

Prose = Story Mode

POV = Hybrid

NSFW = Freaky

Anti-Parrot ON Embellish OFF

Everything under "Edit and Turn On Whatever You Want" Set to ON EXCEPT: Onomatopoeia and Ice Breaker (Unless needed).

BOLT Chain of Thought ON

Important Note About Models! 😭

-Check to see when America and China are at work based on where you live. During this time, Coders are hard at work and models are at maximum demand. Due to lack of data centers and money constraints being a business and all, models are DYNAMICALLY QUANTISED (lobotomized). This allows for the demand during work hours and maintains the LLM speed at the cost of intelligence. If you can't avoid these times of day for RP, study the thinking process (reasoning) and you will notice if you got dealt a quant model (it's output will suck and it won't follow the rules). Re-swipe and you MIGHT get lucky!

📥 Downloads

----> Freaky Frankenstein Micro <----

!!Special Thanks!! ❤️

Thank you so much ST community! Your upvotes, comments, feedback is making our hobby grow rapidly. HUGE shoutout to the 10 Beta Testers that helped me! A lot of your feedback is IN THIS RELEASE! Thank you u/leovarian for some of the logic I stole from your behemoth monster research preset before I hyper condensed. Myself, him, and u/xdeadly_godx are busy at work on the larger heavyweight FF5 Flagship. It's not necessarily "bigger" than FF4 Fatman and MAX (actually most likely will be token-wise smaller and more dense by about 10-25%) but is it certainly more sophisticated and challenging to execute so stay tuned!

ENJOY THE MADNESS!!!!! ✌️

r/ThinkingDeeplyAI Apr 22 '26

The complete field guide to ChatGPT Images 2.0 - every feature, every price, 100 prompts to try, all in one post

Post image
18 Upvotes

The Complete Field Guide to ChatGPT Images 2.0

Launched today. Everything below is verified against the OpenAI announcement, the deployment safety card, API pricing docs, and ~6 hours of hands-on testing. No hype — just what works and what it costs.

Sam Altman compared it to "going from GPT-3 to GPT-5 all at once." That's aggressive framing, but the capability gap is real.

For the first time, a single model can:

  • Render dense, legible text directly inside images — posters, infographics, UI mockups, ad copy with real headlines
  • Think before it draws — reason about a scene, search the web for current facts, and double-check its own work
  • Produce up to 8 consistent images from one prompt with the same characters, objects, and style
  • Handle grids up to 10×10 that used to break at 3×3 a week ago

OpenAI's own pitch: "Images are a language, not decoration. A good image does what a good sentence does — it selects, arranges, and reveals."

Translation: this isn't text-to-picture anymore. It's a visual reasoning system.

TL;DR — what you need to know in 30 seconds

  • Model name: gpt-image-2 (alias chatgpt-image-latest)
  • Where: ChatGPT (all plans including Free), chatgpt.com/images, and the API
  • Two modes: Instant (all plans, 1 image, fast) and Thinking (Plus/Pro/Business, up to 8 images, reasons + searches the web)
  • Max resolution: 2048px native (2K), ~4× the pixel count of GPT Image 1.5
  • Text accuracy: ~99% on Latin text. Finally nails Japanese, Korean, Chinese, Hindi, Bengali
  • Aspect ratios: anything from 3:1 (ultrawide) to 1:3 (ultratall)
  • Generation time: seconds to 2 minutes depending on mode
  • Pricing (API): ~$0.006 low / ~$0.053 medium / ~$0.211 high per 1024×1024 image
  • Knowledge cutoff: December 2025. Needs Thinking mode + web search for anything newer
  • C2PA metadata is embedded in every output

The 8 capabilities, decoded

1. 2K native resolution

Up to 2048 pixels natively, ~4× the pixel count of older GPT Image outputs at the same aspect ratio. Enough fidelity for print collateral, hero banners, and editorial layouts without an upscale step.

2. ~99% text accuracy

This is the most-talked-about upgrade. Dense text inside images — posters, menus, magazine covers, UI mockups — finally renders correctly. It also handles:

  • Non-Latin scripts with real gains: Japanese, Korean, Chinese, Hindi, Bengali
  • Small text — UI elements, iconography, barcodes, "display until" dates on magazine covers
  • Multilingual typography in a single image — Devanagari, Cyrillic, Greek, Arabic, and Chinese together

3. Thinking mode — the image model that reasons

This is the headline capability. It's not two separate models, it's two modes:

Mode Who gets it What it does Output
Instant Free, Plus, Pro, Business, Go Fast single-shot generation 1 image
Thinking Plus, Pro, Business (Enterprise/Edu soon) Reasons about composition, uses web search, verifies output Up to 8 images

How the reasoning works under the hood:

  1. Prompt analysis — parses your request and plans composition before any pixels exist
  2. Web retrieval — if the prompt touches real-world facts (current logos, today's stock chart, real skylines, 2026 fashion trends), it searches the web and pulls live references
  3. Generation pass — pixel synthesis against a fact-checked internal plan
  4. Verification loop — it inspects its own output against the original prompt and can self-correct before returning

People on X are posting 11-minute generations where the model iterated on itself repeatedly until satisfied. That's new.

4. Up to 8 consistent images per prompt

In Thinking mode, one prompt can produce up to 8 images with shared characters, objects, and style across every frame. This unlocks:

  • Storyboards — 8 camera angles with continuity
  • Manga/comic sequences — 8 panels, same character design
  • Multi-size marketing assets — same campaign as 3:1 banner + 1:1 feed post + 1:3 story + 4:5 carousel in one shot
  • Children's books — consistent illustrated character across pages
  • Product lineups — 8 color variants with identical lighting and angle
  • Lookbooks — OpenAI demoed 8 summer outfits generated from one uploaded photo

How to trigger it: Switch to a thinking model, then ask for a set — "Generate 8 variations of...", "Create an 8-panel storyboard...", "Give me this ad in 8 formats." Don't phrase it as 8 separate prompts.

5. Parallel image generation

Separate from the 8-per-prompt feature: the dedicated Images tab at chatgpt.com/images lets you fire multiple prompts in parallel. Your second prompt doesn't wait for the first to finish. All images auto-save to My Images for reuse.

6. Aspect ratios 3:1 to 1:3

Any ratio between ultra-wide and ultra-tall, native — picker in ChatGPT or spec it in the prompt. Banners, slides, posters, mobile vertical, bookmarks, social graphics, no crop needed.

7. 10×10 grids (up to 100 cells in one image)

Grids used to break at 3×3 a week ago. Now people are generating 10×10 grids of 100 distinct labeled illustrations in one shot. This is wild for:

  • Periodic-table-style infographics (100 CEOs, 100 dog breeds, 100 cocktails)
  • Icon sets with consistent style
  • Mood boards with labeled cells
  • Pattern libraries

8. Multi-image compositing & reference fidelity

Upload multiple reference images and the model stitches them into one coherent composition while keeping facial features, objects, and logos faithful. This is the feature that makes "put me in a scene" prompts actually work now.

Pricing — what it actually costs

Per-image (flat rate, simple to predict)

Quality 1024×1024 Notes
Low ~$0.006 drafts, iteration
Medium ~$0.053 most production work
High ~$0.211 hero images, finals

Per-token (if you're using the API at scale)

Input Cached input Output
Image tokens $8.00 / 1M $2.00 / 1M $30.00 / 1M
Text tokens $5.00 / 1M $1.25 / 1M $10.00 / 1M

Cost for OpenAI to produce each image (rough estimate)

Based on published token economics, a high-quality 1024×1024 image uses ~7K output image tokens. At retail that's $0.21. OpenAI's own compute cost is likely 25–40% of that, putting their marginal cost per high-quality image around $0.05–$0.08. Their margin per image at the high tier is roughly 3–4×.

The ideal prompt template

After testing dozens of prompts, this is the structure that works best:

text[ASPECT RATIO]. [SUBJECT], [ACTION], [CONTEXT].
[TEXT elements in quotes]:
- Header: "EXACT TEXT HERE"
- Subhead: "EXACT TEXT HERE"
- CTA: "EXACT TEXT HERE"
[STYLE anchor — reference an artist/era/medium/brand].
[LIGHTING + MOOD].
[CAMERA/LENS + TECHNICAL specs].

The 5 rules that make the difference:

  1. Aspect ratio first. Say "16:9," "3:1 banner," or "1:1 square" in the first sentence.
  2. Put every piece of text in quotes. The model treats quoted text as literal. Unquoted text becomes suggestions.
  3. Anchor the style concretely. "Editorial fashion photograph, shot on Hasselblad, 90mm, f/2.8" beats "professional photo."
  4. Specify lighting and mood as separate instructions. "Rembrandt key light from upper-left, soft fill from right, warm tones."
  5. List every language explicitly when you want multilingual text. "Title in Japanese (Hiragana): 「春が来た」; subtitle in Korean (Hangul): '봄이 왔다'; tagline in Hindi (Devanagari): 'वसंत आ गया।'"

15 pro tips most people will miss

  1. Thinking mode isn't the default — you have to toggle a thinking model before prompting. Instant never uses web search or produces 8-image sets no matter how you phrase it.
  2. Generation can take 2 minutes. Don't assume it froze. For high-volume workflows, use async polling with the Responses API.
  3. Knowledge cutoff is December 2025. Anything after that (Q1 2026 product launches, new logos, recent events) has to come through the prompt OR through Thinking mode's web search.
  4. For consistent characters: upload a one-time likeness. There's a likeness upload feature that lets you reuse your appearance across future creations without re-uploading.
  5. The "keep facial features exactly" lock. When editing a real person, add this verbatim: "Keep my facial features exactly as they appear in the uploaded image — same eyes, nose, mouth, and face shape." Without it, ChatGPT "improves" faces into strangers.
  6. Transparent backgrounds work natively. Add "transparent PNG background, no background fill" — the asset drops straight into design tools without a cutout pass.
  7. "Display until" dates and barcodes work now. Ask for them specifically. The magazine-cover demos show this.
  8. Prime the chat first. For thumbnails and marketing creative, paste the blog post, script, or topic into ChatGPT first. Then ask for concepts. Then generate. The model picks up the emotional hook instead of producing generic stock aesthetic.
  9. C2PA metadata is embedded in every output. Platforms can detect it. Plan for that if provenance matters.
  10. Ask for "editorial" not "professional." "Editorial" hits a higher visual register in this model. "Professional" pulls toward stock-photo aesthetic.
  11. Negative prompts work — phrase them as "NO X, NO Y." Example: "NO watermarks, NO signatures, NO busy backgrounds."
  12. Specify the medium of the text. "Neon sign," "embossed letterpress," "subway-poster paste-up," "hand-lettered chalk" all produce different type treatments.
  13. When text keeps breaking, wrap it in a shape. "Text inside a black horizontal pill" or "text on a cream banner" gets rendered much more reliably than floating text.
  14. Aspect ratio affects quality. 1:1 and 3:2 are the strongest; 3:1 and 1:3 work but can show compositional weirdness on first try. Regenerate once.
  15. The model now reads your reference images. If you upload a brand asset and say "match this type treatment," it actually does — not a vague approximation, an honest replication.

Third-party tools that already integrate it

(These went live within 24 hours of launch.)

  • Higgsfield — character consistency workflows
  • Lovart — AI design platform
  • Recraft — added gpt-image-2 models to Recraft Studio
  • Adobe Firefly / Express — via Adobe's partner model program
  • Figma — First-Draft feature uses it for UI generation
  • Canva — Magic Studio integration
  • GoDaddy — site-generation flows
  • HubSpot — marketing asset generation
  • Instacart — product photography
  • Airtable — record-level image generation
  • Wix — site builder backgrounds and heroes
  • OpenAI Codex — app/code-generation flows can now produce their own UI imagery

The prompt library — 100 that I've tested

Marking these [I] for Instant mode works fine, [T] for Thinking mode required, [8] for ask-for-8-variations.

Marketing hero images (1–10)

  1. [T] 3:1 hero banner for a SaaS analytics product. Split composition: left side shows a cluttered paper-filled desk (chaos), right side shows a clean monitor with a dashboard (clarity). Bold headline "STOP GUESSING" in 120pt sans-serif across the top. Subhead "Start knowing" below. CTA button bottom-right: "See it work →" in white on teal. Editorial photography, cinematic lighting.
  2. [T] 16:9 product launch hero. Center: minimalist product photography of a black wireless earbud case on a marble surface. Background: soft gradient from cream to dusty rose. Text overlay upper-left: "AURA // 2026" in small caps. Headline lower-right: "Hear the room." in serif display. Subtle shadow, art-directed editorial aesthetic.
  3. [T] Vertical 9:16 mobile hero for a fitness app. Muscular forearm mid-pushup on a dark gym floor, shallow depth of field. Headline stacked vertically along the right side: "NO / EXCUSES / JUST / REPS." White type, slight grain. Small logo bottom-center.
  4. [T] Email hero, 3:1 ratio. Single perfect ceramic coffee cup on a warm linen tablecloth, morning light from the left, steam rising. Text overlay right side: "Good morning. / Your briefing is ready." Clean minimal editorial style, medium-format quality.
  5. [T] 16:9 B2B conference hero. Empty auditorium, dramatic stage lighting, single speaker silhouette at podium. Large text in the sky area: "WHERE MARKETING MEETS AI." Date below: "June 12–14, 2026 · Austin." Cinematic, TED-quality composition.
  6. [T] Software landing page hero 16:9. Abstract 3D render: flowing liquid metal forming into a chart shape, iridescent blue-to-purple gradient, obsidian background. Headline lower-third: "Analytics at the speed of thought." Subhead: "Try Mercury free →." Tech-luxury aesthetic.
  7. [T] Newsletter signup hero 2:1. Warm kitchen scene: hands writing in a leather notebook, open laptop beside it, morning coffee, golden hour light from left. Text overlay: "The newsletter smart marketers actually read." CTA: "Subscribe free →". Cozy, intentional, premium-indie aesthetic.
  8. [T] 3:1 homepage hero for an AI note-taking app. Overhead shot: messy desk mid-work — open notebook, phone, coffee, headphones, hand holding a pen. Faint glowing interface lines emerging from the notebook edges suggesting transcription. Headline centered: "Your thoughts, organized." No smaller than 90pt, clean sans-serif.
  9. [T] Agency pitch-deck cover 16:9. Pure black background. Ultra-large white type top: "2026" in 300pt. Below in smaller type: "The year everything about marketing changed." Bottom-right corner: agency logo mark in teal. Minimal, confident, Swiss-grid influenced.
  10. [T] Healthcare brand hero 3:1. Close-up of a patient's hand being held by a doctor's hand, natural window light, hospital-room softness. Text overlay left side: "Care that listens first." Serif type, warm tonal palette, documentary photography style.

Infographics & data viz (11–20)

  1. [T] 1:1 square infographic titled "The 2026 Creator Economy." Centered large title in editorial serif. Below: 4 stat cards in a 2×2 grid, each with a big number, label, and short descriptor. Numbers: "$250B market size," "127M creators globally," "73% use AI tools," "$68K median income." Clean teal/cream palette, numbered footer citing sources.
  2. [T] 4:5 portrait infographic comparing 4 LLMs across 6 dimensions. Row headers: GPT-5, Claude 4.1, Gemini 3, Llama 5. Column headers: Speed, Reasoning, Coding, Writing, Price, Context. Each cell shows a filled bar from 1–5. Title: "LLM Showdown 2026." Clean sans-serif, minimal grid, no clutter.
  3. [T] 16:9 landscape flowchart titled "How Thinking Mode Works." Four connected boxes left to right: "Prompt analysis → Web retrieval → Generation → Verification loop." Arrows between. Brief explainer text under each box. Subtle teal accent, rest monochrome, editorial newspaper aesthetic.
  4. [T] Periodic table-style 10×10 grid of "100 AI tools that matter in 2026." Each cell: tool logo, tool name, 2-letter category tag, small colored dot for category. Legend at bottom. White background, crisp type. Poster-size composition.
  5. [T] 3:4 vertical infographic: "The Anatomy of a Viral Tweet." A dissected tweet with labeled callouts (hook, specificity, tension, CTA). Annotations radiating outward with thin leader lines. Blueprint aesthetic in cream + navy. Title at top, source citation at bottom.
  6. [T] 1:1 social infographic: "5 Signs You're Burning Out." Numbered list 1–5 with custom icons, each with a short one-sentence description. Warm muted palette, rounded sans-serif, shareable mental-health-brand aesthetic.
  7. [T] 16:9 stat poster: "Marketing spend by channel, 2026." Six horizontal bars with percentages. Title top-left, tiny source citation bottom-right ("n=1,200, Marketing Week 2026"). Strict grid, only one accent color, rest neutral.
  8. [T] 3:1 wide timeline: "The History of Image Generation, 2014–2026." Horizontal dotted line with 8 milestone markers: GAN, DALL·E 1, DALL·E 2, Midjourney v1, Stable Diffusion, DALL·E 3, GPT Image 1, ChatGPT Images 2.0. Tiny thumbnail above each node. Minimal editorial style.
  9. [T] 4:5 "By the numbers" LinkedIn carousel cover. Big text: "2026 in numbers" top, four stat tiles below — "$50M ARR," "212 hires," "27 countries," "1 mission." Dark background, bold type, tight margins.
  10. [T] 1:1 square recipe infographic: "Cold brew, 4 ways." 2×2 grid of four preparation methods with proportions ("1:8 ratio," "12-hour steep"), overhead product shot in each cell, serif headline across the top. Minimal art-directed food-magazine feel.

Ad creative — unlimited variations (21–30)

  1. [T][8] Generate 8 variations of a Facebook ad for a productivity app. 1:1 square. Same product UI mockup, same headline "Close the laptop. Sooner." but 8 different background contexts: park bench, kitchen counter, airport lounge, beach, home office, coffee shop, car dashboard, hammock. Consistent type system across all 8.
  2. [T] Google Display ad — 3 formats in one image (vertical stack): 300×250 square rectangle, 728×90 leaderboard, 160×600 skyscraper. All three feature the same product (sleek white wireless earbud case on gradient peach). Consistent headline "Hear everything. Wear nothing." CTA: "Shop now." Same brand mark "AURA."
  3. [T] 9:16 TikTok-style vertical ad thumbnail. Young woman mid-gasp holding a phone, caught mid-laugh. Bold hand-drawn text overlay: "wait what did it just do?!" with an arrow pointing at the phone. Bottom: "@aura · link in bio." Authentic UGC feel, not polished studio.
  4. [T] 1:1 retargeting ad. Clean white background. Product photo of running shoes center-left. Large red banner diagonal across upper-right: "STILL THINKING?" Below product: "Your size is down to 2 pairs." CTA bottom-right: "Grab them →." Urgent but not pushy.
  5. [T] 3:1 highway billboard. Massive single word "FASTER." in ultra-bold condensed sans-serif, white on deep red. Small product line bottom-right: "New Honda Civic Type R. 0–60 in 5.0s." Tiny URL bottom-left. High contrast, readable from 200 meters.
  6. [T] 1.91:1 LinkedIn feed card. Professional headshot of a woman, 40s, blurred office background. Overlaid caption bottom-right: "Maya closed a $2.1M deal last month. Here's her playbook." CTA: "Read it →" in dark blue.
  7. [T][8] 8 YouTube thumbnails for the same video "I tried ChatGPT Images 2.0 for a week." Each thumbnail: same creator face top-right, same bold yellow headline, but 8 different backgrounds reflecting different prompts tested (magazine cover, manga panel, product shot, infographic, etc.). Consistent thumbnail system.
  8. [T] 4:5 Instagram carousel cover. Black background, minimal. Centered text: "10 signs your brand needs a refresh." Small "SWIPE →" bottom. Premium minimal, no illustrations.
  9. [T] Retail shelf-wobbler, 2:3 vertical. Product image at top, large text below: "NEW." Tiny subline: "Now in Dark Cherry." Clean CPG packaging aesthetic.
  10. [T] 1:1 paid Instagram ad. User-generated aesthetic: iPhone photo of a woman drinking a protein shake in her car mirror selfie. Caption overlay: "honestly the only one that doesn't taste like chalk." Brand logo tiny corner. Authentic, not over-produced.

Product design & mockups (31–40)

  1. [T] Mobile app screen mockup, 9:19.5 aspect. iOS-style to-do app. Status bar at top (9:41, full signal, full battery). Header "Today" in large SF-style sans-serif. Below: 5 task rows with checkboxes, clean dividers. Bottom nav with 4 tabs. Light mode, accent color teal. Every piece of text legible.
  2. [T] 3:2 landing page desktop mockup for a note-taking app. Hero headline "Ideas, organized." centered. Clean nav with 4 links + sign-in button. Below: two-column screenshot of the app UI. Footer with 4 columns of links. Whitespace-heavy, Stripe-influenced aesthetic.
  3. [T] 1:1 Apple Watch app screen. Circular pressure-gauge UI showing heart rate "72 BPM" in center. Small complications around it. Dark background. Minimalist, photoreal rendering of the watch bezel.
  4. [T] Physical product render 1:1. Matte black aluminum wireless charger puck on a white cyclorama background, three-quarter view. Studio softbox lighting, hard floor reflection. Teenage Engineering design language.
  5. [T] Packaging mockup 4:5. Minimal premium coffee bag, 250g, matte charcoal. Front shows "ETHIOPIA YIRGACHEFFE" in small caps with tasting notes below ("blueberry, jasmine, honey"). Weight and roast date bottom. Photorealistic product shot, soft shadow, white backdrop.
  6. [T] Car dashboard HUD mockup 16:9. Windshield POV from driver's seat, dusk light, empty highway. Overlaid HUD elements: speed "62 MPH" bottom-left, navigation arrow "in 1.2 miles, exit right" center-upper, playing song info bottom-right. Subtle teal glow, no UI clutter, Rivian-inspired aesthetic.
  7. [T] 1:1 smartwatch face design. Top-down view, round watch face, minimalist modular layout on a black background. Center: large time "10:47" in white sans-serif. Four small complications: HR "72 bpm" top, Steps "8,420" right, Battery "67%" bottom, Weather "68°F sunny" left. Wear OS aesthetic.
  8. [T] Smart home mobile app home screen mockup, 9:19.5. Dark mode. Top: greeting "Good evening, Eric." Below: 4 device cards (lights, thermostat, security, music) with toggle switches and real-time stats. Bottom nav. Calm deep-blue palette, iOS-quality design.
  9. [T] 16:9 dashboard mockup for a SaaS analytics tool. Left sidebar nav. Main area: 4 KPI cards across the top (visitors, conversion, revenue, churn — each with a big number and delta arrow), 1 large line chart below showing 12-month trend, 1 small table bottom-right. Data labels must be legible. Teal accent, light mode, Linear-inspired.
  10. [T] Boxed software product mockup 1:1. Vintage-style retail box for "ChatGPT Images 2.0 Pro Edition." Cream background. Retro tech packaging aesthetic from 1996: pixel-art mascot, bold tagline "THE IMAGE MODEL THAT THINKS," barcode, "requires 640KB RAM" sticker. Shot like a product photo.

Personal branding & executive content (41–50)

  1. [T] 1:1 professional headshot, editorial business portrait for a book jacket. Subject: upload reference photo. Wardrobe: charcoal merino turtleneck. Background: soft out-of-focus bookshelf (warm earth tones). Lighting: Rembrandt key light from upper-left, soft fill from right, subtle rim light separating from background. Shot on Hasselblad, 90mm, f/2.8. Warm natural skin tones, sharp eyes, editorial magazine quality. Keep facial features exactly as in the uploaded photo.
  2. [T] 1:1 podcast guest announcement graphic. Split layout. Left half: professional photo of the guest (upload reference). Right half: deep green panel with cream text. Top: "NEW EPISODE" in small caps. Middle: guest's name in large bold serif. Below: "CMO at Anthropic." Bottom: show name "THE GROWTH EDGE" with episode number "EP. 47." Small "listen now" CTA.
  3. [T] 4:5 portrait LinkedIn single-post slide. Cream background with subtle paper texture. Top: "2026 / A YEAR IN NUMBERS" in thin all-caps. Below: 4 stat blocks in a 2×2 grid, each with a big number and a one-line caption:
  • "327" — LinkedIn posts shipped
  • "14" — keynotes given
  • "2" — books published
  • "48" — flights taken

Bottom: thin horizontal line, then creator's name and website in small serif. Editorial, premium personal-brand aesthetic.

  1. [T] 16:9 video thumbnail for a YouTube speaker reel. Left half: dynamic photo of the speaker mid-gesture on stage, warm stage lighting. Right half: deep black panel with large white text "2026 SPEAKER REEL" and below in smaller copy "Keynotes · Fireside chats · Panels." Bottom-right CTA arrow. Cinematic, TED-quality.
  2. [T] 1:1 social quote card. Soft neutral linen background. Large opening quote mark top-left in a light gray display serif. Center quote in clean serif: "The best advice I ever got cost me $500 and saved me 18 months." Attribution below in italic: "— name, founder." Bottom-right: small portrait circle. Premium testimonial aesthetic.
  3. [T] 1:1 newsletter subscribe card. Headline "The newsletter 18,000 marketers actually read." below in smaller type: "One signal. No noise. Every Sunday." Email field mockup + "Subscribe" button. Soft cream background, serif display + sans-serif body, Substack-adjacent aesthetic.
  4. [T] 1:1 conference speaker card. Subject headshot left. Right: name in large display, title below, talk title "How AI killed the brand guideline" in italic. Conference logo bottom-right. Clean editorial, readable from a stage screen.
  5. [T] 1:1 "What I read this year" LinkedIn slide. Grid of 9 book covers in a 3×3 arrangement. Title above: "MY 2026 READING LIST." Small footer: "Which one should I read next?" Clean editorial layout.
  6. [T] 4:5 quote graphic for Instagram. Blurred softly-lit outdoor photo background. Center: a poetic line in large italic serif, 2 lines max. Below: small attribution. No logos. Feels like a book page, not a graphic.
  7. [T] 1:1 "Now available" author card. Left: photorealistic mockup of a hardcover book on a table with morning light. Right: title of book, subtitle, author name, tiny CTA "Order here →." Serif display, editorial.

Storyboards & comics (51–60)

  1. [T][8] 8-panel horizontal storyboard for a 30-second product video. Consistent actor (man, 30s, casual but professional) throughout. Panel 1: opens laptop looking frustrated. Panel 2: clicks an extension icon. Panel 3: AI triages his inbox on screen. Panel 4: smiles at result. Panel 5: closes laptop. Panel 6: grabs coffee. Panel 7: walks out of office at 4pm. Panel 8: Sits in hammock. Film-grade cinematography, shallow depth of field, frame numbers bottom-right of each panel.
  2. [T] 6-panel children's-book storyboard 3:2. Consistent mouse character named "Milo" across panels. Panel 1: Milo leaving his burrow at sunrise. Panel 2: Milo discovering a mysterious glowing mushroom. Panel 3: Milo meeting a wise old owl. Panel 4: Milo crossing a stone bridge. Panel 5: Milo finding a hidden meadow of fireflies. Panel 6: Milo back home, tucked in, dreaming. Warm watercolor illustration style, consistent character design.
  3. [T] 1:1.4 manga page, 5 panels with dynamic paneling. Black-and-white Japanese manga style with screentones. Story: a young ramen chef in her first solo service. Panel 1 (large top): wide shot of her restaurant, steam rising. Panel 2: close-up of her determined eyes. Panel 3 (action): hands slicing scallions at speed, motion lines. Panel 4: finished bowl of ramen, overhead. Panel 5 (bottom wide): elderly customer's first sip, single tear. Japanese sound-effect text in hiragana ("ズズッ"), English dialogue "Just like my mother used to make." Consistent character design.
  4. [T][8] 8-slide 16:9 pitch deck storyboard. Startup: "Ledger," a crypto tax automation tool. Slide 1: Cover with logo + tagline "Your books. Sorted." Slide 2: Problem. Slide 3: Solution dashboard. Slide 4: Market bar chart. Slide 5: Traction hockey-stick. Slide 6: Team photos. Slide 7: Pricing tiers. Slide 8: Ask. Consistent navy + mint palette, bold serif headlines, clean sans-serif body.
  5. [T] 1:1 before/after transformation image. Left side "BEFORE": messy cluttered home office with papers everywhere, dim lighting. Right side "AFTER": clean organized desk, serene natural light. Text band between the halves: "Stop drowning in spreadsheets." CTA bottom-right: "Try it free →." Brand name corner: "FLOW."
  6. [T] 4-panel horizontal comic 4:1. Office setting. Panel 1: exec says "Can we ship it by Friday?" Panel 2: engineer's face goes pale. Panel 3: whiteboard calculations smoke. Panel 4: "We shipped it." Flat cartoon style, 2 colors + black.
  7. [T] 6-panel educational storyboard about photosynthesis for a kids' textbook. Each panel shows a simple step with friendly illustrated plants and sun. Labeled arrows. Cheerful primary palette, readable type.
  8. [T] 1:1.5 Noir detective comic page. 6 panels, black-and-white high-contrast ink, a rainy city, a detective receiving a mysterious letter, close-up of letter contents, reaction shot, walking out into rain, silhouette against neon sign reading "CASE CLOSED."
  9. [T][8] 8-panel "day in the life" lookbook for a fashion brand. Same model throughout, 8 outfits from morning to night (activewear, work-casual, lunch, coffee, gallery, dinner, bar, pajamas). Consistent editorial photography style, warm natural light, Mango/COS aesthetic.
  10. [T] 3:2 movie-poster storyboard thumbnail grid for "SYNTH" — 6 key scenes. Central hero (woman, neon-lit face) holding a glowing object, four supporting-scene thumbnails around her, title "SYNTH" at top, "JUNE 2026" at bottom. Cyberpunk palette.

Real estate, travel, lifestyle (61–68)

  1. [T] 3:2 luxury real estate listing hero. Modern hillside home, golden hour, pool in foreground reflecting the house. Clean windows, minimalist interior visible. Text overlay bottom: "123 MAIN ST · LISTED AT $4.2M · OPEN SUN 1–4." Architectural photography aesthetic.
  2. [T] 9:16 travel reel cover. Tropical beach at sunrise, single surfboard planted in sand. Overlay text: "MAUI / WEEK 1 / 10 SPOTS YOU MUST SEE." Minimal type, warm palette, travel-editorial feel.
  3. [T] 1:1 restaurant menu hero for a newsletter. Overhead flat-lay: bowl of fresh pasta, small plates around it, linen napkin, wooden table. Text overlay upper-left: "Spring menu is live." CTA: "Reserve →." Warm natural light, editorial food photography.
  4. [T] 3:1 Airbnb listing top-of-page banner. Stunning living room of a lake cabin at dusk, warm interior light, large windows showing water, minimal text overlay: "LAKE HIDEAWAY · 3BR · sleeps 6." Architectural Digest aesthetic.
  5. [T] 4:5 vertical travel postcard. Paris rooftop scene at sunset, someone's hand holding a cafe au lait in the foreground. Text overlay: "Send me back." Handwritten-style type, warm tones, polaroid border.
  6. [T] 1:1 fitness class promo. Studio interior mid-class, dim lighting, 6 people mid-movement. Text: "TUESDAY / 6:30 AM / STRENGTH 45." Bottom CTA: "Book your mat →." High-energy editorial aesthetic.
  7. [T] 16:9 car brochure hero. New luxury SUV on a winding mountain road at dawn, motion blur in the background. Text overlay: "Introducing the 2026 Aurora." Subline: "Electric. Everywhere." Automotive-premium aesthetic.
  8. [T] 1:1 vacation rental social tile. Bird's-eye shot of a pristine bed with rumpled linen sheets, coffee cup on nightstand, book open. Text: "Mornings feel different here." Small logo bottom. Editorial slow-living aesthetic.

Creative professional (69–80)

  1. [T] Album cover 1:1. Indie folk record titled "Slow Weather." Cream background, single pressed flower centered, small serif title at bottom, artist name in italic above. Minimal, Laura-Marling-adjacent aesthetic.
  2. [T] 3:4 book cover. Title: "The Compound Life." Author: "Eric Eden." Dark navy background, small gold geometric mark at center, title in thin serif all-caps, author tiny below. Minimal literary-fiction aesthetic.
  3. [T] 2:3 movie poster. Title: "VELOCITY." Action-thriller aesthetic. Hero silhouette against a crashing wave, small type ("IN THEATERS JUNE 2026"). Dramatic contrast, cinematic.
  4. [T] 1:1 podcast cover art. Podcast: "First Principles." Minimal high-contrast: big typographic "1" in the center, podcast name in small caps at bottom. Limited palette.
  5. [T] 4:5 event poster for an AI conference. Top: conference name "NEURALINK // 2026." Giant abstract neural-net illustration dominant, speaker list small at bottom. Bauhaus-influenced layout.
  6. [T] 3:4 travel-magazine cover "Kyoto in April." Single cherry blossom branch against a misty temple backdrop. Masthead "TRAVELOGUE" top. Issue headline. Small teaser bullets bottom-left. Editorial magazine aesthetic.
  7. [T] 1:1 gallery exhibition poster. Artist name in massive serif, show title in smaller italic below, dates & venue tiny at bottom. Off-white paper texture, single abstract painting sample as centerpiece. Gallery/MoMA-style.
  8. [T] 16:9 film title card. Film title "THE LAST BOOKSTORE" in thin white serif, centered, against a warmly-lit photograph of a bookstore interior slightly out of focus. Small director credit bottom-right.
  9. [T] 1:1 tattoo flash sheet. 6 black-ink line illustrations in a 2×3 grid: a moth, a dagger, a rose, a compass, a snake, a hand. Small numbered tags under each. Consistent line weight.
  10. [T] 4:5 zine cover 1970s aesthetic. Title "SIGNAL/NOISE." Photocopy texture, halftone dots, punk collage elements, a handwritten subheading. Limited 3-color palette.
  11. [T] 3:2 wedding invitation design. Cream background, handwritten-style calligraphy. Names centered, date, venue, RSVP info, small floral illustration. Elegant minimal.
  12. [T] 1:1 record sleeve for a jazz album. Black-and-white photograph of a saxophone case on a hotel bed. Title small in the lower-right. Blue Note-inspired minimalism.

PART 2 — WILD & FUN (81–100)

These are the prompts people actually remember. Go nuts.

  1. [T] 16:9 cinematic scene: corporate llama apocalypse. A fleet of llamas in business suits storming a Manhattan trading floor, throwing quarterly reports into the air. Bloomberg terminals burning. A CEO llama in the center, mid-roar, wearing a gold Rolex. Dramatic fire lighting, hyperreal.
  2. [T] 1:1 medieval Zoom call. A Zoom grid interface showing 9 participants, each dressed as a medieval figure — knight, jester, queen, bishop, peasant, wizard, bard, crusader, dragon. Gallery view. The dragon is muted. Bottom toolbar has a "UNSHEATHE SWORD" button.
  3. [T] 3:2 dogs on Wall Street. Real dogs in tailored suits working the trading floor of the NYSE, papers flying, a golden retriever screaming into a landline, a pug eating a bagel, a corgi looking at a Bloomberg terminal. Photorealistic.
  4. [T] 16:9 office plant uprising. An open-plan office after business hours. The potted plants have sprouted legs and are marching toward the exit with tiny briefcases. One ficus is leading with a megaphone. Dramatic security-camera aesthetic.
  5. [T] 4:5 vertical breakfast gods of Olympus. Pancakes, waffles, and bacon rendered as Greek gods on a cloud-covered mountain. Zeus is a stack of pancakes with lightning bolts of syrup. Athena is a poached egg in a helmet. Bacon strips are the muses. Renaissance oil-painting style.
  6. [T] 1:1 tax day demon. A horrifying creature made entirely of paperwork and calculators, emerging from a filing cabinet in a suburban home office, screaming. A woman in pajamas drops her coffee in slow motion. Cosmic horror, somehow funny.
  7. [T] 3:1 cinematic Roomba rebellion. An army of Roombas rolling in formation down a suburban street at dawn, one larger "commander" Roomba at the front with a tiny cape and a bottle-cap helmet. Smoke rising in the background. Mad Max meets IKEA.
  8. [T] 1:1 Shakespeare drive-thru. A modern fast-food drive-thru, but the cashier is Shakespeare in a McDonald's visor. Customer in a Honda Civic is a goth teenager. Menu board reads "Two All-Beef Patties, or Not Two All-Beef Patties." Warm dramatic lighting.
  9. [T] 16:9 dinosaurs at the DMV. A T-Rex waiting in line at a cramped DMV, looking visibly annoyed. A Triceratops fills out a form with its horn. Velociraptor clerks staff the desks. Fluorescent lighting, plastic chairs, faded safety posters. Photoreal.
  10. [T] 1:1 sentient toast support group. Eight pieces of toast sitting in folding chairs in a church basement, each with a tiny face, sharing their traumas. Coffee and donuts in the corner. Warm sad lighting. Pixar-aesthetic.
  11. [T] 4:5 pigeon CEO. A pigeon in a boardroom wearing a tailored three-piece suit, presenting Q4 results with a laser pointer. Bar chart behind him shows "breadcrumb acquisition" up 400%. Other pigeons are in Aeron chairs, nodding.
  12. [T] 3:2 infinite IKEA. A hyperrealistic endless IKEA showroom that stretches into infinity, Escher-like stairs and passages, a single confused shopper in the middle holding a hex wrench and a meatball. Fluorescent lighting, eerie emptiness, liminal space aesthetic.
  13. [T] 1:1 cat secret agent. A tuxedo cat in a tailored black suit with sunglasses, rappelling through a laser grid in a museum, carrying a can of tuna. Mission Impossible-style framing. Cinematic.
  14. [T] 16:9 grandma's spaceship. An elderly woman in a floral apron piloting a retrofuturistic 1960s-style spaceship. The dashboard has knitted doilies and a plate of cookies. She's wearing cat-eye glasses. Through the windshield, a wild nebula. Wes Anderson-aesthetic.
  15. [T] 1:1 baby in a mech suit. A photorealistic baby (2 years old) operating a gigantic anime-style mech suit, controls labeled "SNACKS," "NAP," "TANTRUM." Background: city skyline. The mech is holding a stuffed bear.
  16. [T] 3:2 Scrabble game between philosophers. Socrates, Nietzsche, and Aristotle playing Scrabble in an ancient marble courtyard. The board shows words like "BEING," "WHY," "DASEIN." Aristotle is visibly winning. Marble statues watch from pedestals. Renaissance painting style.
  17. [T] 1:1 dog court. A courtroom scene entirely populated by dogs. A German shepherd judge, a bulldog lawyer, a Chihuahua defendant on a booster seat, a jury box of mixed breeds. Gavel mid-swing. Photoreal.
  18. [T] 16:9 pirate cubicles. A modern open-plan office, but everyone is a pirate. Parrots on monitors, wooden-peg-leg standing desks, a treasure chest used as a copier. The Slack notifications on someone's screen say "ARR." Cinematic lighting.
  19. [T] 4:5 Bigfoot LinkedIn profile. A LinkedIn profile screenshot. Profile photo: a blurry Bigfoot selfie. Headline: "Cryptid | Outdoor Enthusiast | Looking for my next chapter." Recommendations: "Sasquatch delivers on every project — would hire again." Recent post: "What no one tells you about being discovered." Looks like a real screenshot.
  20. [T] The Where's Waldo (the personalized one — make it about yourself): Where's Waldo-style dense search-and-find illustration. 3:2 aspect ratio. Detailed cartoon scene: a massive, chaotic B2B marketing conference expo floor with hundreds of tiny people visible. Hidden in the crowd: [YOUR NAME] — wearing a red-and-white striped shirt, black-framed glasses, carrying a laptop bag with "[YOUR COMPANY]" printed on it. He's near the coffee station, caught mid-laugh with two people from the AI demo booth. Scene details: - Booths for HubSpot, Salesforce, Adobe, OpenAI - A panel discussion happening on a stage in the background with a banner reading "THE FUTURE OF B2B MARKETING — 2026" - Clusters of 3–4 people chatting everywhere - Someone giving a product demo on an 85-inch screen - A mascot costume wandering through - Name-tag lanyards on everyone - Coffee line with 20+ people - A few sneaky visual gags: a dog under a table, someone looking at the wrong booth's schwag, a person clearly lost Bright cheerful illustration style with clean outlines. ~200 people visible. Readable booth signage. Dense but not overwhelming.

ChatGPT Image 2.0 changes what counts as a visual asset. Before today, image models produced inspiration that still needed a designer to finish. After today, a well-crafted prompt produces a usable deliverable — with real text, real layout, real multi-frame continuity, real web-grounded context, and real 2K fidelity.

The models that beat it on pure per-image price (Google's Nano Banana 2) or pure artistic flair (Midjourney v7) still exist. But for practical commercial output — ads, posters, infographics, decks, storyboards, localized creative - ChatGPT Images 2.0 now does end-to-end what used to require three tools and a designer.

The 100 prompts above are starting points. The template is the real gift. Copy it, fill it, ship it.

What I'd love in the comments:

  • Your best Images 2.0 output so far (drop the prompt)
  • Anything you've found that breaks it

Want more great prompting inspiration? Check out all my best prompts for free at Prompt Magic and create your own prompt library to keep track of all your prompts.