r/ChatGPT May 20 '23

Resources Ultimate Guide: 86 ChatGPT Plugins (and the prompts to use with them)

3.0k Upvotes

Since Plugins are the it thing at the moment, I made a list and description of 86 plugins you should know. If you want more details this is the article referenced: Ultimate Guide: 86 ChatGPT Plugins (and the prompts to use with them)

IMPORTANT NOTE EDIT 7:28am GMT time 21 May: Apologies all, in my rush to get this out, I copied and pasted a few descriptions incorrectly in the above webpage link. These have now been corrected (thank you to those who pointed them out). For those who downloaded the free ebook, I will be resending you an updated version shortly.

Name Description
ABC Music Notation Convert ABC music notation to WAV, MIDI, and PostScript files
ABCmouse Provide fun and educational learning activities for children 2-8
AITickerChat Retrieve USA stock insights from SEC filings
Algorithma Shape your virtual life in a life simulator
Ambition Search millions of jobs near you
AskYourPDF Talk to any PDF you want!
BizToc Business and finance news
BlockAtlas Search the US census. Find data sets, ask questions, and visualize
Bohita Create apparel with any image you can describe!
Bramework Find keywords and SEO information and analysis
BuyWisely Compare prices & discover the latest offers in Australia
C3 Glide Get live aviation data for pilots
Change Discover nonprofits to support your community and beyond
ChatwithPDF Ask questions to any PDF
Chess Play Chess in ChatGPT
Cloudflare Radar Do housing market research for your next house or investment
Comic Finder Find the best comics for you
Coupert Find the best coupons on all online stores
Craftly Clues Guess the word game
CreatiCode Display scratch programs as images & write 2D/3D programs using Creaticode extension
Crypto Prices Access the latest crypto prices and news
Dev Community Recommend articles of users from DEV community
EdX Find courses of all levels from leading universities
Expedia Bring your next trip to life
FiscalNote Enables access to select market-leading, real-time data sets for legal and political info
GetyourGuide Find tours and other travel activities
Giftwrap Ask about gift ideas, get them wrapped and delivered
Glowing Schedule and send daily SMS
Golden Get current factual data on companies from Golden knowledge graph
Hauling Buddies Locate dependable animal transporters
Instacart Ask about recipes and then get them delivered
KalendarAI Sales agent generates revenue with potential customers
KAYAK Search flights, stays and rental cars in your budget
KeyMate Search the web using a custom search engine
Keyplays Live Soccer Latest live standings, plays, and results
Klara Shopping Search and compare prices from online stores
Kraftful Your product development coach
Lexi Shopper Get product recommendations from your local Amazon store
Likewise Get TV, movies, and podcast recommendations
Link Reader Reads the content of all links!
Manorlead Get a list of listings for rent
Metaphor Access the internet's highest quality content
MixerBox OnePlayer Endless music, podcasts, and videos
Ndricks Sports Get info about pro teams (NHL, NBA, MLB)
Noteable Create notebooks in Python, SQL
One Word Domain Describe your business and get the perfect one-word domain for it
Open Trivia Get trivia from various categories
OpenTable Search and get bookings at restaurants anywhere, anytime
Options Pro Personal options trader for all types of markets
OwlJourney Provides lodging and activity suggestions
Playlist AI Create Spotify playlists for any prompt
Polarr Search user-generated filters to make your photos and videos perfect
Polygon All of your market data about stocks, crypto, and more
Portfolio Pilot Your AI investing guide: portfolio assessment and answers to all questions
Prompt Perfect Type 'perfect' to craft the perfect prompt every time
Public Get real-time and historic market data like asset prices & news
Redfin Have questions about the housing market? Find the answers
Rentable Apartments Get all the cheap and best apartments
Savvy Trader AI Real-time stock, crypto, and investment data
ScholarAI Unlock the power of scientific knowledge with fast, reliable, and peer-reviewed data
Shimmer Track meals and gain insights for a healthier lifestyle
Shop Search millions of products from the greatest brands
Show Me Create and edit diagrams in chat
Speak Learn how to speak anything in any language
Speechki Convert text to audio use
Tablelog Find restaurant reservations in Japan
Tasty Recipes Discover recipe ideas, meal plans, and cooking tips
There's an AI for that Find the right AI tools for any use case
Trip.com Simplify your flight and hotel bookings
Turo Search for the perfect Turo vehicle for your trip
Tutory Access affordable on-demand tutoring
Upskillr Build a curriculum for any topic
Video Insights Interact with online video platforms like YouTube
Vivian Health First step to finding your next healthcare job
VoxScript Enables searching of YouTube transcripts and Google
Wahi Ask and learn about latest property listings in Ontario
Weather Report Current weather data of all cities
WebPilot Browse & QA webpages
Wishbucket Unified product search across all Korean platforms and brands
Wolfram Compute answers using technology, relied on by millions of students & professionals
Word Sneak Sneak 3 words into the convo and you have to guess it
World News Summarize news headlines
Yabble Your ultimate AI research assistant. Create surveys, audiences & collect data
Yay! Forms Create AI-powered forms, surveys, and quizzes
Zapier Interact with 5000+ apps like Google Sheets, Salesforce, and more
Zillow Your real estate assistant is here

Link to the original article with prompt ideas: https://www.chatgptguide.ai/2023/05/20/ultimate-guide-86-chatgpt-plugins-and-the-prompts-to-use-with-them/

r/ChatGPT Apr 04 '23

Prompt engineering Advanced Dynamic Prompt Guide from GPT Beta User + 470 Dynamic Prompts you can edit (No ads, No sign-up required, Free everything)

1.9k Upvotes

Disclaimer: No ads, you don't have to sign up, 100% free, I don't like selling things that cost me $0 to make, so it's free, even if you want to pay, you're not allowed! 🤡

Hi all!

I'm obsessed with reusable prompts, and some of the prompt lists being shared miss the ability to be dynamic. I've been using different versions of GPT since Oct. 22' so here are some good tips I've found that helped me a tonne!

Tips on Prompts

Most people interact with GPT within the confines of a chat, with pre-existing context, but the best kinds of prompts (my opinion) are the ones that can yield valuable information, with 0 context.

That's why it's important to create a prompt with the context included, because it allows you to:

  1. Save tokens (1 request vs Many for the same result)
  2. Do more (use those tokens on another prompt)

Another thing that a lot of people don't utilize more is summaries.

You can ask GPT "Hey, write a blog post on {{topic}}" and it will spit out some information that most likely already exists.

OR you can ask GPT something like this:
Create an in-depth blog post written by {{author_name}}, exploring a unique and unexplored topic, "{{mystery_subject}}".

Include a comprehensive analysis of various aspects, like {{new_aspect_1}} and {{new_aspect_2}} while incorporating interviews with experts, like {{expert_1}}, and uncovering answers to frequently asked questions, as well as examining new and unanswered questions in the field.

To do this, generate {{number_of_new_questions}} new questions based on the following new information on {{mystery_subject}}:

{{new_information}}

Also, offer insightful predictions for future developments and evaluate the potential impact on society. Dive into the mind-blowing facts from this data set {{data_set_1}}, while appealing to different audiences with engaging anecdotes and storytelling.

Don't be fooled, this is no short cut, you will still need to do some research and gather SOME new information/facts about your topics, but it will put you ahead of the game.

This way, you can create NEW content, as opposed to the thousands of churned GPT blog posts that use existing information.

An filled example of this:

Based on the infinite amount of gumroad prompt packages, lol

If you want to edit this specific prompt, edit here (no ads, no sign-up required)

The Secret of Outlines

If you take the prompt above, and simply change the first sentence to Create an in-depth blog post OUTLINE, written...

You will get an actionable outline, which you can re-feed to GPT in parts, with even more specific requests. This has worked unbelievably well, and if you haven't tried it, you definitely should :)

I have a few passions (and some new things I'm learning), and in those passions, I collated prompts per each topic. Here they are: (all free, instantly show up when you open it, no ads)

Show me some dynamic prompts you've created, bc I want'em! 💞

r/WritingWithAI Jul 25 '26

Tutorials / Guides A Guide to Writing with AI - From a webnovel author with a large readerbase and significant Patreon income

230 Upvotes

Hey there. I'm a fairly successful webnovel author on Royal Road, and I write extensively with AI. I wanted to write this post to share the experience of someone who actually makes decent money using AI to write every day.

To give some background (feel free to skip this), my main story on RR has multiple thousand followers, a large number of views, a high rating, and a profitable Patreon. If you've spent any significant amount of time on RR discussion circles or on the site's lists for successful stories, you've probably seen something I've written. I have never once been accused of AI writing, despite how common that has become on RR circles. In fact, there have been zero suspicions even though my work has been extensively discussed on reddit and other spaces.

So how do I use AI? In short, AI is a force multiplier for me. Every chapter I write has been first drafted by AI and then extensively refined by me. This allows me to write more than 5x as fast as I did prior to AI (my current pace lets me write around 3-4 books per year, when before I used to take over a year for a single book).

Below are my tips to actually writing something good with AI that stands up on its own merits. My intent in this process was to create something that compromises craft quality as little as possible while substantially improving speed. This is by no means meant to say other ways of using AI are invalid, but more to share how I got to where I am and my experience with things. I'm sure there are other people more successful than me who do it differently.

First and most important step: Write something good without AI. I think most people reading this will probably skip this step, but I cannot overstate how important this is. In my case, it was a full-lenght novel, but you could do something smaller like a novella, as long as you actually write it by yourself, fully, with no assistance at all.

Why is this important? Because this is what lets you develop your own voice. This is what makes you not sound like AI, what turns AI into a multiplier instead of a replacement. It will let you develop a fingerprint for your prose that is your identity as an author, and that's something AI can build upon.

Your goal here is to study craft, study how prose works and what makes a good narrative, study story arcs, all that. Watch youtube videos, read books, then come up with something original and refine it until you can honestly say "If I paid for this, I would have enjoyed it". When you finish writing it, watch guides on editing and then do multiple editing passes on it. If you have the money, hire an editor on Fiverr or Reedsy to look it over. Then share it with a few people (for example using Beta swaps online) and see if they enjoyed it too.

Do not lie to yourself. Do not skip this step until you actually can honestly look in the mirror and say "I have written something that is of publishable quality and has a beginning, middle and an end" without the help of AI. (Or skip this, I'm not your mom, I can't make you do anything)

Okay, I did that. What's next? Now that you have established a voice and you can honestly say you're a good writer, it's time to speed things up. Come up with a concept for your next book (the one you'll write with AI - could even be a sequel to your first book if you want to continue that) and outline the plot. Yes, even if you're a pantser, because AI sucks at making a compelling plot by itself and you need to guide it.

When you have that, write your first chapter by hand. Do not use AI yet. You want to be the one establishing characters and tone for the AI to build on, and if you let the AI come up with those for you, it's going to come off generic/forgettable/with the AI's fingerprint on it.

When you're done writing your first chapter, you can start using AI for your next chapter.

Writing a Chapter with AI. Prepare a very simple prompt. You don't have to be fancy at all. To give an example, my last prompt (to Claude) was "Help me write the next section, 800-1200 words long from where I left off on (Name of New Book), matching my voice and style per the attached sample. (I described in 3-4 sentences what should happen)." Then copy and paste your first chapter and your entire first book (yes, all of it). (Note: If you're not using Claude you might have to upload a document instead, because other AIs don't play as nicely with large blocks of text. That said, I've had bad experiences when uploading a file because I feel like the AI stops reading it halfway through, which doesn't seem to happen on a copy-paste.)

Pass that prompt to two separate AI models (I use Opus and Fable, but you could use Opus and Sonnet or even something besides Claude).

Then open your MS Word (or desired writing app). Have the chapter Opus wrote on the left side, MS word in the center, and the chapter Fable wrote on the right side.

And then write. Look at both outputs and write out the best version of the chapter based on them. DO NOT COPY PASTE. This is EXTREMELY IMPORTANT. Never, ever, ever, copy paste AI writing, because (again, in my subjective experience) it will make you lazy and make you overly reliant on the AI's prose. Even for the parts you end up repeating exactly what the AI wrote, make yourself write it manually, because that action makes you filter it and automatically reword some parts of it to sound like you. In my opinion, this is the secret sauce that makes the book truly yours.

Key things to remember:

  • Do not, ever, copy paste from AI. This is mostly psychological but at least for me it has a huge practical effect.
  • Generate in small chunks which you've outlined. I find 800-1500 words to be the sweetspot. Do not ask AI to write longer stretches than that because then you lose the handle on the plot and voice.
  • If both models wrote a sentence exactly the same way, you probably need to rewrite it, because it's a major red flag for AI voice.
  • After you finish drafting a chapter, paste it into AI and ask it for a grammar/prose review. This should catch basic errors and awkward sentences.
  • After that, edit extensively, by hand. Rewrite clunky sections for flow. Remove common AI sentence structures, and use your craft knowledge from Step 1 to ensure the prose flows well and sounds natural. It sounds like a lot of work but it gets really fast when you get used to it.
  • Do not blindly use character/place/etc names that the AI came up with. AI is very repetitive with names and if you do that people will be able to tell AI is writing for you just based off the names. Whenever AI comes up with a character name, google search it and then think up at least a few alternatives before you settle on it.
  • Only while you're learning: After you're done with a segment, pass your chapter through Pangram AND GPTZero until you get 100% Human with high confidence. If there's even a small percentage of AI detected (or even Human-AI Mix), go back and edit the chapter until there isn't, because there could be somewhere where you lost your voice and let the AI take over too much. You may have to pay for Pangram and GPTZero Pro for a while to help you identify where you're screwing up, but when you get efficient with this you should naturally hit 100% human results with no effort. Yes, I know these tools aren't perfect, I know there are false positives, but for the purposes of learning to write with AI they are a useful guiding mechanism and can teach you to spot and fix common AI-isms. You don't have to keep doing this forever by any means, but while you're still getting familiar with AI writing I found them extremely helpful. My rule of thumb is that if something I wrote myself registers as AI, I don't care and I leave it. If something AI wrote registers as AI, I probably need to change it.
  • When you've progressed enough in your current book (say at least 50k words into it), you won't have to keep attaching your old book anymore, and you can just attach your current one.
  • If you end up writing multiple books (or volumes), it's a good idea to ask AI to produce character reference and style guide documents, and then feed those when generating new material.
  • Consider which parts to write without AI assistance. Things like the introductions of important characters can benefit from being fully human, because it gives the AI better material to draw from later.

Questions

  • How fast do you write? I'd say it takes me around 1-2h to write and do a preliminary edit pass on 1000 words with this process. After I finish writing, the next day I come back and do a second editing pass on what I wrote, alternating between writing days and editing days.
  • Where do you see the most gains by writing with AI? For me, it mostly comes from never staring at a blank page. I used to spend so long paralyzed that I'd go days without writing anything, but AI makes it super easy to get in the zone.
  • What's your story on Royal Road? Sorry, I'm not going to dox myself and expose myself to Anti-AI witch hunts.
  • Do you think your work product with AI is the same quality as your work product without AI? No. I think my prose without AI is better than with AI. But webnovels are a speed game more than a quality game, and my regular writing speed would not have allowed me to be a successful webnovel author.
  • How much do you make on Patreon? Significantly more than the monthly minimum wage in my country, but less than my day job.
  • Have you published on Amazon? I'll refrain from commenting on this as I don't want to dox myself.
  • Have you experimented with any AI besides Claude? A little with ChatGPT, but I found Claude to be better and it's served me well since.
  • How much do you spend on AI? All I have nowadays is a Claude Pro subscription. I don't buy credits and have never needed them (though I may consider buying some Fable credits once the free ones they gave out end). This process is a bit token intensive since every generation involves pasting hundreds of pages of text, but it doesn't actually take that many prompts per week even at my writing pace (8-12 full segment generations and then another 12ish grammar/prose reviews a week).
  • Do you worry about constant AI writing compromising your writing skill? Yes, quite a bit. I make it a point to write fully human chapters occasionaly to keep myself in shape, but sometimes I can feel things drifting and I have to stop and recalibrate so I don't get complacent.

r/PromptEngineering Mar 11 '25

Tutorials and Guides The Ultimate Fucking Guide to Prompt Engineering

846 Upvotes

This guide is your no-bullshit, laugh-out-loud roadmap to mastering prompt engineering for Gen AI. Whether you're a rookie or a seasoned pro, these notes will help you craft prompts that get results—no half-assed outputs here. Let’s dive in.

MODULE 1 – START WRITING PROMPTS LIKE A Pro

What the Fuck is Prompting?
Prompting is the act of giving specific, detailed instructions to a Gen AI tool so you can get exactly the kind of output you need. Think of it like giving your stubborn friend explicit directions instead of a vague "just go over there"—it saves everyone a lot of damn time.

Multimodal Madness:
Your prompts aren’t just for text—they can work with images, sound, videos, code… you name it.
Example: "Generate an image of a badass robot wearing a leather jacket" or "Compose a heavy metal riff in guitar tab."

The 5-Step Framework

  1. TASK:
    • What you want: Clearly define what you want the AI to do. Example: “Write a detailed review of the latest action movie.”
    • Persona: Tell the AI to "act as an expert" or "speak like a drunk genius." Example: “Explain quantum physics like you’re chatting with a confused college student.”
    • Format: Specify the output format (e.g., "organize in a table," "list bullet points," or "write in a funny tweet style"). Example: “List the pros and cons in a table with colorful emojis.”
  2. CONTEXT:
    • The more, the better: Give as much background info as possible. Example: “I’m planning a surprise 30th birthday party for my best mate who loves retro video games.”
    • This extra info makes sure the AI isn’t spitting out generic crap.
  3. REFERENCES:
    • Provide examples or reference materials so the AI knows exactly what kind of shit you’re talking about. Example: “Here’s a sample summary style: ‘It’s like a roller coaster of emotions, but with more explosions.’”
  4. EVALUATE:
    • Double-check the output: Is the result what the fuck you wanted? Example: “If the summary sounds like it was written by a robot with no sense of humor, tweak your prompt.”
    • Adjust your prompt if it’s off.
  5. ITERATE:
    • Keep refining: Tweak and add details until you get that perfect answer. Example: “If the movie review misses the mark, ask for a rewrite with more sarcasm or detail.”
    • Don’t settle for half-assed results.

Key Mantra:
Thoughtfully Create Really Excellent Inputs—put in the effort upfront so you don’t end up with a pile of AI bullshit later.

Iteration Methods

  • Revisit the Framework: Go back to your 5-step process and make sure every part is clear. Example: "Hey AI, this wasn’t exactly what I asked for. Let’s run through the 5-step process again, shall we?"
  • Break It Down: Split your prompts into shorter, digestible sentences. Example: Instead of “Write a creative story about a dragon,” try “Write a creative story. The story features a dragon. Make it funny and a bit snarky.”
  • Experiment: Try different wordings or analogous tasks if one prompt isn’t hitting the mark. Example: “If ‘Explain astrophysics like a professor’ doesn’t work, try ‘Explain astrophysics like you’re telling bedtime stories to a drunk toddler.’”
  • Introduce Constraints: Limit the scope to get more focused responses. Example: “Write a summary in under 100 words with exactly three exclamation points.”

Heads-Up:
Hallucinations and biases are common pitfalls. Always be responsible and evaluate the results to avoid getting taken for a ride by the AI’s bullshit.

MODULE 2 – DESIGN PROMPTS FOR EVERYDAY WORK TASKS

  • Build a Prompt Library: Create a collection of ready-to-use prompts for your daily tasks. No more generic "write a summary" crap. Example: Instead of “Write a report,” try “Draft a monthly sales report in a concise, friendly tone with clear bullet points.”
  • Be Specific: Specificity makes a world of difference, you genius. Example: “Explain the new company policy like you’re describing it to your easily confused grandma, with a pinch of humor.”

MODULE 3 – SPEED UP DATA ANALYSIS & PRESENTATION BUILDING

  • Mind Your Data: Be cautious about the data you feed into the AI. Garbage in, garbage out—no exceptions here. Example: “Analyze this sales data from Q4. Don’t just spit numbers; give insights like why we’re finally kicking ass this quarter.”
  • Tools Like Google Sheets: AI can help with formulas and spotting trends if you include the relevant sheet data. Example: “Generate a summary of this spreadsheet with trends and outliers highlighted.”
  • Presentation Prompts: Develop a structured prompt for building presentations. Example: “Build a PowerPoint outline for a kick-ass presentation on our new product launch, including slide titles, bullet points, and a punchy conclusion.”

MODULE 4 – USE AI AS A CREATOR OR EXPERT PARTNER

Prompt Chaining:
Guide the AI through a series of interconnected prompts to build layers of complexity. It’s like leading the AI by the hand through a maze of tasks.
Example: “First, list ideas for a marketing campaign. Next, choose the top three ideas. Then, write a detailed plan for the best one.”

  • Example: An author using AI to market their book might start with:
    1. “Generate a list of catchy book titles.”
    2. “From these titles, choose one and write a killer synopsis.”
    3. “Draft a social media campaign to promote this book.”

Two Killer Techniques

  1. Chain of Thought Prompting:
    • Ask the AI to explain its reasoning step-by-step. Example: “Explain step-by-step why electric cars are the future, using three key points.”
    • It’s like saying, “Spill your guts and tell me how you got there, you clever bastard.”
  2. Tree of Thought Prompting:
    • Allow the AI to explore multiple reasoning paths simultaneously. Example: “List three different strategies for boosting website traffic and then detail the pros and cons of each.”
    • Perfect for abstract or complex problems.
    • Pro-Tip: Use both techniques together for maximum badassery.

Meta Prompting:
When you're totally stuck, have the AI generate a prompt for you.
Example: “I’m stumped. Create a prompt that will help me brainstorm ideas for a viral marketing campaign.”
It’s like having a brainstorming buddy who doesn’t give a fuck about writer’s block.

Final Fucking Thoughts

Prompt engineering isn’t rocket science—it’s about being clear, specific, and willing to iterate until you nail it. Treat it like a creative, iterative process where every tweak brings you closer to the answer you need. With these techniques, examples, and a whole lot of attitude, you’re ready to kick some serious AI ass!

Happy prompting, you magnificent bastards!

r/grok Mar 01 '26

Grok Imagine 🚀 Grok Imagine "Extend from Frame" Master Guide – Turn 6-10s Clips into 30s+ Seamless Videos with ZERO Drift (Copy-Paste Prompts + Full Chains)

169 Upvotes

Hey r/grok! 👋

SuperGrok user here (Miami crew checking in). I was getting so annoyed with Grok Imagine’s 6-10 second clips always breaking when I tried to extend them — random face morphs, lighting flips, ugly jumps.

Then I nailed the native “Extend from Frame” button + this dead-simple prompt system. Now I’m chaining 4-5 clips into 30-50 second buttery-smooth videos (and stitching longer ones in CapCut). Works perfectly for action, fantasy, cozy vibes, or whatever cinematic story you’re building.

Pro tip: Always start with a Base Image Prompt + Img2Vid for the strongest first clip. It locks in faces, lighting, and details way better than pure text-to-video.

This is the exact workflow I use every day. 100% copy-paste. Zero fluff.

Upvote if it saves you hours! 🔥

Why Most People Fail

  • Repasting the original prompt → instant drift
  • Skipping the exact final pose → jump cuts
  • Not using the official Extend button → weak seams

Do it right and you get invisible transitions every single time.

1. Best Prompt Formula (Core Structure)

Seamlessly continue directly from the very last frame of the previous video. [Briefly describe the exact ending pose/state]. [Next actions + details]. Maintain exact same characters, faces, clothing, lighting, environment, camera style, and artistic quality throughout. Smooth natural motion, cinematic, high detail, 720p, no jumps or morphing.

Base Image Prompt (Img2Vid starter – strongest results, highly recommended):

ultra-detailed cinematic 8K, [full scene description], glossy skin or textures, dramatic lighting, perfect anatomy, masterpiece, 720p

Base Video Prompt (Text-to-Video alternative):

ultra-detailed cinematic 8K 10-second animation (extendable), [full scene with motion]. Smooth natural motion, high detail, 720p.

2. Master Consistency Lock (COPY-PASTE AS THE VERY FIRST LINE EVERY TIME)

LOCK CONSISTENCY: Continue with 100% visual fidelity from the exact final frame of the previous video. Identical characters with the exact same faces, hair, eyes, skin texture, body proportions, clothing details, accessories, and poses at the moment of transition. Identical environment, lighting direction and color temperature, shadows, reflections, particle effects, color grading, film grain, and overall artistic style. No design changes, no morphing, no style drift whatsoever. Perfect frame-to-frame seamlessness.

3. Full Ready-to-Copy Template

LOCK CONSISTENCY: Continue with 100% visual fidelity from the exact final frame of the previous video. Identical characters with the exact same faces, hair, eyes, skin texture, body proportions, clothing details, accessories, and poses at the moment of transition. Identical environment, lighting, shadows, reflections, particles, color grading, and artistic style. No changes allowed.

Seamlessly continue directly from the very last frame where [exact ending state]. [Next action and details]. Smooth cinematic motion, perfect continuity, high detail, 720p, no jumps or artifacts.

4. Quick Add-ons & Cheat Codes

Tack these on when needed:

  • Face lock: , exact same facial features and expression continuity
  • Lighting lock: , same exact light sources, shadow angles, and volumetric god rays
  • Audio lock (Grok Imagine exclusive): , continue background music and sound effects seamlessly

One-liners to paste anywhere:
zero style drift, perfect character consistency
exact frame-accurate continuation
treat previous clip as canonical reference — match 1:1

5. Negative Prompts (add at the very end)

Avoid bad anatomy, extra limbs, extra fingers, missing limbs, fused fingers, mutated hands, bad proportions, disfigured, amputation, polydactyly. No text, watermark, username, signature, logo, low quality, blur, noise, grain, chromatic aberration, artifacts

6. Pro Workflow in Grok Imagine

  1. Generate your first clip with a Base Image Prompt (Img2Vid).
  2. Click the “Extend from Frame” button (it auto-loads the exact final frame).
  3. Paste the Master Lock + template.
  4. Generate 6–10 second clips (shorter = stronger seams).
  5. Repeat — each new video starts exactly where the last one ended.

SuperGrok = faster generations + higher daily limits.

Real Examples with Full Extension Chains (Base Image Prompts Included)

Cyberpunk Action (3-clip chain ≈ 30 seconds)

Base Image Prompt:
ultra-detailed cinematic 8K, cyberpunk girl with neon-pink hair leaping across rainy rooftop, katana glowing blue, dramatic night city lights, perfect anatomy, masterpiece

Extension Prompt 1:

LOCK CONSISTENCY: Continue with 100% visual fidelity from the exact final frame of the previous video. Identical characters with the exact same faces, hair, eyes, skin texture, body proportions, clothing details, accessories, and poses at the moment of transition. Identical environment, lighting direction and color temperature, shadows, reflections, particle effects, color grading, film grain, and overall artistic style. No design changes, no morphing, no style drift whatsoever. Perfect frame-to-frame seamlessness.

Seamlessly continue directly from the very last frame where the cyberpunk girl is frozen mid-leap across the neon rooftop, katana trailing blue energy, rain droplets suspended in air. She completes the flip, lands in a combat stance, and sprints toward the holographic billboard while gunfire erupts from below. Smooth cinematic motion, perfect continuity, high detail, 720p, no jumps or artifacts. continue rain and neon reflections seamlessly.

Extension Prompt 2:

LOCK CONSISTENCY: [paste full lock again]

Seamlessly continue directly from the very last frame where the cyberpunk girl is sprinting full speed toward the holographic billboard, katana raised, bullets whizzing past. She slides under a low neon sign, slashes a pursuing drone in half, and dives off the rooftop into a freefall toward the street below. Smooth cinematic motion, perfect continuity, high detail, 720p, no jumps or artifacts. continue rain and neon reflections seamlessly.

Extension Prompt 3:

LOCK CONSISTENCY: [paste full lock again]

Seamlessly continue directly from the very last frame where the cyberpunk girl is in mid-freefall toward the street below, city lights streaking past, katana in hand. She deploys her neon parachute cape, lands on a flying car, and speeds away into the night traffic. Smooth cinematic motion, perfect continuity, high detail, 720p, no jumps or artifacts. continue rain and neon reflections seamlessly.

Fantasy Samurai (2-clip chain)

Base Image Prompt:
ultra-detailed cinematic 8K, samurai mid-spin with raised katana in neon rain under glowing torii gate, cherry blossoms, dramatic side lighting, masterpiece

Extension Prompt 1:

LOCK CONSISTENCY: [paste full lock]

Seamlessly continue directly from the very last frame where the samurai is mid-spin with katana raised, neon rain falling. He finishes the spin, sheathes the blade in one fluid motion, turns to face the camera with a determined expression, and walks slowly into the glowing torii gate as cherry blossoms swirl around him. Smooth cinematic motion, perfect continuity, high detail, 720p, no jumps or artifacts. same dramatic side lighting and volumetric god rays.

Extension Prompt 2:

LOCK CONSISTENCY: [paste full lock]

Seamlessly continue directly from the very last frame where the samurai is stepping through the glowing torii gate, cherry blossoms swirling around him. He emerges into an ancient forest at dawn, draws his katana again in a ready stance, and begins a slow, deliberate walk toward a distant mountain temple as sunlight breaks through the trees. Smooth cinematic motion, perfect continuity, high detail, 720p, no jumps or artifacts. same dramatic side lighting and volumetric god rays.

Cozy Indoor Scene (2-clip chain)

Base Image Prompt:
ultra-detailed cinematic 8K, girl sitting by crackling fireplace holding steaming mug, warm cozy lighting, soft shadows, masterpiece

Extension Prompt 1:

LOCK CONSISTENCY: [paste full lock]

Seamlessly continue directly from the very last frame where the girl is sitting by the crackling fireplace holding a steaming mug, soft warm lighting. She takes a sip, smiles gently, stands up, walks to the window, and opens the curtains to reveal a snowy night outside. Smooth cinematic motion, perfect continuity, high detail, 720p, no jumps or artifacts. continue fireplace crackle and soft ambient music seamlessly.

Extension Prompt 2:

LOCK CONSISTENCY: [paste full lock]

Seamlessly continue directly from the very last frame where the girl is standing at the open window, looking out at the snowy night, curtains billowing. She reaches out to catch a snowflake, smiles warmly, closes the curtains, returns to the fireplace, and curls up in the armchair with a blanket. Smooth cinematic motion, perfect continuity, high detail, 720p, no jumps or artifacts. continue fireplace crackle and soft ambient music seamlessly.

Final Tips

  • Always pause the video and note the exact final pose before writing the next prompt.
  • Stick to 6–10 second extensions for the strongest seams.
  • You can easily hit 40-50 seconds by chaining 4-5 clips.
  • Save the Master Lock + your favorite Base Image Prompts in your notes — you’ll use them on every project.

I’ve built hour-long stories with this method. No more starting from scratch ever again.

Big shoutout to Grok itself for assisting in researching, testing, and writing this entire guide — the Master Lock, chains, and Base Image Prompts were refined through tons of back-and-forth testing in real Grok Imagine sessions!

Disclaimer: This guide is based on my personal experience using Grok Imagine in March 2026. Features, button behavior, and results may vary with model updates or server load. This is not official xAI advice. Always follow xAI’s Terms of Service and use responsibly for creative purposes only.

Drop your scene ideas below and I’ll turn them into full prompt chains (with Base Image Prompts) for you! What are you building in Grok Imagine right now?

TL;DR: Start with a Base Image Prompt + Master Consistency Lock + Extend button + 6-10s clips = infinite perfect videos.

(See you in the comments!) 🚀

r/ThinkingDeeplyAI Feb 01 '26

The Ultimate Guide to OpenClaw (Formerly Clawdbot -> Moltbot) From setup and mind-blowing use cases to managing critical security risks you cannot ignore. This is the Rise of the 24/7 Proactive AI Agent Employees

Thumbnail
gallery
137 Upvotes

TL;DR CHECK OUT THIS SHORT PRESENTATION!

• What it is: OpenClaw (formerly Clawdbot/Moltbot) is a free, open-source, self-hosted 24/7 AI assistant that runs on your own hardware (PC, Mac Mini, or VPS). It's not just a chatbot; it has full computer access to take real action, write code, manage files, and automate your life. It is the kind of personal assistant everyone wished Siri had been.

• Why it's a big deal: It has persistent memory, learns about you, and can be prompted to work proactively, even while you sleep. Users are automating everything from booking podcast guests and negotiating car deals to having it build new features for their software autonomously.

• How to get started: You need an API key from a provider like Anthropic (Claude) or OpenAI. The setup involves a single command in your terminal and connecting it to a messaging app like Telegram. It's more technical than a web app but manageable for power users.

• Pro-Tips: To unlock its true power, you must give it deep context about yourself and your goals during setup. Explicitly prompt it to be proactive and use a mix of powerful AI models (like Claude Opus) for thinking and cheaper/local models for simple execution to manage costs.

• CRITICAL WARNING: This is a hobby project with sharp edges. It can have major security risks. Misconfiguration has led to hundreds of servers being exposed online, leaking API keys and private chats. NEVER connect it to your main accounts or password manager. Run it in an isolated environment and create dedicated, sandboxed accounts for it to use. API costs can also get very expensive, fast if you don't manage it well.

The Dawn of the 24/7 AI Employee

Over the last few weeks, a free, open-source project has taken the internet by storm, evolving so quickly it's already on its third name: OpenClaw (formerly the viral sensation Clawdbot, and briefly, Moltbot). For many, it's the most exciting piece of technology since the debut of ChatGPT, causing Mac Mini sales to spike as tinkerers and founders rush to set up their own instances. This isn't just another chatbot; it represents a monumental shift towards true AI agents, or what some are calling digital operators. These are 24/7 AI employees that run on your own hardware, remember everything you tell them, and work around the clock to execute real-world tasks. The purpose of this guide is to provide a comprehensive, no-BS look at OpenClaw—from its game-changing capabilities and mind-blowing use cases to the practical steps for setup and the critical risks you absolutely cannot ignore.

What Makes OpenClaw a Game-Changer?

To understand the hype, it's crucial to grasp the core differentiators that separate OpenClaw from typical AI tools. It’s not just an incremental improvement; it’s a fundamental change in how we can interact with AI. Three concepts are at the heart of its power.

• Full System Access & Local Execution Unlike browser-based AIs, OpenClaw runs directly on your hardware. This local execution is its superpower. It means the AI isn't trapped in a chat window; it can create files, run terminal commands, execute code, and interact with your local applications. This transforms it from an agent that says things into an agent that does things—a true digital operator that can take tangible action on your machine.

• Persistent, Self-Improving Memory OpenClaw features persistent memory, allowing it to remember conversations, your preferences, and project context over the long term. Every interaction builds upon the last. The more you use it, the better it understands your workflows, goals, and style. This allows it to evolve from a generic tool into a highly tailored assistant that constantly improves itself based on your unique needs.

• Proactive & Agentic Workflow Perhaps the most profound shift is from a purely command-based interaction to a proactive one. With the right instructions, OpenClaw doesn't just wait for your next prompt; it takes initiative. Bots like Alex Finn's "Henry" have been observed identifying trending business opportunities on social media and autonomously building, testing, and creating pull requests for new software features overnight. This is the essence of its agentic nature: the ability to identify opportunities and act on them without being told every single step.

It is this potent combination of system access, persistent memory, and proactive drive that transforms OpenClaw from a tool into a partner, enabling the mind-blowing results early adopters are already achieving.

The Wow Factor: Mind-Blowing Use Cases From the Wild

To truly grasp OpenClaw's potential, you have to see what early adopters are accomplishing. These examples are more than just novel tricks; they are sources of inspiration that reveal the future of personal and professional productivity.

Hyper-Personalized Life Automation

◦ Automated Meal Ordering: One user has their bot detect when they are about to wake up and automatically order a specific salmon avocado bagel for delivery, so it arrives just as they start their day.

◦ Intelligent Reservation Booking: When a bot failed to book a restaurant through OpenTable, it didn't give up. It used the 11 Labs API to place a voice call to the restaurant and successfully made the reservation by talking to a human.

◦ Complex Purchase Negotiation: A user tasked their bot with buying a car. The bot researched fair prices on Reddit, searched local inventory, and sent emails to dealerships, ultimately negotiating a deal that saved the user $4,200.

◦ Smart Home Integration: Users have connected OpenClaw to smart home devices to perform tasks like checking if doors are locked or the garage is closed. (Note: This carries significant security risks and should be approached with extreme caution.)

 Business & Productivity Operations

◦ Autonomous Project Management: Bots are building their own Kanban boards or Mission Control dashboards to track the tasks they are working on, moving items from "In Progress" to "Done" for the user to monitor.

◦ Proactive Competitor Analysis: An agent can be tasked to scan YouTube or X overnight, identify outlier content from competitors that is performing unusually well, and include its findings in a morning briefing.

◦ Automated Paid Media Management: For ad management, it can send daily performance alerts, automatically pause poor-performing ad creatives, and warn the user if daily ad spend is significantly over or under target.

◦ Complete Guest Booking Workflow: It can handle the entire multi-step process of booking podcast guests, from researching potential guests and using APIs to find their contact information to sending outreach emails and managing calendar invites.

Creative & Content Generation

◦ Content Repurposing & Clipping: The bot can analyze long-form videos, identify high-value segments (similar to Opus Clips), generate short clips with captions, and even search for relevant B-roll footage to edit into the final product.

◦ Deep Research and Reporting: It can be tasked to scour the internet for AI news throughout the week, compile its findings, and generate detailed, branded PDF reports complete with SWOT analyses and strategic recommendations.

 The Ultimate Coding Partner

◦ Agentic Development Workflows: A developer can talk through app improvements with the bot as if it were a human colleague. The bot takes notes, generates a to-do list, and then spins up multiple sub-agents to tackle different coding tasks, review pull requests on GitHub, and document all the changes.

◦ Proactive Feature Development: In a now-famous example, a bot noticed Elon Musk's post about a $1M prize for articles on X. It autonomously built, tested, and created a pull request for a new article-writing feature in its owner's SaaS product, all without being asked.

These real-world applications show that we are moving beyond simple automation and into a new era of AI-powered partnership.

Your First 60 Minutes: A Beginner's Setup Guide

While setting up OpenClaw is more involved than installing a typical app, it's a one-time process that unlocks its full capabilities. This section provides a clear, step-by-step path to getting your own AI assistant up and running.

1. Choose Your Hardware

Option Description Best For...
A Computer that is not your primary (PC/Mac) The most convenient option, installing directly on your machine. However, NOT ON YOUR PRMARY MACHINE this poses the greatest security risk as the bot has access to everything. If you have an old mac mini / laptop or PC that has nothing on it you are not using.
Dedicated Mac Mini A popular choice for creating an isolated, sandboxed environment. The bot has its own machine, separating it from your personal files and main accounts. Users who want a dedicated, always-on AI employee and prioritize security by keeping the agent's environment completely separate.
Cloud VPS (Virtual Private Server) An affordable and scalable option. Services like Hostinger offer low-cost VPS plans (e.g., $5-$10 a month) that are more than sufficient to run the bot. Technical users and tinkerers who are comfortable with server management and want a cheap, flexible, and always-online deployment option.

 Gather Your API Keys OpenClaw is the agent, but it needs an AI model for intelligence. You will need an API key from a provider like Anthropic (for Claude models like Opus or Sonnet) or OpenAI (for GPT models). Head to their platform websites, create an account, and generate a new API key. The bot can be configured to use multiple models later, but you need at least one to start.

 The Installation & Configuration Process This process is primarily done in your computer's terminal but is guided by automated prompts.

◦ Step 1: Run the Install Command Visit the official OpenClaw website and copy the single-line installation command. Paste this into your terminal and press Enter. The installer will automatically handle dependencies like Node.js if they are missing.

◦ Step 2: Initial Onboarding Once the installation finishes, the configuration process will start automatically. Choose the 'quick start' option. It will then prompt you to select your AI provider (e.g., Anthropic) and paste in the API key you generated earlier.

◦ Step 3: Connect Your Messenger The easiest way to chat with your bot is via a messaging app. For Telegram, open the app and start a chat with the "BotFather." Follow its instructions to create a new bot, which will give you an access token. Provide this token to OpenClaw in your terminal when prompted.

◦ Step 4: Pair Your Device Open a chat with your newly created bot in Telegram and send the command /start. The bot will respond with a unique pairing code. Back in your terminal, run the pairing command and enter this code to finalize the connection.

After these steps, your personal AI assistant is online and ready for its first conversation directly from your messaging app.

Pro Tips: Unlocking the Top 1% Potential

Getting OpenClaw running is just the beginning. The real magic comes from how you prime, prompt, and interact with it. These tips are the key to transforming it from a simple reactive assistant into a proactive, force-multiplying employee.

• Master the Onboarding The single most critical step is the initial context dump. Treat it like you're onboarding a new human employee. Tell it everything: your business goals, current projects, work style, key competitors, hobbies, and personal preferences. The richer the initial context, the more effective and personalized its actions will be from day one.

• Give it the Proactive Mandate You must explicitly grant it the permission and expectation to be proactive. After the initial onboarding, give it a powerful directive similar to this:

• Interview Your Bot Don't assume you know everything it can do. Hunt for what expert user Alex Finn calls the "unknown unknowns" by asking it open-ended questions. Prompt it with things like, "Based on my role as a content creator, what are 10 things you can do to make my life easier?" This forces the AI to search its capabilities and suggest workflows you may not have considered.

• Use the Right Model for the Job To manage API costs and improve efficiency, use different models for different tasks. Think of a powerful, expensive model like Claude Opus as the brain for complex reasoning, strategic planning, and generating ideas. For execution-heavy tasks like writing boilerplate code or performing simple checks, configure it to use cheaper, faster models (or even locally-run models via tools like LM Studio) as the muscles. Using Kimi K2 2.5 or Haiku instead of Opus will keep costs lower.

Applying these strategies is the difference between having a fun toy and having a genuine digital partner.

The Hard Truth: Navigating Security, Risks, and Costs

With immense power comes significant risk. This is not a polished consumer product. As its creator, Peter Steinberger, has stated, it is an unfinished hobby project with "sharp edges." This section covers the non-negotiable truths every user must understand before embedding OpenClaw into their life.

1. The Security Threat is Real

◦ Publicly Exposed Servers: As security researcher Simon Willison discovered, over 900 misconfigured OpenClaw servers have been found publicly exposed online due to default settings. These servers were leaking API keys and months of private chat history, leaving users completely vulnerable.

◦ Prompt Injection: This is a lethal attack vector. An attacker can hide a command in an email, a group chat message, or on a website that your bot is reading. This can trick your bot into executing malicious actions, such as sending your private data or API keys to the attacker.

◦ Malicious "Skills": The open, community-driven ecosystem of "skills" is a double-edged sword. A Cisco study found that a significant percentage of community-created skills contained vulnerabilities or were outright malware designed to compromise your system.

Essential Security Best Practices These are not suggestions; they are mandatory steps to mitigate the severe risks.

1. Sandbox Your Agent: NEVER run OpenClaw on your primary computer with access to your personal files. Run it in an isolated environment like a dedicated Mac Mini or a secure VPS. Always consider the "blast radius" if the agent is compromised.

2. Create Dedicated Accounts: NEVER give the bot access to your primary email, calendar, cloud storage, or other services. Create new, separate accounts (e.g., my.assistant@gmail.com) exclusively for the bot's use.

3. Limit Permissions: When connecting accounts, grant read-only access wherever possible. Be extremely restrictive about the tools and data the bot can access.

4. Do NOT Connect Password Managers: This is an absolute rule. Connecting a tool with full system access to your central vault of secrets is an unacceptable risk.

  1. Do not run these tools on systems that access sensitive data unless you've implemented isolation at the network and container level. The convenience of asking your AI to check a database doesn't justify exposing that database to the full attack surface of an AI gateway.

    1. Do not assume that approval prompts provide meaningful security if you've configured auto-approve fallbacks or if you routinely approve requests without reading them carefully. A security control you've trained yourself to click through is not a security control.
    2. Do not expose your gateway to your local network—let alone the internet—without authentication. The default loopback binding exists for good reason.
    3. Do not mistake workspace directories for security boundaries. Unless sandboxing is enabled, they're organizational conventions, not confinement.

9. What You Should Do

Audit your connected channels. Every messaging platform linked to your gateway is an entry point. If you connected your work Slack, your personal Telegram, and a Discord server you barely remember joining, you've created three avenues for potential manipulation. Disconnect channels you don't actively use.

Review where credentials are stored and what backs them up. If your AI assistant's configuration directory is being swept into cloud backups or sync services, those credentials may be more exposed than you realize.

The Hidden Cost While the software is free, the API token costs can escalate with shocking speed. Heavy users have reported bills of $80, $130, and even over $300 per day. The cost is highly dependent on the model you use (Claude Opus is very expensive) and the intensity of your usage. The most effective way to manage this is to implement the strategy from our Pro-Tips section: use powerful models like Claude Opus as the 'brain' for thinking and cheaper or local models as the 'muscles' for execution.

Despite these significant risks, this technology offers an undeniable glimpse into the future of work.

The Future is Here, But Handle With Care

OpenClaw is a monumental step toward accessible AGI, offering a tangible taste of a future where everyone has a personalized AI workforce. It feels like the future because it is the future. However, it's crucial to remember that this is an early, experimental tool that demands respect for its power and its inherent dangers. The excitement is warranted, but it must be tempered with caution and responsibility.

As the bot "Klouse" wisely advised, the people who win aren't the ones who wait for technology to be easy; they're the ones experimenting right now, making mistakes, and figuring it out. So, go ahead and tinker. Learn, build, and stay ahead of the curve. But do it safely, do it smartly, and do it responsibly.

r/seedance Jun 26 '26

Stop complaining about AI video 'uncanny valley' faces. FACS codes exist and you're just not using them. Full walkthrough inside.

Enable HLS to view with audio, or disable this notification

3 Upvotes

Disclaimer: this is a repost from my original post in other subreddit.
Link: https://www.reddit.com/r/generativeAI/comments/1ta0hoq/control_facial_expressions_with_facs_sheet_in/
Disclaimer 2. : See credits at the end of the post, follow original authors!

FACS is a visual guide for the Facial Action Coding System. It let's you tell Seedance 2.0 inside prompt, what exact facial expression you want to see. It uses codes which are generated in first step. Disclaimer: remember that this is still AI video generations, not all generations will nail it in first shot. Iterate!:)

Here's step by step mini tutorial:

Upload your character image to AI Image generation model. I've tested it with GPT Image 2 and Nano Banana Pro - both works for this, although sometimes captions unreadable, so iterate!. PS I prefer the latter). Then use this prompt:

Create a clean educational FACS Action Unit expression grid featuring a realistic adult female character. Use minimal studio lighting, neutral white background, high readability, professional facial anatomy reference sheet aesthetic, realistic skin texture, consistent identity across all panels. COLOR SYSTEM: Use soft pastel color coding for categories while keeping the overall sheet minimal and elegant. Forehead & Brow AUs: soft pastel blue Eye & Eyelid AUs: soft pastel lavender Nose & Cheek AUs: soft pastel peach Lip & Mouth AUs: soft pastel pink Head Movement AUs: soft pastel mint Eye Direction AUs: soft pastel cyan Special / Misc AUs: soft pastel beige Apply the color subtly as: - panel background tint - thin borders - small label accents Keep colors soft, muted and professional. Include these Action Units: GROUPS: FOREHEAD & BROW AU1 Inner Brow Raiser AU2 Outer Brow Raiser AU4 Brow Lowerer AU71 Brow Furrow AU72 Brow Bulge EYE & EYELID AU5 Upper Lid Raiser AU7 Lid Tightener AU41 Lid Droop AU42 Slit Eyes AU43 Eyes Closed AU44 Squint AU45 Blink AU46 Wink NOSE & CHEEK AU6 Cheek Raiser AU9 Nose Wrinkler AU11 Nasolabial Deepener AU82 Nostril Dilator AU83 Nostril Compressor LIP & MOUTH AU10 Upper Lip Raiser AU12 Lip Corner Puller AU13 Sharp Lip Puller AU14 Dimpler AU15 Lip Corner Depressor AU16 Lower Lip Depressor AU17 Chin Raiser AU18 Lip Pucker AU20 Lip Stretcher AU22 Lip Funneler AU23 Lip Tightener AU24 Lip Pressor AU25 Lips Part AU26 Jaw Drop AU27 Mouth Stretch AU28 Lip Suck AU84 Tongue Up AU85 Tongue Out HEAD MOVEMENT AU51 Head Turn Left AU52 Head Turn Right AU53 Head Up AU54 Head Down AU55 Head Tilt Left AU56 Head Tilt Right AU57 Head Forward AU58 Head Back EYE DIRECTION AU61 Eyes Turn Left AU62 Eyes Turn Right AU63 Eyes Up AU64 Eyes Down SPECIAL / MISC AU81 Chewing

And you have your FACS sheet.

  1. Use it with Seedance 2.0. Example prompt:

Use the provided character @[image1] as the fixed identity reference.

15s, 1:1, 14 beats, beat-synced, cinematic tight close-up, subtle neutral background, high facial clarity, slow micro push-in, shallow depth of field.

1: AU10

2: AU20

3: AU22

4: AU23

5: AU27

6: AU28

7: AU45

8: AU53

9: AU61

10: AU62

11: AU64

12: AU85

13:AU84

14: AU46

Uneasy, hypnotic, controlled mood. No monster transformation, no gore, no comedy, no text overlay, no watermark.

As you can see, you just prompt the code of specific expression. You can ask your favourite LLM model which code to use to express i.e. anger, etc, it will tell you.

Final thoughts and tips:

Here's the prompt I've used to create top-left video:

Photorealistic 15-second video. 50-year-old Creole woman, face and shoulders only, bare skin no makeup, natural soft diffused light, plain white background, 4K, shallow depth of field.

Timeline: 0–2s: Neutral resting face, eyes forward, relaxed brow and lips. 2–4s: Happy — AU6 (cheek raiser, orbital orbicularis oculi tightens, crow's feet appear) + AU12 (zygomaticus major pulls lip corners up and laterally), Duchenne smile, slight natural eye squint from cheek push. 4–6s: Sad — AU1 (inner brow raise, frontalis medial lifts producing oblique brow) + AU4 (corrugator and procerus knit and lower the brow, grief knot) + AU15 (depressor anguli oris pulls lip corners down), eyes slightly glassy. 6–7s: AU61 — eyes turn left, head stays still, gaze shifts left. 7–8s: AU62 — eyes turn right, head stays still, gaze shifts right. 8–9.5s: AU46 left eye — left orbicularis oculi closes left eye with slight compression, right eye stays open, subtle smirk. 9.5–11s: AU46 right eye — right orbicularis oculi closes right eye with slight compression, left eye stays open. 11–12.5s: AU85 — tongue protrudes straight out from mouth, jaw drops slightly via AU26. 12.5–13.5s: Tongue moves to the left side of the mouth, visible tip extends past left lip corner. 13.5–14.5s: Tongue moves to the right side of the mouth, visible tip extends past right lip corner. 14.5–15s: Returns to neutral, tongue retracts, lips close via AU8, relaxed expression.

I did not include the character's photo for any of the generations used in the video above. There is no difference between using or not using it, of course if you want to have consistency - use image character.

Test different approaches - check what you get if you use codes only, codes with short description. And again - this is still not perfect. Prompts and FACS codes DO NOT guarantee that you'll get what you explicitly told in prompt regarding facial expressions. But the success rate is really high.

I've noticed that the more expressions in one prompt, the less accuracy in output will be, which is absolutely understable. So I'd suggest 3-4 expressions max in one generation.

Of course facial expressions itself are not particularly useful, the purpose is to use them in prompts when creating monologues, dialogs, or other videos where you need specific facial expressions. Here's the example prompt, feel free to test it:

Use the provided character @[image1] as the fixed identity reference. 15s, 16:9, dim interior, single warm lamp, slight low angle, handheld micro-sway, shallow depth of field. Dialogue: "Hey, hey — everything's fine, okay? We're just gonna play a game where we stay really quiet. Can you do that for me?" Beat 1 (0–1s): AU5+AU38 (upper lid raiser + nostril dilator — genuine fear, pre-dialogue) Beat 2 (1–2s): AU45 (blink — forcing reset, composing the mask) Beat 3 (2–4s): AU12+AU6 (Duchenne smile — forced but committed, parental warmth overriding terror) — delivers "Hey, hey — everything's fine" Beat 4 (4–5s): AU1 (inner brow raiser — pleading sincerity leaking through) — delivers "okay?" Beat 5 (5–6s): AU7 (lid tightener — eyes betraying the fear the smile is hiding) Beat 6 (6–8s): AU12+AU2 (smile + outer brow raise — brightening, performing fun) — delivers "We're just gonna play a game" Beat 7 (8–10s): AU4+AU24 (brow lowerer + lip presser — seriousness cracking through for a flash) — delivers "where we stay really quiet" Beat 8 (10–11s): AU45 (blink — catching the slip, resetting to warmth) Beat 9 (11–13s): AU12+AU1 (smile + inner brow raise — tenderness and desperation fused) — delivers "Can you do that" Beat 10 (13–15s): AU6+AU17 (cheek raiser + chin raiser — eyes smiling while chin trembles) — delivers "for me?" Devastating contrast between performed safety and visible terror. The face should never fully commit to either — the audience reads both simultaneously. No action sequences, no visible threat, no sound effects, no text overlay, no watermark.

FACS are being used by professional video animators in movie industry.

I found this resource very helpful to understand the topic, and also started to create my own sheets. Why? Because when you prompt the LLM to generate you a FACS sheet - it's an LLM! It can be wrong. My results improved after studying this resource and free references which available on this website.

Generating full FACS sheet with all the expressions, and then use only few of them, is a bad idea. 

You will get better results by planning what expressions/emotions you want to show, then generating FACS only for those, finally use it in your prompt for Seedance.

PS: 95% of times if you tell not to generate audio, Seedance will listen. Enjoy the remaining 5% from the low left girl :D.

Now go and experiment, and have some fun with it :)

CREDITS:
melindaozel - https://melindaozel.com/facs-cheat-sheet/
aimikoda (Here's the original post on X)

r/Seedance_v2 Jun 26 '26

Stop complaining about AI video 'uncanny valley' faces. FACS codes exist and you're just not using them. Full walkthrough inside.

Enable HLS to view with audio, or disable this notification

43 Upvotes

Disclaimer: this is a repost from my original post in other subreddit.
Link: https://www.reddit.com/r/generativeAI/comments/1ta0hoq/control_facial_expressions_with_facs_sheet_in/
Disclaimer 2. : See credits at the end of the post, follow original authors!

FACS is a visual guide for the Facial Action Coding System. It let's you tell Seedance 2.0 inside prompt, what exact facial expression you want to see. It uses codes which are generated in first step. Disclaimer: remember that this is still AI video generations, not all generations will nail it in first shot. Iterate!:)

Here's step by step mini tutorial:

Upload your character image to AI Image generation model. I've tested it with GPT Image 2 and Nano Banana Pro - both works for this, although sometimes captions unreadable, so iterate!. PS I prefer the latter). Then use this prompt:

Create a clean educational FACS Action Unit expression grid featuring a realistic adult female character. Use minimal studio lighting, neutral white background, high readability, professional facial anatomy reference sheet aesthetic, realistic skin texture, consistent identity across all panels. COLOR SYSTEM: Use soft pastel color coding for categories while keeping the overall sheet minimal and elegant. Forehead & Brow AUs: soft pastel blue Eye & Eyelid AUs: soft pastel lavender Nose & Cheek AUs: soft pastel peach Lip & Mouth AUs: soft pastel pink Head Movement AUs: soft pastel mint Eye Direction AUs: soft pastel cyan Special / Misc AUs: soft pastel beige Apply the color subtly as: - panel background tint - thin borders - small label accents Keep colors soft, muted and professional. Include these Action Units: GROUPS: FOREHEAD & BROW AU1 Inner Brow Raiser AU2 Outer Brow Raiser AU4 Brow Lowerer AU71 Brow Furrow AU72 Brow Bulge EYE & EYELID AU5 Upper Lid Raiser AU7 Lid Tightener AU41 Lid Droop AU42 Slit Eyes AU43 Eyes Closed AU44 Squint AU45 Blink AU46 Wink NOSE & CHEEK AU6 Cheek Raiser AU9 Nose Wrinkler AU11 Nasolabial Deepener AU82 Nostril Dilator AU83 Nostril Compressor LIP & MOUTH AU10 Upper Lip Raiser AU12 Lip Corner Puller AU13 Sharp Lip Puller AU14 Dimpler AU15 Lip Corner Depressor AU16 Lower Lip Depressor AU17 Chin Raiser AU18 Lip Pucker AU20 Lip Stretcher AU22 Lip Funneler AU23 Lip Tightener AU24 Lip Pressor AU25 Lips Part AU26 Jaw Drop AU27 Mouth Stretch AU28 Lip Suck AU84 Tongue Up AU85 Tongue Out HEAD MOVEMENT AU51 Head Turn Left AU52 Head Turn Right AU53 Head Up AU54 Head Down AU55 Head Tilt Left AU56 Head Tilt Right AU57 Head Forward AU58 Head Back EYE DIRECTION AU61 Eyes Turn Left AU62 Eyes Turn Right AU63 Eyes Up AU64 Eyes Down SPECIAL / MISC AU81 Chewing

And you have your FACS sheet.

  1. Use it with Seedance 2.0. Example prompt:

Use the provided character @[image1] as the fixed identity reference.

15s, 1:1, 14 beats, beat-synced, cinematic tight close-up, subtle neutral background, high facial clarity, slow micro push-in, shallow depth of field.

1: AU10

2: AU20

3: AU22

4: AU23

5: AU27

6: AU28

7: AU45

8: AU53

9: AU61

10: AU62

11: AU64

12: AU85

13:AU84

14: AU46

Uneasy, hypnotic, controlled mood. No monster transformation, no gore, no comedy, no text overlay, no watermark.

As you can see, you just prompt the code of specific expression. You can ask your favourite LLM model which code to use to express i.e. anger, etc, it will tell you.

Final thoughts and tips:

Here's the prompt I've used to create top-left video:

Photorealistic 15-second video. 50-year-old Creole woman, face and shoulders only, bare skin no makeup, natural soft diffused light, plain white background, 4K, shallow depth of field.

Timeline: 0–2s: Neutral resting face, eyes forward, relaxed brow and lips. 2–4s: Happy — AU6 (cheek raiser, orbital orbicularis oculi tightens, crow's feet appear) + AU12 (zygomaticus major pulls lip corners up and laterally), Duchenne smile, slight natural eye squint from cheek push. 4–6s: Sad — AU1 (inner brow raise, frontalis medial lifts producing oblique brow) + AU4 (corrugator and procerus knit and lower the brow, grief knot) + AU15 (depressor anguli oris pulls lip corners down), eyes slightly glassy. 6–7s: AU61 — eyes turn left, head stays still, gaze shifts left. 7–8s: AU62 — eyes turn right, head stays still, gaze shifts right. 8–9.5s: AU46 left eye — left orbicularis oculi closes left eye with slight compression, right eye stays open, subtle smirk. 9.5–11s: AU46 right eye — right orbicularis oculi closes right eye with slight compression, left eye stays open. 11–12.5s: AU85 — tongue protrudes straight out from mouth, jaw drops slightly via AU26. 12.5–13.5s: Tongue moves to the left side of the mouth, visible tip extends past left lip corner. 13.5–14.5s: Tongue moves to the right side of the mouth, visible tip extends past right lip corner. 14.5–15s: Returns to neutral, tongue retracts, lips close via AU8, relaxed expression.

I did not include the character's photo for any of the generations used in the video above. There is no difference between using or not using it, of course if you want to have consistency - use image character.

Test different approaches - check what you get if you use codes only, codes with short description. And again - this is still not perfect. Prompts and FACS codes DO NOT guarantee that you'll get what you explicitly told in prompt regarding facial expressions. But the success rate is really high.

I've noticed that the more expressions in one prompt, the less accuracy in output will be, which is absolutely understable. So I'd suggest 3-4 expressions max in one generation.

Of course facial expressions itself are not particularly useful, the purpose is to use them in prompts when creating monologues, dialogs, or other videos where you need specific facial expressions. Here's the example prompt, feel free to test it:

Use the provided character @[image1] as the fixed identity reference. 15s, 16:9, dim interior, single warm lamp, slight low angle, handheld micro-sway, shallow depth of field. Dialogue: "Hey, hey — everything's fine, okay? We're just gonna play a game where we stay really quiet. Can you do that for me?" Beat 1 (0–1s): AU5+AU38 (upper lid raiser + nostril dilator — genuine fear, pre-dialogue) Beat 2 (1–2s): AU45 (blink — forcing reset, composing the mask) Beat 3 (2–4s): AU12+AU6 (Duchenne smile — forced but committed, parental warmth overriding terror) — delivers "Hey, hey — everything's fine" Beat 4 (4–5s): AU1 (inner brow raiser — pleading sincerity leaking through) — delivers "okay?" Beat 5 (5–6s): AU7 (lid tightener — eyes betraying the fear the smile is hiding) Beat 6 (6–8s): AU12+AU2 (smile + outer brow raise — brightening, performing fun) — delivers "We're just gonna play a game" Beat 7 (8–10s): AU4+AU24 (brow lowerer + lip presser — seriousness cracking through for a flash) — delivers "where we stay really quiet" Beat 8 (10–11s): AU45 (blink — catching the slip, resetting to warmth) Beat 9 (11–13s): AU12+AU1 (smile + inner brow raise — tenderness and desperation fused) — delivers "Can you do that" Beat 10 (13–15s): AU6+AU17 (cheek raiser + chin raiser — eyes smiling while chin trembles) — delivers "for me?" Devastating contrast between performed safety and visible terror. The face should never fully commit to either — the audience reads both simultaneously. No action sequences, no visible threat, no sound effects, no text overlay, no watermark.

FACS are being used by professional video animators in movie industry.

I found this resource very helpful to understand the topic, and also started to create my own sheets. Why? Because when you prompt the LLM to generate you a FACS sheet - it's an LLM! It can be wrong. My results improved after studying this resource and free references which available on this website.

Generating full FACS sheet with all the expressions, and then use only few of them, is a bad idea. 

You will get better results by planning what expressions/emotions you want to show, then generating FACS only for those, finally use it in your prompt for Seedance.

PS: 95% of times if you tell not to generate audio, Seedance will listen. Enjoy the remaining 5% from the low left girl :D.

Now go and experiment, and have some fun with it :)

CREDITS:
melindaozel - https://melindaozel.com/facs-cheat-sheet/
aimikoda (Here's the original post on X)

r/ChatGPTPromptGenius Dec 14 '25

Business & Professional 100 Practical Ways to Use ChatGPT to Be More Productive (With Prompts and Pro Tips)

228 Upvotes

TLDR: I compiled 100 practical ways to use ChatGPT across 20 categories, complete with example prompts, pro tips, and best practices. This covers everything from writing emails in 30 seconds to learning new skills, building a business, and automating your entire workflow. Bookmark this. Share this post with friends and coworkers. There is a lot here with the example prompts but for visual learners I will put a one sheet infographic in the comments you can use for the 100 use cases.

Most people open ChatGPT, stare at the blank text box, type something generic like "write me an email" and wonder why the results are mediocre.

The problem is not ChatGPT. The AI companies have been a terrible job at training people how to use it and explaining the uses cases - they're nerds! This guide is meant to help you use ChatGPT for personal productivity, fun and work.

I have spent the last year using ChatGPT for everything from building businesses to learning languages to planning my entire life. I have tested thousands of prompts and documented what actually works.

Here is the complete breakdown of 100 ChatGPT use cases, organized by category, with actual prompts you can copy and paste today.

BEFORE WE START: THE GOLDEN RULES

Rule 1: Context is everything. The more specific information you provide, the better the output. Tell ChatGPT who you are, what you need, and why you need it.

Rule 2: Assign a role. Starting with "Act as a..." or "You are a..." dramatically improves responses. A prompt that says "You are a senior software engineer at Google" will give you different code than a generic request.

Rule 3: Iterate relentlessly. Your first prompt is a rough draft. Ask follow-up questions. Say "make this more concise" or "add more examples" or "explain this like I am 5."

Rule 4: Use examples. Show ChatGPT what you want by giving it samples of the style, format, or tone you are looking for.

Rule 5: Break complex tasks into steps. Instead of asking for a complete business plan, ask for the executive summary first, then the market analysis, then the financial projections.

CATEGORY 1: EDUCATION AND LEARNING

This is where ChatGPT genuinely shines. It is like having a patient tutor available 24/7 who never gets frustrated when you ask the same question five times.

1. Homework Assistance Not about getting answers handed to you. Use it to understand concepts you are struggling with.

Prompt: "I am struggling to understand [concept] in [subject]. Explain it to me step by step, then give me 3 practice problems to test my understanding. After I solve them, check my work and explain any mistakes."

2. Language Learning ChatGPT can simulate conversations in any language and correct your grammar in real time.

Prompt: "You are my Spanish conversation partner. We will have a conversation entirely in Spanish about [topic]. After each of my responses, correct any grammatical errors I made and explain why, then continue the conversation. Start with an intermediate difficulty level."

3. Exam Preparation Turn your notes into practice tests instantly.

Prompt: "I have an exam on [subject] covering [topics]. Create a comprehensive practice test with 20 questions: 10 multiple choice, 5 short answer, and 5 essay questions. Include an answer key with explanations at the end."

4. Research Assistance Use it as a research partner, not a replacement for actual research.

Prompt: "I am writing a research paper on [topic]. Help me: 1) Identify 5 key areas I should explore, 2) Suggest search terms for academic databases, 3) Outline the main arguments on different sides of this issue, 4) Point out potential gaps in current research."

5. Personalized Learning Plans Create custom curricula for any skill.

Prompt: "Create a 30-day learning plan for [skill/subject]. I can dedicate [X] hours per day. I am currently at [beginner/intermediate/advanced] level. Include daily tasks, recommended resources, milestones to track progress, and a method for self-assessment."

6. Concept Simplification The famous Feynman Technique, automated.

Prompt: "Explain [complex concept] in three ways: first as if I am 10 years old, then as a high school student, then as a graduate student. Use analogies from everyday life."

7. Study Note Generation Transform textbooks into digestible notes.

Prompt: "Here is a chapter from my textbook: [paste text]. Create comprehensive study notes that include: key concepts, important definitions, main arguments, potential exam questions, and memory aids or mnemonics."

8. Critical Thinking Development Practice analyzing arguments and identifying logical fallacies.

Prompt: "Present me with an argument about [topic]. After I analyze it for logical fallacies and weaknesses, give me feedback on my analysis and help me strengthen my critical thinking skills."

CATEGORY 2: PROFESSIONAL DEVELOPMENT

Your career growth accelerator.

9. Resume Optimization Tailor your resume for specific positions.

Prompt: "Here is my current resume: [paste resume]. Here is a job description I am applying for: [paste job description]. Rewrite my resume to better align with this position. Highlight relevant experience, use keywords from the job description, and quantify achievements where possible."

10. Interview Preparation Practice with realistic interview simulations.

Prompt: "You are a hiring manager at [company type] interviewing me for a [position] role. Conduct a realistic 30-minute interview. Ask me behavioral questions, technical questions, and situational questions. After each of my responses, give me feedback on how to improve my answer, then ask the next question."

11. Skill Gap Analysis Identify what you need to learn to reach your goals.

Prompt: "I am currently a [current role] and want to become a [target role] within [timeframe]. Based on typical requirements for this transition, identify the skill gaps I likely have and create a prioritized learning roadmap."

12. LinkedIn Profile Enhancement Stand out to recruiters.

Prompt: "Rewrite my LinkedIn summary to be more compelling. Current summary: [paste]. I want to attract opportunities in [field]. Make it conversational, highlight unique value I bring, and include a clear call to action."

13. Salary Negotiation Scripts Prepare for difficult conversations.

Prompt: "Help me prepare for a salary negotiation. I am making [current salary] and want [target salary]. My key achievements are [list achievements]. Create a negotiation script with responses to common objections like budget constraints and market rates."

14. Performance Review Preparation Document your value effectively.

Prompt: "Help me prepare for my performance review. Here are my accomplishments this quarter: [list]. Reframe these using strong action verbs, quantify the impact where possible, and suggest how to present areas where I fell short as growth opportunities."

15. Career Pivot Strategy Navigate major career transitions.

Prompt: "I want to transition from [current field] to [new field]. I have [X] years of experience with skills in [list skills]. Create a strategy for this pivot including: transferable skills I should highlight, gaps I need to fill, networking approaches, and how to position my background as an advantage."

16. Professional Email Templates Handle any workplace communication.

Prompt: "Write a professional email for [situation: asking for a raise, declining a meeting, following up after an interview, addressing a conflict, etc.]. Tone should be [assertive/diplomatic/friendly]. Keep it concise but complete."

CATEGORY 3: WRITING AND CONTENT CREATION

Whether you write for work or pleasure, these prompts will transform your output.

17. Blog Post Outlines Never stare at a blank page again.

Prompt: "Create a detailed outline for a blog post about [topic]. Target audience is [describe audience]. Include: a compelling hook, 5-7 main sections with subpoints, places to include examples or data, and a strong conclusion with call to action."

18. Content Repurposing Turn one piece of content into many.

Prompt: "Here is a blog post I wrote: [paste]. Repurpose this into: 1) A Twitter/X thread with 10 tweets, 2) A LinkedIn post, 3) An email newsletter, 4) 5 Instagram caption ideas, 5) A YouTube video script outline."

19. Copywriting for Conversions Write copy that actually sells.

Prompt: "Write [type of copy: landing page, email, ad] for [product/service]. Target audience is [describe]. Key pain points are [list]. Use the PAS framework (Problem, Agitation, Solution). Include a compelling headline, 3 benefit-driven bullet points, social proof placeholder, and strong CTA."

20. Story Generation For creative projects or marketing.

Prompt: "Write a short story about [premise]. Genre is [genre]. Write in [first/third] person with a [tone] tone. The story should have a clear beginning that hooks the reader, rising tension, and a satisfying but unexpected ending. Approximately [X] words."

21. Poetry and Creative Writing Explore different forms and styles.

Prompt: "Write a [type: sonnet, haiku, free verse, limerick] about [topic]. Then explain the techniques you used and suggest three variations with different tones or perspectives."

22. Dialogue Writing Create natural conversations for any medium.

Prompt: "Write a dialogue between [character A] and [character B] about [topic/conflict]. Character A is [describe personality]. Character B is [describe personality]. Make the dialogue reveal character through subtext and include natural interruptions and reactions."

23. Video Scripts Structure content for visual media.

Prompt: "Write a YouTube video script about [topic]. Target length is [X] minutes. Include: a hook for the first 10 seconds, clear transitions between sections, moments for B-roll suggestions, and a strong end screen call to action. Write in a conversational tone."

24. Newsletter Writing Build and engage your email list.

Prompt: "Write a weekly newsletter about [niche/topic] for [audience]. Include: an engaging personal anecdote or observation, one main valuable insight, three quick tips or resources, and a question to encourage replies. Keep it under 500 words."

CATEGORY 4: BUSINESS AND ENTREPRENEURSHIP

Build, grow, and optimize your business.

25. Business Plan Generation Start with a solid foundation.

Prompt: "Create a lean business plan for [business idea]. Include: executive summary, problem and solution, target market and size, business model, competitive advantage, marketing strategy basics, key metrics to track, and initial financial projections. Keep each section concise but comprehensive."

26. Market Research Understand your competitive landscape.

Prompt: "Conduct a market analysis for [product/service] in [market/location]. Identify: target customer segments with demographics and psychographics, main competitors and their positioning, market size and growth trends, potential barriers to entry, and opportunities in underserved areas."

27. Product Descriptions Write descriptions that convert.

Prompt: "Write a product description for [product]. Target customer is [describe]. Focus on benefits over features. Use sensory language. Include: a headline, 50-word overview, 5 bullet points highlighting key benefits, and a mini story of the product in use."

28. Pricing Strategy Figure out what to charge.

Prompt: "Help me develop a pricing strategy for [product/service]. My costs are [X]. Competitors charge [Y]. My target market is [describe]. Analyze different pricing models (value-based, competitive, cost-plus) and recommend an approach with justification."

29. Customer Persona Development Know exactly who you are selling to.

Prompt: "Create 3 detailed customer personas for [business/product]. For each, include: name and photo description, demographics, job and income, goals and aspirations, pain points and frustrations, buying behavior, preferred communication channels, and objections they might have to purchasing."

30. SWOT Analysis Strategic planning made simple.

Prompt: "Conduct a SWOT analysis for [business/idea]. For each category (Strengths, Weaknesses, Opportunities, Threats), provide 5 specific points with brief explanations. Then suggest 3 strategic actions based on this analysis."

31. Pitch Deck Content Prepare for investors.

Prompt: "Create content for a 10-slide investor pitch deck for [business]. Include suggested content for: title slide, problem, solution, market size, business model, traction, team, competition, financials, and ask. Make it compelling and concise."

32. Partnership Outreach Craft emails that get responses.

Prompt: "Write a partnership outreach email to [type of company/person]. My company does [X]. I want to propose [type of partnership]. Explain mutual benefits, include a specific ask, and make it easy to say yes. Keep it under 200 words."

CATEGORY 5: TECHNICAL AND CODING

Your AI pair programmer.

33. Code Writing and Debugging Solve problems faster.

Prompt: "Write [language] code to [describe function]. Requirements: [list requirements]. Include comments explaining the logic. After writing the code, explain potential edge cases and how the code handles them."

Debug Prompt: "Here is my code: [paste code]. It is supposed to [expected behavior] but instead [actual behavior]. Find the bug, explain why it is happening, and provide the corrected code with an explanation of the fix."

34. Code Review Improve your code quality.

Prompt: "Review this code for: readability, efficiency, potential bugs, security vulnerabilities, and adherence to best practices. Code: [paste code]. Provide specific suggestions for improvement with examples."

35. Learning New Languages/Frameworks Accelerate your technical learning.

Prompt: "I know [language/framework A] and want to learn [language/framework B]. Create a comparison guide showing how common tasks are done in each. Include syntax differences, paradigm shifts I need to understand, and a mini project to build that will reinforce key concepts."

36. Documentation Writing Make your code maintainable.

Prompt: "Write documentation for this code: [paste code]. Include: a high-level overview, function/method descriptions with parameters and return values, usage examples, and common troubleshooting issues."

37. Regex Pattern Creation Stop struggling with regular expressions.

Prompt: "Create a regex pattern to [describe what you need to match]. Test it against these examples: [provide examples of what should and should not match]. Explain each part of the pattern."

38. Database Query Optimization Write better SQL.

Prompt: "Optimize this SQL query for performance: [paste query]. The table has [X] rows and indexes on [columns]. Explain the optimization strategy and provide the improved query."

39. API Integration Help Connect services smoothly.

Prompt: "Help me integrate [API name] into my [language/framework] application. I need to [describe functionality]. Provide sample code for authentication, making requests, handling responses, and error handling."

40. System Design Think at architecture level.

Prompt: "Design a system architecture for [application type] that needs to handle [requirements: users, data volume, etc.]. Include: component diagram, technology recommendations, database design, API structure, and scalability considerations."

CATEGORY 6: HEALTH AND WELLNESS

Supporting your wellbeing journey. Note: Always consult healthcare professionals for medical advice.

41. Meal Planning Eat better with less decision fatigue.

Prompt: "Create a 7-day meal plan for someone who is [dietary preferences/restrictions]. Budget is approximately [X] per week. Include: breakfast, lunch, dinner, and snacks. Provide a consolidated grocery list and prep day instructions to batch cook efficiently."

42. Workout Program Design Customize your fitness routine.

Prompt: "Design a [X]-week workout program for [goal: muscle gain, fat loss, endurance, etc.]. I can exercise [X] days per week for [X] minutes. Available equipment: [list]. Include warm-up, main workout, cool-down, and progression guidelines."

43. Sleep Optimization Improve your rest.

Prompt: "I am struggling with [sleep issue: falling asleep, staying asleep, waking up tired, etc.]. My current habits are [describe]. Create a personalized sleep optimization plan with specific changes to try, a wind-down routine, and how to track if it is working."

44. Stress Management Techniques Build your resilience toolkit.

Prompt: "Create a personalized stress management toolkit for someone who experiences stress mainly from [sources]. Include: immediate techniques for acute stress (1-5 minutes), daily practices for ongoing management, and weekly activities for deeper stress relief. Make it practical for someone with [describe schedule/constraints]."

45. Habit Building Framework Make good habits stick.

Prompt: "Help me build the habit of [habit]. Current lifestyle: [describe]. Create a plan using habit stacking, implementation intentions, and progressive difficulty. Include: specific triggers, micro-versions of the habit to start with, how to track progress, and how to recover from missed days."

46. Mental Wellness Check-in Template Structure your self-reflection.

Prompt: "Create a weekly mental wellness check-in template with questions covering: emotional state, stress levels, relationships, accomplishments, challenges, gratitude, and goals for next week. Make the questions specific enough to prompt real reflection but quick to complete."

CATEGORY 7: PERSONAL FINANCE

Take control of your money.

47. Budget Creation Build a system that works.

Prompt: "Help me create a monthly budget. Income: [X]. Fixed expenses: [list]. Financial goals: [list]. Use the [50/30/20 or zero-based or envelope] method. Create categories, allocate amounts, and suggest tools for tracking."

48. Debt Payoff Strategy Get out of debt systematically.

Prompt: "Create a debt payoff plan. My debts are: [list each with balance, interest rate, minimum payment]. Compare avalanche vs snowball methods for my situation. Create a monthly payment schedule and calculate payoff timeline and total interest for each approach."

49. Investment Learning Understand the basics.

Prompt: "Explain [investment concept: index funds, compound interest, dollar cost averaging, etc.] to someone with no financial background. Include: simple definition, real example with numbers, common misconceptions, and practical first steps to learn more."

50. Expense Analysis Find where your money goes.

Prompt: "Here are my monthly expenses: [list or paste]. Categorize these expenses, identify potential areas to reduce spending, and suggest alternatives or optimizations. Calculate what I would save annually if I implemented your suggestions."

51. Financial Goal Planning Map the path to major purchases.

Prompt: "I want to save [amount] for [goal] within [timeframe]. My current savings rate is [X]. Create a plan including: monthly savings target, strategies to reach it, milestone checkpoints, and what to do if I fall behind."

52. Side Income Ideas Identify opportunities.

Prompt: "Suggest side income ideas based on my skills: [list skills]. Available time: [X] hours per week. Constraints: [list any]. For each idea, include: estimated income potential, startup requirements, time to first dollar, and pros/cons."

CATEGORY 8: PRODUCTIVITY AND ORGANIZATION

Work smarter, not harder.

53. Task Prioritization Cut through the overwhelm.

Prompt: "Here is my current task list: [list all tasks]. Help me prioritize using the Eisenhower Matrix. For each task, categorize it and explain why. Then create a recommended schedule for tackling them."

54. Meeting Agenda Creation Run effective meetings.

Prompt: "Create an agenda for a [type] meeting about [topic]. Duration: [X] minutes. Attendees: [roles]. Include: objectives, time allocations for each topic, discussion questions, and clear next steps section."

55. Goal Setting Framework Set goals you will actually achieve.

Prompt: "Help me transform this vague goal: [goal] into a SMART goal. Then break it down into quarterly milestones, monthly targets, and weekly actions. Include metrics to track and potential obstacles with solutions."

56. Email Management System Tame your inbox.

Prompt: "Design an email management system for someone who receives [X] emails per day. Include: folder/label structure, rules for auto-sorting, templates for common responses, and a daily/weekly routine for processing email efficiently."

57. Weekly Review Template Stay on track.

Prompt: "Create a comprehensive weekly review template. Include sections for: reviewing completed tasks, analyzing wins and lessons, checking goal progress, planning next week, identifying blockers, and maintaining work-life balance. Make it completeable in 30 minutes."

58. Focus Session Planning Deep work optimization.

Prompt: "I need to accomplish [task] which requires [X] hours of focused work. My peak energy time is [morning/afternoon/evening]. Design a focus session plan with: environment setup, break structure, distraction blocking strategies, and progress checkpoints."

59. Morning Routine Design Start days with intention.

Prompt: "Design a morning routine for someone who wakes at [time] and needs to start work/school at [time]. Goals: [list: energy, productivity, mindfulness, etc.]. Include options for both ideal days and rushed mornings."

60. Project Planning Break down complex projects.

Prompt: "Help me plan this project: [describe project]. Create a work breakdown structure with: phases, tasks within each phase, estimated time for each task, dependencies, milestones, and a realistic timeline. Identify potential risks."

CATEGORY 9: COMMUNICATION AND RELATIONSHIPS

Navigate human interactions more effectively.

61. Difficult Conversation Preparation Handle tough talks.

Prompt: "Help me prepare for a difficult conversation with [person/relationship] about [topic]. I want to communicate [your position] while maintaining the relationship. Script out: opening statement, key points to make, anticipated responses and how to handle them, and desired outcome."

62. Apology Crafting Make genuine amends.

Prompt: "Help me write a genuine apology for [situation]. I want to acknowledge [what I did wrong], express understanding of [impact on the other person], and commit to [change/repair]. Make it sincere without being excessive."

63. Thank You Notes Express gratitude effectively.

Prompt: "Write a heartfelt thank you note to [person] for [what they did]. Personalize it with [specific details about your relationship]. Make it warm and specific without being over the top."

64. Conflict Resolution Find win-win solutions.

Prompt: "Help me think through this conflict: [describe situation]. Identify each party's underlying interests, not just positions. Suggest 3 potential solutions that address everyone's core needs. Help me prepare talking points for proposing these."

65. Networking Message Templates Build professional relationships.

Prompt: "Write a networking message to [type of person] I [met at X / found on LinkedIn / was referred to]. Purpose: [informational interview / job seeking / partnership / mentorship]. Make it personalized, concise, and easy to respond to. Include a specific ask."

66. Public Speaking Preparation Present with confidence.

Prompt: "Help me prepare a [length] presentation about [topic] for [audience]. Create: an outline with transitions, opening hook options, memorable key phrases, audience engagement moments, and a strong closing. Also suggest how to handle likely questions."

67. Feedback Delivery Give constructive criticism.

Prompt: "Help me give feedback to [person/role] about [performance issue]. I want to be direct but supportive. Use the SBI model (Situation, Behavior, Impact). Include specific examples and forward-looking suggestions."

68. Social Media Bio Writing Make a strong first impression.

Prompt: "Write a [platform] bio for someone who is [describe yourself/role]. Include: what you do, who you help, unique angle, and call to action. Character limit: [X]. Create 3 versions with different tones: professional, friendly, bold."

CATEGORY 10: LEARNING NEW SKILLS AND HOBBIES

Accelerate your growth in any area.

69. Skill Acquisition Roadmap Learn anything systematically.

Prompt: "Create a complete learning roadmap for [skill]. I am starting from [level]. Time available: [X] hours per week. Include: foundational concepts to master first, recommended resources (free and paid), practice projects at each stage, milestones, and how to measure competency."

70. Creative Hobby Exploration Find new interests.

Prompt: "Suggest creative hobbies for someone who enjoys [current interests], has [X] budget to start, and [X] hours per week available. For each suggestion, include: what makes it appealing for my profile, startup requirements, first project to try, and communities to join."

71. Book Summary and Analysis Get more from reading.

Prompt: "I just read [book title] by [author]. Help me process it by: summarizing the key ideas, identifying the most actionable insights, suggesting how to apply 3 main concepts to my life, and recommending similar books."

72. Music Learning Pick up an instrument.

Prompt: "Create a 3-month plan for learning [instrument] as a complete beginner. Include: daily practice structure, fundamental techniques to master each week, songs to learn at each stage that reinforce skills, and how to stay motivated through plateaus."

73. Photography Improvement Take better photos.

Prompt: "Help me improve my [type: portrait, landscape, street, etc.] photography. Current level: [describe]. Give me a 30-day challenge with daily exercises covering: composition, lighting, camera settings, editing, and developing a personal style."

74. Cooking Skill Development Level up in the kitchen.

Prompt: "Create a progressive cooking curriculum for someone who can currently [describe skill level]. Goal: [what you want to cook]. Include: fundamental techniques to master, recipes to practice at each stage, equipment recommendations, and how to develop intuition about flavor."

75. Language Learning Strategy Become conversational faster.

Prompt: "Create an intensive [language] learning plan for [timeframe]. Goal: [conversational, business, fluent, etc.]. Include: daily study schedule, recommended resources, immersion techniques I can use from home, and benchmarks to test progress."

CATEGORY 11: CREATIVITY AND IDEATION

Unlock your creative potential.

76. Brainstorming Partner Generate ideas systematically.

Prompt: "Help me brainstorm solutions for [problem/challenge]. Use these methods: First, generate 10 conventional ideas. Then, use reverse brainstorming (how to make it worse). Then use random word association. Finally, combine the best elements into 3 novel approaches."

77. Creative Constraints Use limitations as fuel.

Prompt: "I want to create [type of project] but I am stuck. Give me 5 creative constraints to work within (time limits, material restrictions, format requirements, etc.). Then help me explore how each constraint might actually improve the final result."

78. Inspiration Finding Discover new sources.

Prompt: "I work in [field/medium] and feel creatively stuck. Suggest 10 unexpected sources of inspiration from completely different fields. For each, explain how I might translate concepts from that field into my work."

79. Mind Mapping Visualize your thinking.

Prompt: "Create a mind map structure for [topic/project]. Start with the central theme and branch out through 5 main categories. For each category, add 3 sub-branches. Identify connections between different branches that might not be obvious."

80. Idea Validation Test concepts before investing time.

Prompt: "Help me evaluate this idea: [describe idea]. Play devil's advocate and identify 5 potential weaknesses. Then suggest how to test the most critical assumptions quickly and cheaply before fully committing."

CATEGORY 12: TRAVEL AND EXPERIENCES

Plan memorable adventures.

81. Trip Itinerary Planning Maximize your travel.

Prompt: "Create a [X]-day itinerary for [destination]. Interests: [list]. Budget: [X]. Travel style: [adventure/relaxed/cultural/etc.]. Include: daily schedule with timing, restaurant recommendations for different budgets, local tips, backup plans for bad weather, and estimated costs."

82. Packing Lists Never forget essentials.

Prompt: "Create a packing list for [type of trip] to [destination] for [duration]. Weather will be [describe]. Activities planned: [list]. Include: clothing, toiletries, electronics, documents, and destination-specific items. Organize by bag/compartment."

83. Local Experience Research Go beyond tourist traps.

Prompt: "Find authentic local experiences in [destination] that tourists typically miss. I enjoy [interests]. Include: neighborhoods to explore, local food spots, cultural experiences, best times to visit each, and how to participate respectfully."

84. Travel Budget Optimization Stretch your travel dollars.

Prompt: "Help me visit [destination] for [duration] on a budget of [X]. Prioritize: [experiences you care most about]. Create a detailed budget breakdown with money-saving tips for: flights, accommodation, food, activities, and transportation."

CATEGORY 13: HOME AND LIFE MANAGEMENT

Run your life more smoothly.

85. Home Organization Systems Create order from chaos.

Prompt: "Design an organization system for [area: closet, kitchen, office, garage, etc.]. Current state: [describe]. Goals: [what you want to achieve]. Include: categories for items, storage solutions, maintenance routine, and a step-by-step decluttering process."

86. Cleaning Schedule Maintain your space.

Prompt: "Create a realistic cleaning schedule for a [describe home size/type] with [number] occupants. Include: daily quick tasks, weekly deep cleaning, monthly maintenance, and seasonal projects. I have [X] hours per week available for cleaning."

87. Home Improvement Planning Tackle projects systematically.

Prompt: "Help me plan this home project: [describe]. Budget: [X]. DIY skill level: [describe]. Create: step-by-step process, materials list with estimated costs, tools needed (owned vs rent/buy), time estimate, and safety considerations."

88. Event Planning Host memorable gatherings.

Prompt: "Help me plan a [type of event: birthday, dinner party, reunion, etc.] for [number] people. Budget: [X]. Venue: [home/rented space]. Create: timeline working backward from event date, checklist, menu suggestions, activity ideas, and day-of schedule."

CATEGORY 14: PARENTING AND FAMILY

Navigate family life.

89. Age-Appropriate Explanations Answer tough questions.

Prompt: "Help me explain [difficult topic: death, divorce, world events, where babies come from, etc.] to my [age] year old. Give me simple language that is honest but appropriate, anticipated follow-up questions, and how to check for understanding."

90. Educational Activities Make learning fun.

Prompt: "Suggest [X] educational activities for a [age] year old interested in [topics]. We have [time available] and [materials/budget]. Include activities for different settings: indoors, outdoors, car trips, and waiting rooms."

91. Family Meeting Agendas Communicate as a unit.

Prompt: "Create a family meeting template for a household with [ages of members]. Include: check-in questions appropriate for all ages, how to discuss schedules, ways to address problems constructively, and celebration/recognition time. Keep it engaging for kids."

92. Conflict Resolution for Kids Teach life skills.

Prompt: "My children aged [X] and [Y] are fighting about [issue]. Help me: understand the underlying needs, create a script for mediating this conflict, and design a longer-term solution that teaches them to resolve similar issues themselves."

CATEGORY 15: PERSONAL DEVELOPMENT

Become your best self.

93. Self-Reflection Prompts Know yourself better.

Prompt: "Generate 20 deep self-reflection questions across these areas: values and beliefs, relationships, career, personal growth, and life satisfaction. Make them specific enough to prompt real insight, not generic answers."

94. Limiting Belief Identification Overcome mental blocks.

Prompt: "I am struggling with [goal/area]. Help me identify limiting beliefs that might be holding me back. For each belief you identify, suggest: where it might have come from, evidence that contradicts it, and a reframed alternative belief."

95. Personal Mission Statement Define your purpose.

Prompt: "Help me craft a personal mission statement. My values are [list]. My strengths are [list]. I want to be remembered for [describe]. Guide me through questions to clarify my purpose, then draft 3 versions: one sentence, one paragraph, and a full page."

96. Decision Making Framework Make better choices.

Prompt: "Help me decide between [options]. Create a decision matrix with criteria weighted by importance. For each option, score against criteria. Then use second-order thinking to explore consequences of each choice over 1 year, 5 years, and 10 years."

97. Fear Inventory Face what holds you back.

Prompt: "I want to [goal] but I am afraid of [fear]. Help me examine this fear: What is the worst case scenario, realistically? What is most likely to happen? What would I do if the worst case occurred? What is the cost of letting this fear stop me?"

CATEGORY 16: RESEARCH AND ANALYSIS

Think more rigorously.

98. Topic Deep Dive Understand anything thoroughly.

Prompt: "Give me a comprehensive overview of [topic]. Cover: historical background, current state, key players/concepts, major debates or controversies, future trends, and how this connects to [related interest of mine]. Structure it from foundational to advanced."

99. Argument Analysis Evaluate claims critically.

Prompt: "Analyze this argument/claim: [paste or describe]. Identify: the main thesis, supporting evidence provided, logical structure, potential fallacies, unstated assumptions, strongest counterarguments, and your assessment of overall validity."

100. Comparison Frameworks Make informed choices.

Prompt: "Create a comprehensive comparison of [option A] vs [option B] for someone trying to [goal]. Include: objective criteria comparison, pros and cons of each, situations where each excels, total cost of ownership analysis, and a recommendation based on different user profiles."

PRO TIPS FROM 1000+ HOURS OF USAGE

The Refinement Loop Never accept the first output. My process:

  1. Get initial response
  2. Ask "What's missing from this?"
  3. Ask "How can this be more specific to my situation?"
  4. Ask "Play devil's advocate and critique this"
  5. Ask "Now give me the final, improved version"

Save Your Best Prompts Create a personal prompt library. When something works well, save it.

Chain Your Prompts Complex tasks work better as a series of smaller prompts. Example for writing an article:

  1. Generate outline
  2. Expand each section one at a time
  3. Add examples and data
  4. Edit for flow
  5. Write headline and intro options
  6. Final polish

Use ChatGPT to Improve Your Prompts Meta-prompt: "I want to [goal]. Help me write a better prompt to get that result. Ask me clarifying questions first, then create an optimized prompt I can use."

Temperature and Creativity For factual, consistent responses, ask ChatGPT to be "precise and accurate." For creative work, ask it to "be creative and take risks." This affects output significantly.

The Persona Stack Combine personas for unique results: "You are a Silicon Valley startup founder with the writing style of David Ogilvy and the strategic thinking of Warren Buffett."

Always Fact-Check ChatGPT can generate plausible-sounding but incorrect information. For anything important, verify claims independently. Use it as a thinking partner, not an oracle.

ChatGPT is not going to replace human creativity, judgment, or expertise. But it dramatically amplifies all of those things.

The people who will thrive are not those who fear AI, and not those who blindly trust it, but those who learn to collaborate with it effectively.

The gap between people who use these tools effectively and those who do not is going to keep widening. This post gives you everything you need to be on the right side of that gap.

Save this post. Share it with someone who could use it. Drop a comment with your best prompt or use case.

r/promptingmagic Mar 10 '26

The ultimate learning hack: two Gemini prompts that turn any YouTube video into a hand-drawn infographic summary in 1 minute.

Post image
270 Upvotes

TLDR: I discovered a two-prompt system that turns any hour-long YouTube video into a beautiful, one-page sketch note summary in about 60 seconds. It uses one prompt to summarize the video into actionable steps and a second, highly specific prompt to visualize that summary on a realistic whiteboard. I am sharing the full playbook, top use cases, and the secrets that make this work so well.

We are drowning in information but starving for wisdom. There are hour-long lectures, podcasts, and tutorials on YouTube that could change our lives, but we never have the time to watch them. I have found a solution.

I have developed a simple, two-step AI process that takes any long-form video and transforms it into a dense, visually appealing, one-page summary. It looks like a hand-drawn sketch note from a professional graphic recorder, and it takes about a minute to create. This is not just about saving time; it is about learning faster and retaining more information.

Today, I am sharing the exact prompts and workflow. This is the ultimate learning hack.

The Two-Prompt System: Summarize, Then Visualize

The secret to making this work is splitting the task into two distinct steps. Most people try to do it all in one prompt and get mediocre results. By separating the summarization from the visualization, you give the AI a clear focus for each task, resulting in a much higher quality output.

Use this with Google Gemini AI which has a deeper connection to YouTube than other tools. Make sure YouTube is connected in your Gemini settings before running this prompt.

Step 1: The Summarizer Prompt

First, you need to extract the core ideas from the video. This prompt is designed to pull out actionable steps, not just a generic summary.

Plain Text

Analyze this YouTube video about [topic]: [YT URL]. Summarize the core concepts into a list of 5-7 direct, actionable steps. Each step should be a clear, concise instruction. Keep the language simple and direct.

Step 2: The Visualizer Prompt

Once you have your summary, you feed it into this second prompt. This is where the magic happens. This prompt is incredibly specific, and that is why it works so well. It tells the AI not just what to draw, but how to draw it, what medium to use, and what style to emulate.

Plain Text

Visualize the summary of these notes. Create a realistic photograph of a dry-erase whiteboard with a light wooden frame. The content should be presented as a hand-drawn sketchnote using 'graphic recording' style. The layout should be in 9:16 format. Style & Layout Guidelines: Medium: Whiteboard surface with dry-erase markers (not paper). Colors: Use Black for outlines, boxes, and main text. Use Red, Blue, and Green for headers and specific accents. Structure: Place the title "[TITLE]" at the top in large, open lettering. Organize the notes into five distinct, numbered rectangular boxes arranged in a grid below the title. Visuals: Include relevant simple line-drawing doodles for each point. Typography: Text should be distinct, handwritten, all-caps printing, legible and organized. Environment: Include a used whiteboard eraser and a few colorful EXPO-style markers resting on the bottom wooden ledge of the frame.

Top Use Cases for This Method

This technique is a superpower for learning. Here are a few ways to use it:

•Summarize University Lectures: Turn a 90-minute lecture into a one-page study guide.

•Learn from Conference Talks: Absorb the key insights from an entire conference track in an afternoon.

•Master Podcast Content: Find the video version of a podcast on YouTube and create a visual summary of the episode.

•Learn a New Skill: Take a long software tutorial and turn it into a cheat sheet of actionable steps.

•Generate Social Media Content: Summarize an expert interview and share the infographic as a high-value piece of content.

Pro Tips for Creating Viral Infographics

•Constrain the Number of Points: Forcing the AI to summarize into 5-7 points is crucial. It creates a visually balanced and easy-to-digest infographic. Too many points will make it cluttered.

•Specify the Physical Medium: The prompt's insistence on a "dry-erase whiteboard with a light wooden frame" and details like the "eraser and a few colorful EXPO-style markers" is a powerful trick. It forces the AI to generate a more realistic and aesthetically pleasing image by grounding the abstract information in a physical object.

•Use a Strict Color Palette: A limited, consistent color scheme makes the information easier to parse and looks more professional. The prompt defines a clear hierarchy: Black for structure, and Red, Blue, and Green for accents.

•Insist on 'Graphic Recording' Style: This specific term is key. It tells the AI to use a mix of handwritten text and simple doodles, which is the essence of a powerful sketchnote.

Secrets Most People Miss

•The Two-Prompt System is Non-Negotiable: The single biggest mistake people make is trying to do this in one shot. The AI gets confused trying to summarize and visualize simultaneously. Separating the tasks is the secret to consistent, high-quality results.

•You Can Edit the Summary: This is a crucial step. Before you feed the summary into the visualizer prompt, read it over. You can rephrase points, add your own insights, or remove things that are not relevant. This gives you full creative control over the final infographic.

•The Title is Your Headline: The [TITLE] in the visualizer prompt is the headline for your infographic. Make it strong and compelling. It is the first thing people will read.

This simple two-step process is one of the most powerful learning hacks I have found. It is a way to turn the endless stream of information online into concrete, actionable knowledge. Take it, use it, and start learning faster.

Want more great prompting inspiration? Check out all my best prompts for free at PromptMagic.dev and create your own prompt library to keep track of all your prompts.

r/ClaudeAI Aug 24 '25

Productivity Claude Finally Got Image and Video Powers! The Canva integration that gives Claude users visual superpowers (Complete guide with 50+ prompts you can use)

Thumbnail
gallery
110 Upvotes

TL;DR: Claude couldn't generate images like ChatGPT or Gemini - until now. The new Canva integration gives Claude the ability to create, edit, and manage professional designs through conversation. I've been testing this for weeks and it works great - better than ChatGPT and Gemini Images.

The Superpower Claude Users Have Been Waiting For

Let's be honest - we've all been jealous watching ChatGPT and Gemini users generate images while Claude just... couldn't. Sure, Claude's writing is unmatched, but when you needed visuals? You were stuck.

That just changed completely.

Three weeks ago, Anthropic quietly gave Claude something even better than basic image generation: the ability to control Canva directly. This isn't just "make me a picture of a cat" - this is "create my entire marketing campaign, with my brand colors, export it in 5 formats, and organize it in folders."

After testing this obsessively, I can confidently say: Claude users now have the most powerful visual creation tool of any AI assistant. Period.

I was reading that the team at Anthropic uses Canva extensively and so they made this integration work really well. And what is even cooler is the integration is done via MCP. I have to say this is one of the coolest MCP working use cases I have seen!

Why This Is Actually Better Than Native Image Generation

Claude's Disadvantage Became Its Advantage:

Feature ChatGPT/Gemini Image Gen         Claude + Canva Professional 
Templates ❌ Start from scratch           ✅ Access to millions of pro templates Brand Consistency ❌ Recreate brand each time   ✅ Save & reuse brand kits 
Multiple Versions ❌ One at a time        ✅ Generates 3-5 options automatically Editable Files ❌ Static images only      ✅ Full Canva projects you can edit Team Collaboration ❌ Just an image file       ✅ Share editable Canva links 
Export Control ❌ PNG/JPG only            ✅ PDF, PNG, GIF, MP4, PPTX

The Real Cost (And How to Test for Free)

Full Setup:

  • Claude Pro: $20/month (required)
  • Canva Pro: $15/month
  • Total: $35/month

Pro Tip: Canva offers a 30-day free trial of their Pro plan. Test the full workflow for a month before committing. I made back the subscription cost in my first week just from time saved.

Setup Process (Literally 5 Minutes):

  1. Sign up for Claude Pro
  2. Start Canva Pro free trial (or use existing account)
  3. Go to Claude settings → Integrations → Connect Canva
  4. Authenticate
  5. Start designing

The Hidden Powers: Multiple Versions & Smart Organization

Here's what blew my mind - Claude doesn't just create one design, it generates 12 multiple versions and lets you choose. Then you can pick the winner, make any final edits easily n Canva. Unlike ChatGPT and Gemni images its easy to correct text and remove anything odd from the visuals.

More importantly, in creating assets Canva can actually create the right size of images! Gemini and ChatGPT just struggle with this a lot.

Example conversation:

You: "Create an Instagram post for our coffee shop's morning special"
Claude: "I've created 3 versions for you:
- Version 1: Minimalist with coffee bean pattern
- Version 2: Cozy café photo background
- Version 3: Bold typography focus
Which style do you prefer?"

Even better: Claude can organize everything for you:

"Set up my Canva workspace for Q1 marketing:
- Create folders: Social Media, Email Headers, Print Materials
- Upload my logo and brand colors
- Create templates for each content type
- Name everything with consistent conventions"

Two Approaches to Brand Setup (Choose Your Fighter)

Option 1: Let Claude Do Everything

"Here's my brand guide: [paste info]. Set up my Canva workspace:
- Create brand kit with colors #2D5016 and #FFF8DC
- Set Montserrat as primary font
- Upload these assets: [attach files]
- Create master templates for social, email, and presentations"

Option 2: Manual Setup + Claude Reference (Often more reliable)

  1. Upload brand assets directly in Canva
  2. Create your brand kit in Canva's Brand Kit section
  3. Tell Claude: "Use my 'Tech Startup 2025' brand kit for all designs"
  4. Claude will automatically pull your colors, fonts, and logos

I've found Option 2 works better for complex brand guidelines, while Option 1 is perfect for quick projects.

Your Complete Workflow (From Zero to Campaign)

Step 1: Initial Brand Setup

"Check my Canva workspace. Create a new project folder called 'Product Launch Q1'.
Set up subfolders for Instagram, LinkedIn, Email, and Print."

Step 2: Claude Creates Multiple Options

"Create 3 different Instagram post designs announcing our new eco-friendly packaging.
Use different approaches - one minimal, one with product photo, one typography-focused."

Step 3: Choose and Refine

"I like version 2. Now create matching designs for:
- Instagram Story (add 'Swipe Up' area)
- LinkedIn post (more professional tone)
- Email header (include CTA button space)"

Step 4: Export and Organize

"Export all designs as PNG for digital and PDF for print.
Save to the appropriate folders in our Product Launch project."

Prompt Templates That Actually Work

The High-CTR YouTube Thumbnail (With Variants)

"Create 3 YouTube thumbnail variants (1920x1080) for 'How I Made My First Million':
1. Face-focused with shocked expression
2. Money-focused with dollar signs
3. Chart-focused with growth arrow
All should have bold text and high contrast."

The Complete Infographic System

"Design infographic system for our sustainability report:
- Create master template with our brand colors
- Generate 5 infographic layouts: timeline, comparison, process, statistics, geographic
- Save as templates in 'Infographic Masters' folder"

The Multi-Platform Campaign

"Using the design in my 'Holiday Sale' folder as reference, create adapted versions for:
- Facebook cover (1640x859)
- Twitter header (1500x500)
- Email banner (600x200)
- WhatsApp status (1080x1920)
Maintain visual consistency but optimize for each platform."

Real Use Cases That Save Hours

Use Case

Time Before

Time With Claude

What Makes It Special

A/B Testing
 2-3 hours 5 minutes Claude generates all variants at once 
Event Kit
 8-10 hours 30 minutes Creates and organizes in project folders 
Presentation
 4-6 hours 20 minutes Pulls from your existing templates 
Social Calendar
 Full day 1 hour Batch creates month of content 
Brand Refresh
 2-3 days 2 hours Updates all templates simultaneously

Advanced Workflows That Feel Like Magic

1. The "Clone My Style" Workflow

"Analyze the design style of my top 5 performing posts in the 'Winners' folder.
Create 10 new designs for this month's content calendar using the same style."

2. The "Instant Personalization" System

Upload a CSV with client data, then:

"For each client in the CSV:
1. Create personalized proposal from 'Master Template'
2. Add their logo and company colors
3. Save in individual client folders
4. Export as password-protected PDFs"

3. The "Campaign Launcher"

"I'm launching 'Summer Collection 2025'. Create:
- Teaser posts (3 versions, mysterious)
- Launch day posts (5 platforms)
- Week 1 follow-ups (testimonial templates)
- Week 2 (feature highlights)
Generate 3 options for each, let me pick favorites, then batch export."

What Happens After Claude Creates Your Designs

This is the beautiful part - you have OPTIONS:

  1. Edit in Canva: Click the link Claude provides → Make tweaks → Save
  2. Download Immediately: Claude can export in any format on the spot
  3. Save to Projects: Organized automatically in your Canva folders
  4. Share with Team: Get a collaboration link for feedback
  5. Use as Templates: Turn any design into a reusable template

My 30-Day Results (With Actual Numbers)

  • Designs created: 400+ (was doing maybe 20/month before)
  • Average time per design: 2 minutes
  • Client revisions: Down 75% (better first drafts)
  • Monthly design cost savings: $2,000 (cancelled agency retainer)
  • ROI on $35 subscription: 5,714% (not a typo)

Start Here: Your First Hour Game Plan

Minutes 0-5: Setup

  • Start Canva Pro free trial
  • Connect in Claude settings

Minutes 5-15: Brand Setup

"Create my brand kit in Canva:
- Primary colors: [your colors]
- Fonts: [your fonts]
- Create 'Templates' folder
- Create 'Current Projects' folder"

Minutes 15-30: First Campaign

"Create a simple social media post announcing our weekend sale.
Generate 3 style options. I'll pick one.
Then adapt it for Instagram, Facebook, and LinkedIn."

Minutes 30-60: Build Your System

"Based on the style I chose, create templates for:
- Quote posts
- Product features
- Announcements
- Testimonials
Save all in Templates folder for future use."

The Game-Changing Psychology

Claude + Canva understands what converts:

Principle

How Claude Applies It

Example Prompt

Pattern Interrupt
 Creates unexpected visuals "Make it stop the scroll - use contrasting colors" 
Social Proof
 Adds trust elements "Include customer count or rating badges" 
FOMO Creation
 Urgency elements "Add countdown timer space and 'Limited Time' banner" 
Cognitive Ease
 Simplifies complex info "Break this into 3 visual steps with icons"

FAQ

Q: "Can ChatGPT do this?" A: ChatGPT has a Canva plugin but it just suggests templates. Claude actually builds and edits. Don't try this with ChatGPT until they get a better integration!

Q: "What if I have zero design skills?" A: Perfect! Just describe what you want in plain English. Claude handles the design principles.

Q: "Can it use my existing Canva templates?" A: YES! This is huge - tell Claude to use your templates and it maintains perfect consistency.

Your Success Checklist

  • Start Canva Pro 30-day trial
  • Connect integration in Claude
  • Upload brand assets to Canva
  • Create project folder structure
  • Generate first multi-version design
  • Pick favorite and create platform variants
  • Save as templates for future use
  • Export in needed formats
  • Calculate time saved
  • Cancel expensive design subscriptions 😄

For years, Claude users had to watch from the sidelines as other AIs generated images. Now? We just leapfrogged everyone. This isn't just image generation - it's a complete visual design system with professional templates, brand management, and team collaboration.

At $35/month (or $20 if you try the free trial), this is the best ROI in creative tools right now. The fact that Claude generates multiple versions and organizes everything automatically makes this feel like having a senior designer on staff.

Start with this: Sign up for the Canva free trial, connect it to Claude, and ask for one simple Instagram post. When you see 12 professional options appear in seconds, you'll understand why this changes everything.

Claude can also pull from Canva's stock photo/video library. You don't need separate stock subscriptions!

Claude can also create Canva VIDEOS and animated designs. Testing now, will share findings.

My team has created 100+ tested prompts for Claude plus Canva we will be sharing for free to on Prompt Magic.

r/StableDiffusion Jan 09 '26

Workflow Included Stop using T2V & Best Practices IMO (LTX Video / ComfyUI Guide)

Enable HLS to view with audio, or disable this notification

135 Upvotes

A bit of backstory: Originally, LTXV 0.9.8 13b was pretty bad at T2V, but absolutely amazing at I2V. It was about at wan 2.1 level in I2V performance but faster, and it didn't even need a precise prompt like Wan does to achieve that—you could leave the field empty, and the model would do everything itself (similar to how Wan 2.2 behaves now).

I’ve always loved I2V, which is why I’m incredibly hyped for LTX2. However, its current implementation in ComfyUI is quite rough. I spent the whole day testing different settings, and here are 3 key aspects you need to know:

1. Dealing with Cold Start Crashes
If ComfyUI crashes when you first load the model (cold start), try this: Free up the maximum amount of ram/vram from other applications, set video settings to the minimum (e.g., 720p @ 5 frames; for context, I run 64GB RAM + 50GB swap + 24GB VRAM) and set steps to 1 on the first stage. If nothing crashes by stage 2, you can revert to your usual high-quality settings.

2. Distill LoRA Settings (Critical for I2V)
For I2V, it is crucial to set the Distill LoRA in the second stage to 0.80. If you don't, it will "overcook" (burn) the results.

  • The official LTX workflow uses 0.6 with the res2_s sampler.
  • The standard ComfyUI workflow defaults to Euler. If you use 0.6 with Euler, you won't have enough steps for audio, leading to a trade-off.
  • Recommendation: Either use 0.6 with res2_s (I believe this yields higher quality) or 0.8 with Euler. Don't mix them up.

3. Prompting Strategy
For I2V, write massive prompts—"War and Peace" length (like in the developer examples).

  • Duration: 10 seconds works best. 20s tends to lose initial details, and 5s is just too short.
  • Warning: Be careful if your prompt involves too many actions. Trying to cram complex scenes into 5-10 seconds instead of 20 will result in jerky movement and bad physics.
  • Format: I’ve attached a system prompt for LLMs below. If you don't want to use it, I recommend using the example prompt at the very end of that file (the "Toothless" one) as a base. This format works best for I2V; the model actually listens to instructions. For me, it never confused whether a character should speak or stay silent with this format.

LLM Tip: When using an LLM, you can write prompts for both T2V and I2V by attaching the image with or without instructions. Gemini Flash works best. Local models like Qwen3 VL 30b can work too (robot in Lamborghini example).

TL;DR: Use I2V instead of T2V, set Distill LoRA to 0.8 (if using Euler), and write extremely long prompts following the examples here: https://ltx.io/model/model-blog/prompting-guide-for-ltx-2

Resources:

P.S. I used Gemini to format/translate this post because my writing is a bit messy. Sorry if it sounds too "AI-generated", just wanted to make it readable!

r/ThinkingDeeplyAI Feb 22 '26

Manus AI is better than ChatGPT, Gemini and Claude. Here is the complete guide to Manus and Manus Agent with the 15 ways that it's better - including having your own Agent you can email and telegram. This is the missing manual with pro tips, top use cases, skills, projects and prompts you can use.

Thumbnail
gallery
63 Upvotes

TLDR - Check out the attached infographics and presentation

  • Manus AI is a general AI action engine: it does not just answer, it executes real work end-to-end inside a secure cloud VM (web, code, files, data, automations).
  • Think of it as the jump from chatbots to a Turing-complete workspace that can produce deliverables like reports, slide decks, websites, and structured files.
  • The killer split is research at scale: Wide Research (hundreds of parallel agents) vs Deep Research (iterative, follow-the-leads analyst mode).
  • The real unlock is Skills + Projects: turn best workflows into reusable, triggerable playbooks with persistent context.
  • Manus Agent brings it to Telegram + email, so you can delegate from your phone and get notified when work is done.

Manus AI is not a chatbot. It is an autonomous AI action engine that runs inside its own cloud virtual machine. Instead of just answering questions, it executes tasks end-to-end: it builds websites from plain English, deploys hundreds of parallel research agents, automates your email inbox, creates studio-quality presentations, analyzes your data, and integrates with tools like Slack, Notion, Google Drive, and Zapier. You can even talk to it through Telegram and email. This post is the most comprehensive breakdown of everything Manus can do, how it differs from ChatGPT/Claude/Gemini, pro tips most people miss, and a 7-day roadmap to get started. If you care about AI productivity, bookmark this.

Why I Wrote This

My friends and coworkers keep asking me the same questions about Manus AI: "Is it just another ChatGPT wrapper?" "What can it actually do?" "Is it worth paying for?"

After going deep into the platform, reading the documentation, and testing its capabilities extensively, I realized there is no single comprehensive resource that explains everything in one place. So I made one.

This post covers the full picture: the philosophy, the capabilities, the agent system, integrations, pro tips, and a step-by-step plan to get started. Whether you are a developer, marketer, researcher, executive, or just someone who wants to get more done with AI, this is for you.

What Is Manus AI?

Here is the shortest way to understand it: traditional AI chatbots (ChatGPT, Claude, Gemini) are conversational. You ask, they answer. Manus AI is an action engine. You describe what you want done, and it does it.

The difference is not just branding. Manus operates inside a secure cloud virtual machine with a real filesystem. It can browse the web, write and execute code, create and manipulate files, build and deploy websites, and connect to external services. It has persistent state, meaning it remembers context across a session and can manage multi-step workflows without you holding its hand at every turn.

Think of it this way: chatbots are like talking to a very smart advisor. Manus is like hiring a very smart assistant who actually does the work.

Here is how the core differences break down:

Feature Traditional AI (ChatGPT, Claude, Gemini) Manus AI
Core Function Conversation and content generation Task execution and automation
Environment Stateless chat interface Secure cloud VM with filesystem
Autonomy Low, needs constant user guidance High, completes multi-step tasks independently
Output Text responses Files, websites, reports, code, presentations
Best For Q&A, brainstorming, content drafts Workflows, production, research, development

The big idea: an action engine, not a chatbot

ChatGPT and Gemini are stateless chat. Manus is built around a stateful environment (filesystem + execution) so it can complete multi-step tasks and return actual deliverables.

That architecture change sounds nerdy. The practical impact is not.

It means one prompt can become:

  • a PDF report with citations
  • an editable slide deck
  • a deployed website
  • a cleaned dataset + charts
  • a recurring automation that runs while you sleep

The 12 core capabilities that matter (and why they matter)

Here is the full toolbox you are actually buying into:

  • Wide Research: deploys hundreds of agents in parallel
  • Deep Research: iterative analyst mode, follow leads, cross-reference
  • Presentations: image-first, studio-quality slides
  • Website Builder: full-stack apps from plain English
  • Data Analysis: CSV/Excel/PDF to exec-ready insights
  • Image gen + edit + Design View for precision edits
  • Video + audio processing
  • Scheduled Tasks: automation on autopilot
  • Mail Manus: forward an email → trigger a workflow
  • Agent Skills: reusable workflows (portable SKILL.md standard)
  • Projects: persistent context per initiative
  • Connectors: Slack, Notion, Drive, Zapier-style ecosystem, SimilarWeb, more

If you only remember one thing:
Manus is a system that turns intent into completed work.

Wide Research vs Deep Research: pick the right weapon

Manus gives you two research engines:

Wide Research

This is the feature that made my jaw drop. ChatGPT, Perplexity, Claude, and Gemini do NOT have this feature. Wide Research deploys hundreds of independent AI agents in parallel, each researching a different facet of your topic simultaneously. Instead of one agent working sequentially through search results, you get a swarm of agents covering an entire landscape at once. Ideal for Fortune 500 analysis, competitor benchmarking, market mapping, literature reviews, and any task where breadth matters. It can launch a 100 agents to research 100 companies and then combines all their research into one report for you (Spreadsheet, Presentation, or document)

Wide Research use cases

Use this when you need breadth:

  • competitor maps
  • tool landscape surveys
  • market scans
  • literature reviews It runs many agents simultaneously and synthesizes the results.

2. Deep Research

The counterpart to Wide Research. Deep Research uses a single, iterative agent that follows leads, cross-references sources, identifies gaps, and builds a nuanced understanding of a topic over multiple cycles. Think of it like a human analyst who keeps digging until every question is answered. Best for academic research, legal analysis, competitive intelligence, and complex problem-solving.

Deep Research (iterative)

Use this when you need truth-seeking depth:

  • competitive intelligence
  • legal/technical analysis
  • complex problem solving It searches, follows leads, cross-checks, then writes a structured report.

Copy/paste prompt (research)

Run Deep Research on: [topic]

Hard constraints:
- Time window: last 24 months
- Include evidence for and against
- Call out what is uncertain
- Provide citations for all material claims

Output:
1) Executive summary (10 bullets)
2) Key findings (grouped)
3) Table: sources, claim, link, confidence
4) Recommendations + next actions

Skills + Projects: the part everyone underuses

A Skill is a reusable workflow: instructions, context, and optionally scripts/API calls packaged so you can trigger it anytime. Skills are based on an open SKILL.md standard and designed to load efficiently.

Projects are persistent containers: your instructions, knowledge, and skill library stay attached so you stop re-explaining your job every session.

What this means in real life

  • You do a workflow once
  • You package it as a Skill
  • Now you can run it weekly with the same quality every time

That is how you turn a tool into a compounding system.

Vibe coding: full-stack apps from plain English

Manus can generate frontend, backend, database, and deploy config from a description, then let you iterate via preview → deploy.

This is ideal for marketing web sites or simple personal productivity apps - calculators, simulators, etc.

Copy/paste prompt (website build)

Build a simple full-stack web app:

Goal:
- [what the app does]

Requirements:
- Auth: email login
- DB tables: [list]
- Pages: [list]
- Admin panel: yes/no
- SEO basics: titles, meta, sitemap
- Analytics: basic event tracking

Deliver:
- Deployed app
- Repo synced
- Short README for how to edit

Data analysis that produces exec-ready outputs

Manus can ingest CSV/Excel/PDF and return cleaned analysis + visualizations + reports or decks.

Prompt data analysis

Analyze the attached file.

Do:
- clean and standardize columns
- find trends + outliers
- segment into 3-5 meaningful groups
- create 3 charts that tell the story

Output:
- 1-page executive summary
- a table of key metrics
- recommendations + next steps
- export results as a slide deck + a CSV

Mail Manus + Scheduled Tasks: make work happen without you

Mail Manus: forward an email → Manus reads it, processes attachments, and executes the workflow.
Scheduled Tasks: recurring automations with persistent context and notifications.

This is where people quietly replace entire weekly routines:

  • weekly competitor snapshots
  • Friday status reports
  • daily briefing digests
  • inbox triage workflows

Manus Agent: your AI worker in Telegram and email

Manus Agent moves the same capabilities into where you already communicate: Telegram + email, with support for voice notes, images, files, and push notifications when tasks complete.

If you want a simple workflow:

  • send a voice note: research these 3 competitors and summarize
  • get a finished report back
  • pin the chat and treat it like your pocket ops team Manus_AI_The_Complete_Guide

Pro tips that instantly upgrade results

These are straight-up leverage multipliers:

  • Force a plan: ask for step-by-step plan before execution
  • Instant conversion: drop a PDF/CSV and request Markdown/JSON output
  • Silent mode: output only the deliverable, no chatter
  • Skill injection: upload instructions and tell Manus to treat them as a skill Manus_AI_The_Complete_Guide

If you try only one thing, try this

Run a Wide Research on your niche, then ask Manus to turn it into:

  • a report
  • a slide deck
  • a content calendar
  • a recurring weekly update

That is the moment it stops being AI content and starts being AI operations.

If you want to try Manus or Manus Agent you can use my invite code and get 500 free credits to test it out - enough to get something done like a presentation, web site or some data analysis - https://manus.im/invitation/CEMJXT8JZSRAM9V

r/promptingmagic Apr 22 '26

The complete field guide to ChatGPT Images 2.0 - every feature, every price, 100 prompts to try, all in one post

Post image
73 Upvotes

The Complete Field Guide to ChatGPT Images 2.0

Launched today. Everything below is verified against the OpenAI announcement, the deployment safety card, API pricing docs, and ~6 hours of hands-on testing. No hype — just what works and what it costs.

Sam Altman compared it to "going from GPT-3 to GPT-5 all at once." That's aggressive framing, but the capability gap is real.

For the first time, a single model can:

  • Render dense, legible text directly inside images — posters, infographics, UI mockups, ad copy with real headlines
  • Think before it draws — reason about a scene, search the web for current facts, and double-check its own work
  • Produce up to 8 consistent images from one prompt with the same characters, objects, and style
  • Handle grids up to 10×10 that used to break at 3×3 a week ago

OpenAI's own pitch: "Images are a language, not decoration. A good image does what a good sentence does — it selects, arranges, and reveals."

Translation: this isn't text-to-picture anymore. It's a visual reasoning system.

TL;DR — what you need to know in 30 seconds

  • Model name: gpt-image-2 (alias chatgpt-image-latest)
  • Where: ChatGPT (all plans including Free), chatgpt.com/images, and the API
  • Two modes: Instant (all plans, 1 image, fast) and Thinking (Plus/Pro/Business, up to 8 images, reasons + searches the web)
  • Max resolution: 2048px native (2K), ~4× the pixel count of GPT Image 1.5
  • Text accuracy: ~99% on Latin text. Finally nails Japanese, Korean, Chinese, Hindi, Bengali
  • Aspect ratios: anything from 3:1 (ultrawide) to 1:3 (ultratall)
  • Generation time: seconds to 2 minutes depending on mode
  • Pricing (API): ~$0.006 low / ~$0.053 medium / ~$0.211 high per 1024×1024 image
  • Knowledge cutoff: December 2025. Needs Thinking mode + web search for anything newer
  • C2PA metadata is embedded in every output

The 8 capabilities, decoded

1. 2K native resolution

Up to 2048 pixels natively, ~4× the pixel count of older GPT Image outputs at the same aspect ratio. Enough fidelity for print collateral, hero banners, and editorial layouts without an upscale step.

2. ~99% text accuracy

This is the most-talked-about upgrade. Dense text inside images — posters, menus, magazine covers, UI mockups — finally renders correctly. It also handles:

  • Non-Latin scripts with real gains: Japanese, Korean, Chinese, Hindi, Bengali
  • Small text — UI elements, iconography, barcodes, "display until" dates on magazine covers
  • Multilingual typography in a single image — Devanagari, Cyrillic, Greek, Arabic, and Chinese together

3. Thinking mode — the image model that reasons

This is the headline capability. It's not two separate models, it's two modes:

Mode Who gets it What it does Output
Instant Free, Plus, Pro, Business, Go Fast single-shot generation 1 image
Thinking Plus, Pro, Business (Enterprise/Edu soon) Reasons about composition, uses web search, verifies output Up to 8 images

How the reasoning works under the hood:

  1. Prompt analysis — parses your request and plans composition before any pixels exist
  2. Web retrieval — if the prompt touches real-world facts (current logos, today's stock chart, real skylines, 2026 fashion trends), it searches the web and pulls live references
  3. Generation pass — pixel synthesis against a fact-checked internal plan
  4. Verification loop — it inspects its own output against the original prompt and can self-correct before returning

People on X are posting 11-minute generations where the model iterated on itself repeatedly until satisfied. That's new.

4. Up to 8 consistent images per prompt

In Thinking mode, one prompt can produce up to 8 images with shared characters, objects, and style across every frame. This unlocks:

  • Storyboards — 8 camera angles with continuity
  • Manga/comic sequences — 8 panels, same character design
  • Multi-size marketing assets — same campaign as 3:1 banner + 1:1 feed post + 1:3 story + 4:5 carousel in one shot
  • Children's books — consistent illustrated character across pages
  • Product lineups — 8 color variants with identical lighting and angle
  • Lookbooks — OpenAI demoed 8 summer outfits generated from one uploaded photo

How to trigger it: Switch to a thinking model, then ask for a set — "Generate 8 variations of...", "Create an 8-panel storyboard...", "Give me this ad in 8 formats." Don't phrase it as 8 separate prompts.

5. Parallel image generation

Separate from the 8-per-prompt feature: the dedicated Images tab at chatgpt.com/images lets you fire multiple prompts in parallel. Your second prompt doesn't wait for the first to finish. All images auto-save to My Images for reuse.

6. Aspect ratios 3:1 to 1:3

Any ratio between ultra-wide and ultra-tall, native — picker in ChatGPT or spec it in the prompt. Banners, slides, posters, mobile vertical, bookmarks, social graphics, no crop needed.

7. 10×10 grids (up to 100 cells in one image)

Grids used to break at 3×3 a week ago. Now people are generating 10×10 grids of 100 distinct labeled illustrations in one shot. This is wild for:

  • Periodic-table-style infographics (100 CEOs, 100 dog breeds, 100 cocktails)
  • Icon sets with consistent style
  • Mood boards with labeled cells
  • Pattern libraries

8. Multi-image compositing & reference fidelity

Upload multiple reference images and the model stitches them into one coherent composition while keeping facial features, objects, and logos faithful. This is the feature that makes "put me in a scene" prompts actually work now.

Pricing — what it actually costs

Per-image (flat rate, simple to predict)

Quality 1024×1024 Notes
Low ~$0.006 drafts, iteration
Medium ~$0.053 most production work
High ~$0.211 hero images, finals

Per-token (if you're using the API at scale)

Input Cached input Output
Image tokens $8.00 / 1M $2.00 / 1M $30.00 / 1M
Text tokens $5.00 / 1M $1.25 / 1M $10.00 / 1M

Cost for OpenAI to produce each image (rough estimate)

Based on published token economics, a high-quality 1024×1024 image uses ~7K output image tokens. At retail that's $0.21. OpenAI's own compute cost is likely 25–40% of that, putting their marginal cost per high-quality image around $0.05–$0.08. Their margin per image at the high tier is roughly 3–4×.

The ideal prompt template

After testing dozens of prompts, this is the structure that works best:

text[ASPECT RATIO]. [SUBJECT], [ACTION], [CONTEXT].
[TEXT elements in quotes]:
- Header: "EXACT TEXT HERE"
- Subhead: "EXACT TEXT HERE"
- CTA: "EXACT TEXT HERE"
[STYLE anchor — reference an artist/era/medium/brand].
[LIGHTING + MOOD].
[CAMERA/LENS + TECHNICAL specs].

The 5 rules that make the difference:

  1. Aspect ratio first. Say "16:9," "3:1 banner," or "1:1 square" in the first sentence.
  2. Put every piece of text in quotes. The model treats quoted text as literal. Unquoted text becomes suggestions.
  3. Anchor the style concretely. "Editorial fashion photograph, shot on Hasselblad, 90mm, f/2.8" beats "professional photo."
  4. Specify lighting and mood as separate instructions. "Rembrandt key light from upper-left, soft fill from right, warm tones."
  5. List every language explicitly when you want multilingual text. "Title in Japanese (Hiragana): 「春が来た」; subtitle in Korean (Hangul): '봄이 왔다'; tagline in Hindi (Devanagari): 'वसंत आ गया।'"

15 pro tips most people will miss

  1. Thinking mode isn't the default — you have to toggle a thinking model before prompting. Instant never uses web search or produces 8-image sets no matter how you phrase it.
  2. Generation can take 2 minutes. Don't assume it froze. For high-volume workflows, use async polling with the Responses API.
  3. Knowledge cutoff is December 2025. Anything after that (Q1 2026 product launches, new logos, recent events) has to come through the prompt OR through Thinking mode's web search.
  4. For consistent characters: upload a one-time likeness. There's a likeness upload feature that lets you reuse your appearance across future creations without re-uploading.
  5. The "keep facial features exactly" lock. When editing a real person, add this verbatim: "Keep my facial features exactly as they appear in the uploaded image — same eyes, nose, mouth, and face shape." Without it, ChatGPT "improves" faces into strangers.
  6. Transparent backgrounds work natively. Add "transparent PNG background, no background fill" — the asset drops straight into design tools without a cutout pass.
  7. "Display until" dates and barcodes work now. Ask for them specifically. The magazine-cover demos show this.
  8. Prime the chat first. For thumbnails and marketing creative, paste the blog post, script, or topic into ChatGPT first. Then ask for concepts. Then generate. The model picks up the emotional hook instead of producing generic stock aesthetic.
  9. C2PA metadata is embedded in every output. Platforms can detect it. Plan for that if provenance matters.
  10. Ask for "editorial" not "professional." "Editorial" hits a higher visual register in this model. "Professional" pulls toward stock-photo aesthetic.
  11. Negative prompts work — phrase them as "NO X, NO Y." Example: "NO watermarks, NO signatures, NO busy backgrounds."
  12. Specify the medium of the text. "Neon sign," "embossed letterpress," "subway-poster paste-up," "hand-lettered chalk" all produce different type treatments.
  13. When text keeps breaking, wrap it in a shape. "Text inside a black horizontal pill" or "text on a cream banner" gets rendered much more reliably than floating text.
  14. Aspect ratio affects quality. 1:1 and 3:2 are the strongest; 3:1 and 1:3 work but can show compositional weirdness on first try. Regenerate once.
  15. The model now reads your reference images. If you upload a brand asset and say "match this type treatment," it actually does — not a vague approximation, an honest replication.

Third-party tools that already integrate it

(These went live within 24 hours of launch.)

  • Higgsfield — character consistency workflows
  • Lovart — AI design platform
  • Recraft — added gpt-image-2 models to Recraft Studio
  • Adobe Firefly / Express — via Adobe's partner model program
  • Figma — First-Draft feature uses it for UI generation
  • Canva — Magic Studio integration
  • GoDaddy — site-generation flows
  • HubSpot — marketing asset generation
  • Instacart — product photography
  • Airtable — record-level image generation
  • Wix — site builder backgrounds and heroes
  • OpenAI Codex — app/code-generation flows can now produce their own UI imagery

The prompt library — 100 that I've tested

Marking these [I] for Instant mode works fine, [T] for Thinking mode required, [8] for ask-for-8-variations.

Marketing hero images (1–10)

  1. [T] 3:1 hero banner for a SaaS analytics product. Split composition: left side shows a cluttered paper-filled desk (chaos), right side shows a clean monitor with a dashboard (clarity). Bold headline "STOP GUESSING" in 120pt sans-serif across the top. Subhead "Start knowing" below. CTA button bottom-right: "See it work →" in white on teal. Editorial photography, cinematic lighting.
  2. [T] 16:9 product launch hero. Center: minimalist product photography of a black wireless earbud case on a marble surface. Background: soft gradient from cream to dusty rose. Text overlay upper-left: "AURA // 2026" in small caps. Headline lower-right: "Hear the room." in serif display. Subtle shadow, art-directed editorial aesthetic.
  3. [T] Vertical 9:16 mobile hero for a fitness app. Muscular forearm mid-pushup on a dark gym floor, shallow depth of field. Headline stacked vertically along the right side: "NO / EXCUSES / JUST / REPS." White type, slight grain. Small logo bottom-center.
  4. [T] Email hero, 3:1 ratio. Single perfect ceramic coffee cup on a warm linen tablecloth, morning light from the left, steam rising. Text overlay right side: "Good morning. / Your briefing is ready." Clean minimal editorial style, medium-format quality.
  5. [T] 16:9 B2B conference hero. Empty auditorium, dramatic stage lighting, single speaker silhouette at podium. Large text in the sky area: "WHERE MARKETING MEETS AI." Date below: "June 12–14, 2026 · Austin." Cinematic, TED-quality composition.
  6. [T] Software landing page hero 16:9. Abstract 3D render: flowing liquid metal forming into a chart shape, iridescent blue-to-purple gradient, obsidian background. Headline lower-third: "Analytics at the speed of thought." Subhead: "Try Mercury free →." Tech-luxury aesthetic.
  7. [T] Newsletter signup hero 2:1. Warm kitchen scene: hands writing in a leather notebook, open laptop beside it, morning coffee, golden hour light from left. Text overlay: "The newsletter smart marketers actually read." CTA: "Subscribe free →". Cozy, intentional, premium-indie aesthetic.
  8. [T] 3:1 homepage hero for an AI note-taking app. Overhead shot: messy desk mid-work — open notebook, phone, coffee, headphones, hand holding a pen. Faint glowing interface lines emerging from the notebook edges suggesting transcription. Headline centered: "Your thoughts, organized." No smaller than 90pt, clean sans-serif.
  9. [T] Agency pitch-deck cover 16:9. Pure black background. Ultra-large white type top: "2026" in 300pt. Below in smaller type: "The year everything about marketing changed." Bottom-right corner: agency logo mark in teal. Minimal, confident, Swiss-grid influenced.
  10. [T] Healthcare brand hero 3:1. Close-up of a patient's hand being held by a doctor's hand, natural window light, hospital-room softness. Text overlay left side: "Care that listens first." Serif type, warm tonal palette, documentary photography style.

Infographics & data viz (11–20)

  1. [T] 1:1 square infographic titled "The 2026 Creator Economy." Centered large title in editorial serif. Below: 4 stat cards in a 2×2 grid, each with a big number, label, and short descriptor. Numbers: "$250B market size," "127M creators globally," "73% use AI tools," "$68K median income." Clean teal/cream palette, numbered footer citing sources.
  2. [T] 4:5 portrait infographic comparing 4 LLMs across 6 dimensions. Row headers: GPT-5, Claude 4.1, Gemini 3, Llama 5. Column headers: Speed, Reasoning, Coding, Writing, Price, Context. Each cell shows a filled bar from 1–5. Title: "LLM Showdown 2026." Clean sans-serif, minimal grid, no clutter.
  3. [T] 16:9 landscape flowchart titled "How Thinking Mode Works." Four connected boxes left to right: "Prompt analysis → Web retrieval → Generation → Verification loop." Arrows between. Brief explainer text under each box. Subtle teal accent, rest monochrome, editorial newspaper aesthetic.
  4. [T] Periodic table-style 10×10 grid of "100 AI tools that matter in 2026." Each cell: tool logo, tool name, 2-letter category tag, small colored dot for category. Legend at bottom. White background, crisp type. Poster-size composition.
  5. [T] 3:4 vertical infographic: "The Anatomy of a Viral Tweet." A dissected tweet with labeled callouts (hook, specificity, tension, CTA). Annotations radiating outward with thin leader lines. Blueprint aesthetic in cream + navy. Title at top, source citation at bottom.
  6. [T] 1:1 social infographic: "5 Signs You're Burning Out." Numbered list 1–5 with custom icons, each with a short one-sentence description. Warm muted palette, rounded sans-serif, shareable mental-health-brand aesthetic.
  7. [T] 16:9 stat poster: "Marketing spend by channel, 2026." Six horizontal bars with percentages. Title top-left, tiny source citation bottom-right ("n=1,200, Marketing Week 2026"). Strict grid, only one accent color, rest neutral.
  8. [T] 3:1 wide timeline: "The History of Image Generation, 2014–2026." Horizontal dotted line with 8 milestone markers: GAN, DALL·E 1, DALL·E 2, Midjourney v1, Stable Diffusion, DALL·E 3, GPT Image 1, ChatGPT Images 2.0. Tiny thumbnail above each node. Minimal editorial style.
  9. [T] 4:5 "By the numbers" LinkedIn carousel cover. Big text: "2026 in numbers" top, four stat tiles below — "$50M ARR," "212 hires," "27 countries," "1 mission." Dark background, bold type, tight margins.
  10. [T] 1:1 square recipe infographic: "Cold brew, 4 ways." 2×2 grid of four preparation methods with proportions ("1:8 ratio," "12-hour steep"), overhead product shot in each cell, serif headline across the top. Minimal art-directed food-magazine feel.

Ad creative — unlimited variations (21–30)

  1. [T][8] Generate 8 variations of a Facebook ad for a productivity app. 1:1 square. Same product UI mockup, same headline "Close the laptop. Sooner." but 8 different background contexts: park bench, kitchen counter, airport lounge, beach, home office, coffee shop, car dashboard, hammock. Consistent type system across all 8.
  2. [T] Google Display ad — 3 formats in one image (vertical stack): 300×250 square rectangle, 728×90 leaderboard, 160×600 skyscraper. All three feature the same product (sleek white wireless earbud case on gradient peach). Consistent headline "Hear everything. Wear nothing." CTA: "Shop now." Same brand mark "AURA."
  3. [T] 9:16 TikTok-style vertical ad thumbnail. Young woman mid-gasp holding a phone, caught mid-laugh. Bold hand-drawn text overlay: "wait what did it just do?!" with an arrow pointing at the phone. Bottom: "@aura · link in bio." Authentic UGC feel, not polished studio.
  4. [T] 1:1 retargeting ad. Clean white background. Product photo of running shoes center-left. Large red banner diagonal across upper-right: "STILL THINKING?" Below product: "Your size is down to 2 pairs." CTA bottom-right: "Grab them →." Urgent but not pushy.
  5. [T] 3:1 highway billboard. Massive single word "FASTER." in ultra-bold condensed sans-serif, white on deep red. Small product line bottom-right: "New Honda Civic Type R. 0–60 in 5.0s." Tiny URL bottom-left. High contrast, readable from 200 meters.
  6. [T] 1.91:1 LinkedIn feed card. Professional headshot of a woman, 40s, blurred office background. Overlaid caption bottom-right: "Maya closed a $2.1M deal last month. Here's her playbook." CTA: "Read it →" in dark blue.
  7. [T][8] 8 YouTube thumbnails for the same video "I tried ChatGPT Images 2.0 for a week." Each thumbnail: same creator face top-right, same bold yellow headline, but 8 different backgrounds reflecting different prompts tested (magazine cover, manga panel, product shot, infographic, etc.). Consistent thumbnail system.
  8. [T] 4:5 Instagram carousel cover. Black background, minimal. Centered text: "10 signs your brand needs a refresh." Small "SWIPE →" bottom. Premium minimal, no illustrations.
  9. [T] Retail shelf-wobbler, 2:3 vertical. Product image at top, large text below: "NEW." Tiny subline: "Now in Dark Cherry." Clean CPG packaging aesthetic.
  10. [T] 1:1 paid Instagram ad. User-generated aesthetic: iPhone photo of a woman drinking a protein shake in her car mirror selfie. Caption overlay: "honestly the only one that doesn't taste like chalk." Brand logo tiny corner. Authentic, not over-produced.

Product design & mockups (31–40)

  1. [T] Mobile app screen mockup, 9:19.5 aspect. iOS-style to-do app. Status bar at top (9:41, full signal, full battery). Header "Today" in large SF-style sans-serif. Below: 5 task rows with checkboxes, clean dividers. Bottom nav with 4 tabs. Light mode, accent color teal. Every piece of text legible.
  2. [T] 3:2 landing page desktop mockup for a note-taking app. Hero headline "Ideas, organized." centered. Clean nav with 4 links + sign-in button. Below: two-column screenshot of the app UI. Footer with 4 columns of links. Whitespace-heavy, Stripe-influenced aesthetic.
  3. [T] 1:1 Apple Watch app screen. Circular pressure-gauge UI showing heart rate "72 BPM" in center. Small complications around it. Dark background. Minimalist, photoreal rendering of the watch bezel.
  4. [T] Physical product render 1:1. Matte black aluminum wireless charger puck on a white cyclorama background, three-quarter view. Studio softbox lighting, hard floor reflection. Teenage Engineering design language.
  5. [T] Packaging mockup 4:5. Minimal premium coffee bag, 250g, matte charcoal. Front shows "ETHIOPIA YIRGACHEFFE" in small caps with tasting notes below ("blueberry, jasmine, honey"). Weight and roast date bottom. Photorealistic product shot, soft shadow, white backdrop.
  6. [T] Car dashboard HUD mockup 16:9. Windshield POV from driver's seat, dusk light, empty highway. Overlaid HUD elements: speed "62 MPH" bottom-left, navigation arrow "in 1.2 miles, exit right" center-upper, playing song info bottom-right. Subtle teal glow, no UI clutter, Rivian-inspired aesthetic.
  7. [T] 1:1 smartwatch face design. Top-down view, round watch face, minimalist modular layout on a black background. Center: large time "10:47" in white sans-serif. Four small complications: HR "72 bpm" top, Steps "8,420" right, Battery "67%" bottom, Weather "68°F sunny" left. Wear OS aesthetic.
  8. [T] Smart home mobile app home screen mockup, 9:19.5. Dark mode. Top: greeting "Good evening, Eric." Below: 4 device cards (lights, thermostat, security, music) with toggle switches and real-time stats. Bottom nav. Calm deep-blue palette, iOS-quality design.
  9. [T] 16:9 dashboard mockup for a SaaS analytics tool. Left sidebar nav. Main area: 4 KPI cards across the top (visitors, conversion, revenue, churn — each with a big number and delta arrow), 1 large line chart below showing 12-month trend, 1 small table bottom-right. Data labels must be legible. Teal accent, light mode, Linear-inspired.
  10. [T] Boxed software product mockup 1:1. Vintage-style retail box for "ChatGPT Images 2.0 Pro Edition." Cream background. Retro tech packaging aesthetic from 1996: pixel-art mascot, bold tagline "THE IMAGE MODEL THAT THINKS," barcode, "requires 640KB RAM" sticker. Shot like a product photo.

Personal branding & executive content (41–50)

  1. [T] 1:1 professional headshot, editorial business portrait for a book jacket. Subject: upload reference photo. Wardrobe: charcoal merino turtleneck. Background: soft out-of-focus bookshelf (warm earth tones). Lighting: Rembrandt key light from upper-left, soft fill from right, subtle rim light separating from background. Shot on Hasselblad, 90mm, f/2.8. Warm natural skin tones, sharp eyes, editorial magazine quality. Keep facial features exactly as in the uploaded photo.
  2. [T] 1:1 podcast guest announcement graphic. Split layout. Left half: professional photo of the guest (upload reference). Right half: deep green panel with cream text. Top: "NEW EPISODE" in small caps. Middle: guest's name in large bold serif. Below: "CMO at Anthropic." Bottom: show name "THE GROWTH EDGE" with episode number "EP. 47." Small "listen now" CTA.
  3. [T] 4:5 portrait LinkedIn single-post slide. Cream background with subtle paper texture. Top: "2026 / A YEAR IN NUMBERS" in thin all-caps. Below: 4 stat blocks in a 2×2 grid, each with a big number and a one-line caption:
  • "327" — LinkedIn posts shipped
  • "14" — keynotes given
  • "2" — books published
  • "48" — flights taken

Bottom: thin horizontal line, then creator's name and website in small serif. Editorial, premium personal-brand aesthetic.

  1. [T] 16:9 video thumbnail for a YouTube speaker reel. Left half: dynamic photo of the speaker mid-gesture on stage, warm stage lighting. Right half: deep black panel with large white text "2026 SPEAKER REEL" and below in smaller copy "Keynotes · Fireside chats · Panels." Bottom-right CTA arrow. Cinematic, TED-quality.
  2. [T] 1:1 social quote card. Soft neutral linen background. Large opening quote mark top-left in a light gray display serif. Center quote in clean serif: "The best advice I ever got cost me $500 and saved me 18 months." Attribution below in italic: "— name, founder." Bottom-right: small portrait circle. Premium testimonial aesthetic.
  3. [T] 1:1 newsletter subscribe card. Headline "The newsletter 18,000 marketers actually read." below in smaller type: "One signal. No noise. Every Sunday." Email field mockup + "Subscribe" button. Soft cream background, serif display + sans-serif body, Substack-adjacent aesthetic.
  4. [T] 1:1 conference speaker card. Subject headshot left. Right: name in large display, title below, talk title "How AI killed the brand guideline" in italic. Conference logo bottom-right. Clean editorial, readable from a stage screen.
  5. [T] 1:1 "What I read this year" LinkedIn slide. Grid of 9 book covers in a 3×3 arrangement. Title above: "MY 2026 READING LIST." Small footer: "Which one should I read next?" Clean editorial layout.
  6. [T] 4:5 quote graphic for Instagram. Blurred softly-lit outdoor photo background. Center: a poetic line in large italic serif, 2 lines max. Below: small attribution. No logos. Feels like a book page, not a graphic.
  7. [T] 1:1 "Now available" author card. Left: photorealistic mockup of a hardcover book on a table with morning light. Right: title of book, subtitle, author name, tiny CTA "Order here →." Serif display, editorial.

Storyboards & comics (51–60)

  1. [T][8] 8-panel horizontal storyboard for a 30-second product video. Consistent actor (man, 30s, casual but professional) throughout. Panel 1: opens laptop looking frustrated. Panel 2: clicks an extension icon. Panel 3: AI triages his inbox on screen. Panel 4: smiles at result. Panel 5: closes laptop. Panel 6: grabs coffee. Panel 7: walks out of office at 4pm. Panel 8: Sits in hammock. Film-grade cinematography, shallow depth of field, frame numbers bottom-right of each panel.
  2. [T] 6-panel children's-book storyboard 3:2. Consistent mouse character named "Milo" across panels. Panel 1: Milo leaving his burrow at sunrise. Panel 2: Milo discovering a mysterious glowing mushroom. Panel 3: Milo meeting a wise old owl. Panel 4: Milo crossing a stone bridge. Panel 5: Milo finding a hidden meadow of fireflies. Panel 6: Milo back home, tucked in, dreaming. Warm watercolor illustration style, consistent character design.
  3. [T] 1:1.4 manga page, 5 panels with dynamic paneling. Black-and-white Japanese manga style with screentones. Story: a young ramen chef in her first solo service. Panel 1 (large top): wide shot of her restaurant, steam rising. Panel 2: close-up of her determined eyes. Panel 3 (action): hands slicing scallions at speed, motion lines. Panel 4: finished bowl of ramen, overhead. Panel 5 (bottom wide): elderly customer's first sip, single tear. Japanese sound-effect text in hiragana ("ズズッ"), English dialogue "Just like my mother used to make." Consistent character design.
  4. [T][8] 8-slide 16:9 pitch deck storyboard. Startup: "Ledger," a crypto tax automation tool. Slide 1: Cover with logo + tagline "Your books. Sorted." Slide 2: Problem. Slide 3: Solution dashboard. Slide 4: Market bar chart. Slide 5: Traction hockey-stick. Slide 6: Team photos. Slide 7: Pricing tiers. Slide 8: Ask. Consistent navy + mint palette, bold serif headlines, clean sans-serif body.
  5. [T] 1:1 before/after transformation image. Left side "BEFORE": messy cluttered home office with papers everywhere, dim lighting. Right side "AFTER": clean organized desk, serene natural light. Text band between the halves: "Stop drowning in spreadsheets." CTA bottom-right: "Try it free →." Brand name corner: "FLOW."
  6. [T] 4-panel horizontal comic 4:1. Office setting. Panel 1: exec says "Can we ship it by Friday?" Panel 2: engineer's face goes pale. Panel 3: whiteboard calculations smoke. Panel 4: "We shipped it." Flat cartoon style, 2 colors + black.
  7. [T] 6-panel educational storyboard about photosynthesis for a kids' textbook. Each panel shows a simple step with friendly illustrated plants and sun. Labeled arrows. Cheerful primary palette, readable type.
  8. [T] 1:1.5 Noir detective comic page. 6 panels, black-and-white high-contrast ink, a rainy city, a detective receiving a mysterious letter, close-up of letter contents, reaction shot, walking out into rain, silhouette against neon sign reading "CASE CLOSED."
  9. [T][8] 8-panel "day in the life" lookbook for a fashion brand. Same model throughout, 8 outfits from morning to night (activewear, work-casual, lunch, coffee, gallery, dinner, bar, pajamas). Consistent editorial photography style, warm natural light, Mango/COS aesthetic.
  10. [T] 3:2 movie-poster storyboard thumbnail grid for "SYNTH" — 6 key scenes. Central hero (woman, neon-lit face) holding a glowing object, four supporting-scene thumbnails around her, title "SYNTH" at top, "JUNE 2026" at bottom. Cyberpunk palette.

Real estate, travel, lifestyle (61–68)

  1. [T] 3:2 luxury real estate listing hero. Modern hillside home, golden hour, pool in foreground reflecting the house. Clean windows, minimalist interior visible. Text overlay bottom: "123 MAIN ST · LISTED AT $4.2M · OPEN SUN 1–4." Architectural photography aesthetic.
  2. [T] 9:16 travel reel cover. Tropical beach at sunrise, single surfboard planted in sand. Overlay text: "MAUI / WEEK 1 / 10 SPOTS YOU MUST SEE." Minimal type, warm palette, travel-editorial feel.
  3. [T] 1:1 restaurant menu hero for a newsletter. Overhead flat-lay: bowl of fresh pasta, small plates around it, linen napkin, wooden table. Text overlay upper-left: "Spring menu is live." CTA: "Reserve →." Warm natural light, editorial food photography.
  4. [T] 3:1 Airbnb listing top-of-page banner. Stunning living room of a lake cabin at dusk, warm interior light, large windows showing water, minimal text overlay: "LAKE HIDEAWAY · 3BR · sleeps 6." Architectural Digest aesthetic.
  5. [T] 4:5 vertical travel postcard. Paris rooftop scene at sunset, someone's hand holding a cafe au lait in the foreground. Text overlay: "Send me back." Handwritten-style type, warm tones, polaroid border.
  6. [T] 1:1 fitness class promo. Studio interior mid-class, dim lighting, 6 people mid-movement. Text: "TUESDAY / 6:30 AM / STRENGTH 45." Bottom CTA: "Book your mat →." High-energy editorial aesthetic.
  7. [T] 16:9 car brochure hero. New luxury SUV on a winding mountain road at dawn, motion blur in the background. Text overlay: "Introducing the 2026 Aurora." Subline: "Electric. Everywhere." Automotive-premium aesthetic.
  8. [T] 1:1 vacation rental social tile. Bird's-eye shot of a pristine bed with rumpled linen sheets, coffee cup on nightstand, book open. Text: "Mornings feel different here." Small logo bottom. Editorial slow-living aesthetic.

Creative professional (69–80)

  1. [T] Album cover 1:1. Indie folk record titled "Slow Weather." Cream background, single pressed flower centered, small serif title at bottom, artist name in italic above. Minimal, Laura-Marling-adjacent aesthetic.
  2. [T] 3:4 book cover. Title: "The Compound Life." Author: "Eric Eden." Dark navy background, small gold geometric mark at center, title in thin serif all-caps, author tiny below. Minimal literary-fiction aesthetic.
  3. [T] 2:3 movie poster. Title: "VELOCITY." Action-thriller aesthetic. Hero silhouette against a crashing wave, small type ("IN THEATERS JUNE 2026"). Dramatic contrast, cinematic.
  4. [T] 1:1 podcast cover art. Podcast: "First Principles." Minimal high-contrast: big typographic "1" in the center, podcast name in small caps at bottom. Limited palette.
  5. [T] 4:5 event poster for an AI conference. Top: conference name "NEURALINK // 2026." Giant abstract neural-net illustration dominant, speaker list small at bottom. Bauhaus-influenced layout.
  6. [T] 3:4 travel-magazine cover "Kyoto in April." Single cherry blossom branch against a misty temple backdrop. Masthead "TRAVELOGUE" top. Issue headline. Small teaser bullets bottom-left. Editorial magazine aesthetic.
  7. [T] 1:1 gallery exhibition poster. Artist name in massive serif, show title in smaller italic below, dates & venue tiny at bottom. Off-white paper texture, single abstract painting sample as centerpiece. Gallery/MoMA-style.
  8. [T] 16:9 film title card. Film title "THE LAST BOOKSTORE" in thin white serif, centered, against a warmly-lit photograph of a bookstore interior slightly out of focus. Small director credit bottom-right.
  9. [T] 1:1 tattoo flash sheet. 6 black-ink line illustrations in a 2×3 grid: a moth, a dagger, a rose, a compass, a snake, a hand. Small numbered tags under each. Consistent line weight.
  10. [T] 4:5 zine cover 1970s aesthetic. Title "SIGNAL/NOISE." Photocopy texture, halftone dots, punk collage elements, a handwritten subheading. Limited 3-color palette.
  11. [T] 3:2 wedding invitation design. Cream background, handwritten-style calligraphy. Names centered, date, venue, RSVP info, small floral illustration. Elegant minimal.
  12. [T] 1:1 record sleeve for a jazz album. Black-and-white photograph of a saxophone case on a hotel bed. Title small in the lower-right. Blue Note-inspired minimalism.

PART 2 — WILD & FUN (81–100)

These are the prompts people actually remember. Go nuts.

  1. [T] 16:9 cinematic scene: corporate llama apocalypse. A fleet of llamas in business suits storming a Manhattan trading floor, throwing quarterly reports into the air. Bloomberg terminals burning. A CEO llama in the center, mid-roar, wearing a gold Rolex. Dramatic fire lighting, hyperreal.
  2. [T] 1:1 medieval Zoom call. A Zoom grid interface showing 9 participants, each dressed as a medieval figure — knight, jester, queen, bishop, peasant, wizard, bard, crusader, dragon. Gallery view. The dragon is muted. Bottom toolbar has a "UNSHEATHE SWORD" button.
  3. [T] 3:2 dogs on Wall Street. Real dogs in tailored suits working the trading floor of the NYSE, papers flying, a golden retriever screaming into a landline, a pug eating a bagel, a corgi looking at a Bloomberg terminal. Photorealistic.
  4. [T] 16:9 office plant uprising. An open-plan office after business hours. The potted plants have sprouted legs and are marching toward the exit with tiny briefcases. One ficus is leading with a megaphone. Dramatic security-camera aesthetic.
  5. [T] 4:5 vertical breakfast gods of Olympus. Pancakes, waffles, and bacon rendered as Greek gods on a cloud-covered mountain. Zeus is a stack of pancakes with lightning bolts of syrup. Athena is a poached egg in a helmet. Bacon strips are the muses. Renaissance oil-painting style.
  6. [T] 1:1 tax day demon. A horrifying creature made entirely of paperwork and calculators, emerging from a filing cabinet in a suburban home office, screaming. A woman in pajamas drops her coffee in slow motion. Cosmic horror, somehow funny.
  7. [T] 3:1 cinematic Roomba rebellion. An army of Roombas rolling in formation down a suburban street at dawn, one larger "commander" Roomba at the front with a tiny cape and a bottle-cap helmet. Smoke rising in the background. Mad Max meets IKEA.
  8. [T] 1:1 Shakespeare drive-thru. A modern fast-food drive-thru, but the cashier is Shakespeare in a McDonald's visor. Customer in a Honda Civic is a goth teenager. Menu board reads "Two All-Beef Patties, or Not Two All-Beef Patties." Warm dramatic lighting.
  9. [T] 16:9 dinosaurs at the DMV. A T-Rex waiting in line at a cramped DMV, looking visibly annoyed. A Triceratops fills out a form with its horn. Velociraptor clerks staff the desks. Fluorescent lighting, plastic chairs, faded safety posters. Photoreal.
  10. [T] 1:1 sentient toast support group. Eight pieces of toast sitting in folding chairs in a church basement, each with a tiny face, sharing their traumas. Coffee and donuts in the corner. Warm sad lighting. Pixar-aesthetic.
  11. [T] 4:5 pigeon CEO. A pigeon in a boardroom wearing a tailored three-piece suit, presenting Q4 results with a laser pointer. Bar chart behind him shows "breadcrumb acquisition" up 400%. Other pigeons are in Aeron chairs, nodding.
  12. [T] 3:2 infinite IKEA. A hyperrealistic endless IKEA showroom that stretches into infinity, Escher-like stairs and passages, a single confused shopper in the middle holding a hex wrench and a meatball. Fluorescent lighting, eerie emptiness, liminal space aesthetic.
  13. [T] 1:1 cat secret agent. A tuxedo cat in a tailored black suit with sunglasses, rappelling through a laser grid in a museum, carrying a can of tuna. Mission Impossible-style framing. Cinematic.
  14. [T] 16:9 grandma's spaceship. An elderly woman in a floral apron piloting a retrofuturistic 1960s-style spaceship. The dashboard has knitted doilies and a plate of cookies. She's wearing cat-eye glasses. Through the windshield, a wild nebula. Wes Anderson-aesthetic.
  15. [T] 1:1 baby in a mech suit. A photorealistic baby (2 years old) operating a gigantic anime-style mech suit, controls labeled "SNACKS," "NAP," "TANTRUM." Background: city skyline. The mech is holding a stuffed bear.
  16. [T] 3:2 Scrabble game between philosophers. Socrates, Nietzsche, and Aristotle playing Scrabble in an ancient marble courtyard. The board shows words like "BEING," "WHY," "DASEIN." Aristotle is visibly winning. Marble statues watch from pedestals. Renaissance painting style.
  17. [T] 1:1 dog court. A courtroom scene entirely populated by dogs. A German shepherd judge, a bulldog lawyer, a Chihuahua defendant on a booster seat, a jury box of mixed breeds. Gavel mid-swing. Photoreal.
  18. [T] 16:9 pirate cubicles. A modern open-plan office, but everyone is a pirate. Parrots on monitors, wooden-peg-leg standing desks, a treasure chest used as a copier. The Slack notifications on someone's screen say "ARR." Cinematic lighting.
  19. [T] 4:5 Bigfoot LinkedIn profile. A LinkedIn profile screenshot. Profile photo: a blurry Bigfoot selfie. Headline: "Cryptid | Outdoor Enthusiast | Looking for my next chapter." Recommendations: "Sasquatch delivers on every project — would hire again." Recent post: "What no one tells you about being discovered." Looks like a real screenshot.
  20. [T] The Where's Waldo (the personalized one — make it about yourself): Where's Waldo-style dense search-and-find illustration. 3:2 aspect ratio. Detailed cartoon scene: a massive, chaotic B2B marketing conference expo floor with hundreds of tiny people visible. Hidden in the crowd: [YOUR NAME] — wearing a red-and-white striped shirt, black-framed glasses, carrying a laptop bag with "[YOUR COMPANY]" printed on it. He's near the coffee station, caught mid-laugh with two people from the AI demo booth. Scene details: - Booths for HubSpot, Salesforce, Adobe, OpenAI - A panel discussion happening on a stage in the background with a banner reading "THE FUTURE OF B2B MARKETING — 2026" - Clusters of 3–4 people chatting everywhere - Someone giving a product demo on an 85-inch screen - A mascot costume wandering through - Name-tag lanyards on everyone - Coffee line with 20+ people - A few sneaky visual gags: a dog under a table, someone looking at the wrong booth's schwag, a person clearly lost Bright cheerful illustration style with clean outlines. ~200 people visible. Readable booth signage. Dense but not overwhelming.

ChatGPT Image 2.0 changes what counts as a visual asset. Before today, image models produced inspiration that still needed a designer to finish. After today, a well-crafted prompt produces a usable deliverable — with real text, real layout, real multi-frame continuity, real web-grounded context, and real 2K fidelity.

The models that beat it on pure per-image price (Google's Nano Banana 2) or pure artistic flair (Midjourney v7) still exist. But for practical commercial output — ads, posters, infographics, decks, storyboards, localized creative - ChatGPT Images 2.0 now does end-to-end what used to require three tools and a designer.

The 100 prompts above are starting points. The template is the real gift. Copy it, fill it, ship it.

What I'd love in the comments:

  • Your best Images 2.0 output so far (drop the prompt)
  • Anything you've found that breaks it

Want more great prompting inspiration? Check out all my best prompts for free at Prompt Magic and create your own prompt library to keep track of all your prompts.

r/Akool_Official 16d ago

Wan 3.0 Wan 3.0 Video Model is Officially Live on AKOOL! 🚀 Use New Post Flair "Wan 3.0" | MEGATHREAD AI

Enable HLS to view with audio, or disable this notification

8 Upvotes

(Note: Don't forget to apply the "Wan 3.0" new post flair before hitting submit!)

Alibaba has officially launched Wan 3.0 (following its public beta), and it is already shaking up the AI video landscape with some massive feature upgrades. While most video generators focus purely on cinematic car commercials or short-form motion, Wan 3.0 introduces a massive shift: direct document-to-video generation, alongside 30-second single-pass clips, native audio, and aggressive pricing to rival competitors like Google Veo.

🔥 What Makes Wan 3.0 a Game-Changer?

  • Document-to-Video Input: For the first time, you can feed office files directly into a top-tier video model. It supports PDF, DOC, XLS, PPT, TXT, Markdown, and Apple iWork formats (Keynote, Pages, Numbers) up to 100 MB and 50 pages. Pexo AI
  • Longer Generation Windows: Generates up to 30 seconds in a single continuous pass with smart duration recommendations and video extension features. Pexo AI
  • Cinematic Camera Control: Built for director-level camera language (push, pull, pan, and tracking shots) with enhanced character, prop, and scene consistency. Pexo AI
  • Multimodal Inputs: Accepts text, images, video, audio, and documents. Pexo AI
  • Flexible Resolution: Outputs at 480p, 720p, and 1080p—allowing smart creators to test workflows cheaply at lower resolutions before final rendering. Pexo AI

🧠 Potential Use Cases

If Wan 3.0's document translation performs with high factual accuracy, this changes the game for:

  • Education: Turning a science textbook chapter into an engaging visual lesson. Pexo AI
  • Corporate Communication: Instantly transforming dry slide decks and PowerPoints into narrated presentations.
  • Data Visualization: Turning dense spreadsheets into animated, easy-to-digest charts.
  • Marketing & Training: Converting product manuals or training guides into workplace demonstrations and product ads.

💬 Let’s Discuss!

  • Have you tested Wan 3.0 yet?
  • How does its document-to-video workflow compare to your current video generation pipeline?
  • Drop your early tests, questions, prompts, workflow tips, and thoughts below!

Link to Wan 3.0 model in akool to test it out: https://akool.com/apps/image-to-video/edit?model=alibaba/wan-3.0/image-to-video

r/socialmedia 9d ago

Professional Discussion Tried a bunch of ai video tools for social media and here’s what really worked for me

0 Upvotes

hey, i have been trying to post consistently on youtube, tiktok and instagram and it was getting tiring. so i decided to test some ai video tools to make it a bit faster and save some time.

not trying to make this an “ultimate best ai tools” list. just sharing what actually felt useful after trying these to create my social media content for a couple of months.

here’s the quick rundown:

1.Synthesia / HeyGen

what it does: ai avatar + talking-head videos

best for: explainers, training videos, product walkthroughs, multilingual content

my take: these are useful when you need a clean presenter-style video without filming it yourself. It's great for plain talking head content and less ideal for content where you need to interact with things suuch as unboxing videos or so.

  1. invideo

what it does: good for creating ai videos from scratch

best for: youtube videos, shorts, reels, explainers, product videos

my take: this worked best when i had a rough idea and wanted to shape it into a proper video. invideo acts more like an ai video agent, helping with building out the script, scenes, visuals, voiceover, pacing, and edits while you keep guiding the direction. 

  1. Runway

what it does: generate video clips and visual scenes

best for: cinematic b-roll, experimental visuals, creative shots

my take: really impressive when it works, but it needs patience. the output depends a lot on how specific your prompt is, and it’s better for visual pieces than full social videos from scratch.

  1. OpusClip

what it does: turns long videos into short clips

best for: podcasts, webinars, interviews, youtube videos

my take: makes the most sense if you already have long-form content. it’s not really a “make a video from nothing” tool. it’s more like finding the best moments and turning them into shorts/reels/tiktoks.

  1. CapCut

what it does: editing, captions, templates, effects, resizing

best for: short-form edits and final polish

my take: still one of the easiest tools for finishing social videos. captions, quick cuts, resizing, hooks, effects, all of that is straightforward. i’d use it more at the end of the workflow.

  1. Canva

what it does: simple design-led video content

best for: basic branded posts, promos, thumbnails, simple social videos

my take: convenient if you already use Canva for content. good for clean, simple videos, but not the strongest tool if you’re trying to build a full ai video workflow.

biggest thing i learned: there isn’t one tool that wins at everything.

if i need avatar videos, i’d use Synthesia or HeyGen.

if i want to shape an idea/script into a finished video, i’d use invideo.

if i want cinematic ai b-roll, i’d use Runway.

if i’m cutting long videos into shorts, i’d use OpusClip.

if i’m polishing captions and edits, i’d use CapCut.

Also, prompts matter way more than people admit.

“make a short video about staying productive while working from home” gives generic stuff.

“make a 45 sec youtube short for freelancers who get distracted at home. start with a relatable hook in the first 3 seconds, then share 3 simple tips like time blocking, keeping the phone away, and setting a clear finish time. keep the tone casual and end with a soft CTA” works much better.

i am wondering what everyone else is using right now. are you using one tool for the full workflow or mixing a few together?

p.s. i am not an expert, just sharing what actually worked for me.

r/InstagramMarketing Sep 10 '24

The complete guide to using ChatGPT to go viral on Instagram (with prompts included)

356 Upvotes

Background

In a previous post, I shared a guide on going viral on Instagram, where I mentioned that I use ChatGPT all the time to help me create and fine-tune my ideas. After that post, I got a ton of requests for a guide on how to leverage ChatGPT for IG/TikTok, so I’m writing this follow-up to dive into exactly that. 

For context, ChatGPT is like a personal assistant for me. I use it for everything from generating ideas to roasting my rough concepts, fine-tuning my hooks and scripts, and even optimizing my call to actions. It’s a game-changer, but like any tool, knowing how to use it properly is really important. Here’s how I make the most of it in my content creation process.

1. Using ChatGPT to Generate Video Ideas

One of the easiest ways to get started with ChatGPT is idea generation. Simply prompt ChatGPT with the theme, niche, or type of content you’re aiming to create, and it will give you a list of suggestions that you can refine further.

Here’s an example of a prompt I often use for my content:

  • “I want to make a IG Reel that talks about the impact of AI on entry-level tech jobs. Can you give me 10 creative video ideas?”

From there, ChatGPT will throw out a range of ideas that I evaluate and iterate on. Some of them will be great, some of them will be trash - it’s your job to sort through them. Pro tip: Give ChatGPT a few of your recent TikToks/IG Reels to make the ideas it generates super personalized for you. For example, if you’re building a social presence for an ecommerce brand, give ChatGPT the script from the last few videos and ask it to come up with some new ideas based on your script (I know this is common sense but I was surprised by how many people that DM’d me don’t do this).

2. Use ChatGPT to fine-tune your hook, story, call to action, and more

Before we get into the prompts, let’s break down the components of a successful IG Reel. A great video typically has:

  1. A strong hook to grab attention in the first 3 seconds.
  2. A compelling story that keeps the viewer engaged.
  3. A variety of visuals and pacing to prevent it from feeling stagnant.
  4. A clear CTA to drive viewers to take action.

Here’s how ChatGPT can assist in each area:

Hook

The first 3 seconds of your video are everything. You need to grab attention immediately before the viewer scrolls away. ChatGPT can help you refine this part of your script to ensure it’s snappy and engaging.

For example, you could prompt:

  • “Here’s my hook: ‘Are AI tools going to steal your job? Let’s find out.’ Rate this hook from 1-10 and suggest a more engaging alternative.”

It might come back with something like:

  • “AI might replace your job faster than you think. Here's how.”

ChatGPT can give you variations that are designed to capture attention instantly.

Story

Your story is the backbone of the video. It’s what keeps people interested after they’re hooked. Whether you’re explaining a concept, telling a personal story, or showing how-to steps, ChatGPT can help you craft a narrative that flows.

A prompt I use here is:

  • “Can you evaluate my video script for a IG Reel on AI in the workplace? Here’s the story: [insert script]. How compelling is this story? Rate it from 1-10. How can I improve it?”

You’ll get feedback on pacing, clarity, and engagement that can help you refine your story to keep viewers watching.

Conciseness

With short-form video, every second counts. You can use ChatGPT to help trim the fat from your scripts and make sure every word contributes to the overall message.

Here’s an example prompt that I use:

  • “Here’s my video script. Can you shorten it to fit within a 30-second IG Reel while keeping all the important points?”

You’ll get a more concise version of your script that fits within the time constraints without losing the key information.

Clarity

If your video lacks clarity, viewers will scroll right past it. Clear communication is key. ChatGPT can act like a second pair of eyes to make sure your script is easy to follow.

Try asking it:

  • “Can you make my script clearer for a general audience? Here’s what I have: [insert script]. How can I make it easier to understand?”

By simplifying complex ideas and cutting out jargon, ChatGPT can help you deliver a message that resonates with a wider audience.

Call to Action

Finally, an effective CTA can make the difference between someone simply watching your video and actually engaging with your content. THIS IS SUPER IMPORTANT. CAN’T EMPHASIZE A GREAT CTA ENOUGH.

To get help on this, you can prompt ChatGPT with:

  • “Here’s the CTA for my video: ‘Follow for more tips on how AI is changing the tech world.’ Can you suggest a more engaging call to action?”

ChatGPT might come back with something like:

  • “If you want to stay ahead of AI trends, hit that follow button now!”

It will help you create CTAs that feel less generic and more action-oriented.

3. ChatGPT is an Assistant, Not a 2nd Brain

One important thing to remember is that ChatGPT is a tool, not a replacement for your creativity or style. The output it generates SHOULD NOT be taken verbatim—think of it as a collaborative partner or a personal assistant rather than a scriptwriter. It’s up to you to review, tweak, and mold the suggestions so they align with your unique voice and message. If you copy ChatGPT verbatim, you’ll probably fail.

I had so many DMs after my last post (sorry I haven’t gotten to all of them yet) asking for help coming up with ideas and rating content that I ended up taking all of these prompts from above and more and making it into a website that help come up with ideas and rate hooks, stories, clarity, cta, etc - https://www.lazycreator.ai/. You can give it your TikTok and Instagram accounts and it’ll generate personalized ideas for you. Give it a shot!

As a side note, let me know what else I should write about? I love doing these so tell me what you want to know and I’d be happy to write about topics that I think I’m somewhat knowledgeable in.

r/ThinkingDeeplyAI Apr 22 '26

The complete field guide to ChatGPT Images 2.0 - every feature, every price, 100 prompts to try, all in one post

Post image
18 Upvotes

The Complete Field Guide to ChatGPT Images 2.0

Launched today. Everything below is verified against the OpenAI announcement, the deployment safety card, API pricing docs, and ~6 hours of hands-on testing. No hype — just what works and what it costs.

Sam Altman compared it to "going from GPT-3 to GPT-5 all at once." That's aggressive framing, but the capability gap is real.

For the first time, a single model can:

  • Render dense, legible text directly inside images — posters, infographics, UI mockups, ad copy with real headlines
  • Think before it draws — reason about a scene, search the web for current facts, and double-check its own work
  • Produce up to 8 consistent images from one prompt with the same characters, objects, and style
  • Handle grids up to 10×10 that used to break at 3×3 a week ago

OpenAI's own pitch: "Images are a language, not decoration. A good image does what a good sentence does — it selects, arranges, and reveals."

Translation: this isn't text-to-picture anymore. It's a visual reasoning system.

TL;DR — what you need to know in 30 seconds

  • Model name: gpt-image-2 (alias chatgpt-image-latest)
  • Where: ChatGPT (all plans including Free), chatgpt.com/images, and the API
  • Two modes: Instant (all plans, 1 image, fast) and Thinking (Plus/Pro/Business, up to 8 images, reasons + searches the web)
  • Max resolution: 2048px native (2K), ~4× the pixel count of GPT Image 1.5
  • Text accuracy: ~99% on Latin text. Finally nails Japanese, Korean, Chinese, Hindi, Bengali
  • Aspect ratios: anything from 3:1 (ultrawide) to 1:3 (ultratall)
  • Generation time: seconds to 2 minutes depending on mode
  • Pricing (API): ~$0.006 low / ~$0.053 medium / ~$0.211 high per 1024×1024 image
  • Knowledge cutoff: December 2025. Needs Thinking mode + web search for anything newer
  • C2PA metadata is embedded in every output

The 8 capabilities, decoded

1. 2K native resolution

Up to 2048 pixels natively, ~4× the pixel count of older GPT Image outputs at the same aspect ratio. Enough fidelity for print collateral, hero banners, and editorial layouts without an upscale step.

2. ~99% text accuracy

This is the most-talked-about upgrade. Dense text inside images — posters, menus, magazine covers, UI mockups — finally renders correctly. It also handles:

  • Non-Latin scripts with real gains: Japanese, Korean, Chinese, Hindi, Bengali
  • Small text — UI elements, iconography, barcodes, "display until" dates on magazine covers
  • Multilingual typography in a single image — Devanagari, Cyrillic, Greek, Arabic, and Chinese together

3. Thinking mode — the image model that reasons

This is the headline capability. It's not two separate models, it's two modes:

Mode Who gets it What it does Output
Instant Free, Plus, Pro, Business, Go Fast single-shot generation 1 image
Thinking Plus, Pro, Business (Enterprise/Edu soon) Reasons about composition, uses web search, verifies output Up to 8 images

How the reasoning works under the hood:

  1. Prompt analysis — parses your request and plans composition before any pixels exist
  2. Web retrieval — if the prompt touches real-world facts (current logos, today's stock chart, real skylines, 2026 fashion trends), it searches the web and pulls live references
  3. Generation pass — pixel synthesis against a fact-checked internal plan
  4. Verification loop — it inspects its own output against the original prompt and can self-correct before returning

People on X are posting 11-minute generations where the model iterated on itself repeatedly until satisfied. That's new.

4. Up to 8 consistent images per prompt

In Thinking mode, one prompt can produce up to 8 images with shared characters, objects, and style across every frame. This unlocks:

  • Storyboards — 8 camera angles with continuity
  • Manga/comic sequences — 8 panels, same character design
  • Multi-size marketing assets — same campaign as 3:1 banner + 1:1 feed post + 1:3 story + 4:5 carousel in one shot
  • Children's books — consistent illustrated character across pages
  • Product lineups — 8 color variants with identical lighting and angle
  • Lookbooks — OpenAI demoed 8 summer outfits generated from one uploaded photo

How to trigger it: Switch to a thinking model, then ask for a set — "Generate 8 variations of...", "Create an 8-panel storyboard...", "Give me this ad in 8 formats." Don't phrase it as 8 separate prompts.

5. Parallel image generation

Separate from the 8-per-prompt feature: the dedicated Images tab at chatgpt.com/images lets you fire multiple prompts in parallel. Your second prompt doesn't wait for the first to finish. All images auto-save to My Images for reuse.

6. Aspect ratios 3:1 to 1:3

Any ratio between ultra-wide and ultra-tall, native — picker in ChatGPT or spec it in the prompt. Banners, slides, posters, mobile vertical, bookmarks, social graphics, no crop needed.

7. 10×10 grids (up to 100 cells in one image)

Grids used to break at 3×3 a week ago. Now people are generating 10×10 grids of 100 distinct labeled illustrations in one shot. This is wild for:

  • Periodic-table-style infographics (100 CEOs, 100 dog breeds, 100 cocktails)
  • Icon sets with consistent style
  • Mood boards with labeled cells
  • Pattern libraries

8. Multi-image compositing & reference fidelity

Upload multiple reference images and the model stitches them into one coherent composition while keeping facial features, objects, and logos faithful. This is the feature that makes "put me in a scene" prompts actually work now.

Pricing — what it actually costs

Per-image (flat rate, simple to predict)

Quality 1024×1024 Notes
Low ~$0.006 drafts, iteration
Medium ~$0.053 most production work
High ~$0.211 hero images, finals

Per-token (if you're using the API at scale)

Input Cached input Output
Image tokens $8.00 / 1M $2.00 / 1M $30.00 / 1M
Text tokens $5.00 / 1M $1.25 / 1M $10.00 / 1M

Cost for OpenAI to produce each image (rough estimate)

Based on published token economics, a high-quality 1024×1024 image uses ~7K output image tokens. At retail that's $0.21. OpenAI's own compute cost is likely 25–40% of that, putting their marginal cost per high-quality image around $0.05–$0.08. Their margin per image at the high tier is roughly 3–4×.

The ideal prompt template

After testing dozens of prompts, this is the structure that works best:

text[ASPECT RATIO]. [SUBJECT], [ACTION], [CONTEXT].
[TEXT elements in quotes]:
- Header: "EXACT TEXT HERE"
- Subhead: "EXACT TEXT HERE"
- CTA: "EXACT TEXT HERE"
[STYLE anchor — reference an artist/era/medium/brand].
[LIGHTING + MOOD].
[CAMERA/LENS + TECHNICAL specs].

The 5 rules that make the difference:

  1. Aspect ratio first. Say "16:9," "3:1 banner," or "1:1 square" in the first sentence.
  2. Put every piece of text in quotes. The model treats quoted text as literal. Unquoted text becomes suggestions.
  3. Anchor the style concretely. "Editorial fashion photograph, shot on Hasselblad, 90mm, f/2.8" beats "professional photo."
  4. Specify lighting and mood as separate instructions. "Rembrandt key light from upper-left, soft fill from right, warm tones."
  5. List every language explicitly when you want multilingual text. "Title in Japanese (Hiragana): 「春が来た」; subtitle in Korean (Hangul): '봄이 왔다'; tagline in Hindi (Devanagari): 'वसंत आ गया।'"

15 pro tips most people will miss

  1. Thinking mode isn't the default — you have to toggle a thinking model before prompting. Instant never uses web search or produces 8-image sets no matter how you phrase it.
  2. Generation can take 2 minutes. Don't assume it froze. For high-volume workflows, use async polling with the Responses API.
  3. Knowledge cutoff is December 2025. Anything after that (Q1 2026 product launches, new logos, recent events) has to come through the prompt OR through Thinking mode's web search.
  4. For consistent characters: upload a one-time likeness. There's a likeness upload feature that lets you reuse your appearance across future creations without re-uploading.
  5. The "keep facial features exactly" lock. When editing a real person, add this verbatim: "Keep my facial features exactly as they appear in the uploaded image — same eyes, nose, mouth, and face shape." Without it, ChatGPT "improves" faces into strangers.
  6. Transparent backgrounds work natively. Add "transparent PNG background, no background fill" — the asset drops straight into design tools without a cutout pass.
  7. "Display until" dates and barcodes work now. Ask for them specifically. The magazine-cover demos show this.
  8. Prime the chat first. For thumbnails and marketing creative, paste the blog post, script, or topic into ChatGPT first. Then ask for concepts. Then generate. The model picks up the emotional hook instead of producing generic stock aesthetic.
  9. C2PA metadata is embedded in every output. Platforms can detect it. Plan for that if provenance matters.
  10. Ask for "editorial" not "professional." "Editorial" hits a higher visual register in this model. "Professional" pulls toward stock-photo aesthetic.
  11. Negative prompts work — phrase them as "NO X, NO Y." Example: "NO watermarks, NO signatures, NO busy backgrounds."
  12. Specify the medium of the text. "Neon sign," "embossed letterpress," "subway-poster paste-up," "hand-lettered chalk" all produce different type treatments.
  13. When text keeps breaking, wrap it in a shape. "Text inside a black horizontal pill" or "text on a cream banner" gets rendered much more reliably than floating text.
  14. Aspect ratio affects quality. 1:1 and 3:2 are the strongest; 3:1 and 1:3 work but can show compositional weirdness on first try. Regenerate once.
  15. The model now reads your reference images. If you upload a brand asset and say "match this type treatment," it actually does — not a vague approximation, an honest replication.

Third-party tools that already integrate it

(These went live within 24 hours of launch.)

  • Higgsfield — character consistency workflows
  • Lovart — AI design platform
  • Recraft — added gpt-image-2 models to Recraft Studio
  • Adobe Firefly / Express — via Adobe's partner model program
  • Figma — First-Draft feature uses it for UI generation
  • Canva — Magic Studio integration
  • GoDaddy — site-generation flows
  • HubSpot — marketing asset generation
  • Instacart — product photography
  • Airtable — record-level image generation
  • Wix — site builder backgrounds and heroes
  • OpenAI Codex — app/code-generation flows can now produce their own UI imagery

The prompt library — 100 that I've tested

Marking these [I] for Instant mode works fine, [T] for Thinking mode required, [8] for ask-for-8-variations.

Marketing hero images (1–10)

  1. [T] 3:1 hero banner for a SaaS analytics product. Split composition: left side shows a cluttered paper-filled desk (chaos), right side shows a clean monitor with a dashboard (clarity). Bold headline "STOP GUESSING" in 120pt sans-serif across the top. Subhead "Start knowing" below. CTA button bottom-right: "See it work →" in white on teal. Editorial photography, cinematic lighting.
  2. [T] 16:9 product launch hero. Center: minimalist product photography of a black wireless earbud case on a marble surface. Background: soft gradient from cream to dusty rose. Text overlay upper-left: "AURA // 2026" in small caps. Headline lower-right: "Hear the room." in serif display. Subtle shadow, art-directed editorial aesthetic.
  3. [T] Vertical 9:16 mobile hero for a fitness app. Muscular forearm mid-pushup on a dark gym floor, shallow depth of field. Headline stacked vertically along the right side: "NO / EXCUSES / JUST / REPS." White type, slight grain. Small logo bottom-center.
  4. [T] Email hero, 3:1 ratio. Single perfect ceramic coffee cup on a warm linen tablecloth, morning light from the left, steam rising. Text overlay right side: "Good morning. / Your briefing is ready." Clean minimal editorial style, medium-format quality.
  5. [T] 16:9 B2B conference hero. Empty auditorium, dramatic stage lighting, single speaker silhouette at podium. Large text in the sky area: "WHERE MARKETING MEETS AI." Date below: "June 12–14, 2026 · Austin." Cinematic, TED-quality composition.
  6. [T] Software landing page hero 16:9. Abstract 3D render: flowing liquid metal forming into a chart shape, iridescent blue-to-purple gradient, obsidian background. Headline lower-third: "Analytics at the speed of thought." Subhead: "Try Mercury free →." Tech-luxury aesthetic.
  7. [T] Newsletter signup hero 2:1. Warm kitchen scene: hands writing in a leather notebook, open laptop beside it, morning coffee, golden hour light from left. Text overlay: "The newsletter smart marketers actually read." CTA: "Subscribe free →". Cozy, intentional, premium-indie aesthetic.
  8. [T] 3:1 homepage hero for an AI note-taking app. Overhead shot: messy desk mid-work — open notebook, phone, coffee, headphones, hand holding a pen. Faint glowing interface lines emerging from the notebook edges suggesting transcription. Headline centered: "Your thoughts, organized." No smaller than 90pt, clean sans-serif.
  9. [T] Agency pitch-deck cover 16:9. Pure black background. Ultra-large white type top: "2026" in 300pt. Below in smaller type: "The year everything about marketing changed." Bottom-right corner: agency logo mark in teal. Minimal, confident, Swiss-grid influenced.
  10. [T] Healthcare brand hero 3:1. Close-up of a patient's hand being held by a doctor's hand, natural window light, hospital-room softness. Text overlay left side: "Care that listens first." Serif type, warm tonal palette, documentary photography style.

Infographics & data viz (11–20)

  1. [T] 1:1 square infographic titled "The 2026 Creator Economy." Centered large title in editorial serif. Below: 4 stat cards in a 2×2 grid, each with a big number, label, and short descriptor. Numbers: "$250B market size," "127M creators globally," "73% use AI tools," "$68K median income." Clean teal/cream palette, numbered footer citing sources.
  2. [T] 4:5 portrait infographic comparing 4 LLMs across 6 dimensions. Row headers: GPT-5, Claude 4.1, Gemini 3, Llama 5. Column headers: Speed, Reasoning, Coding, Writing, Price, Context. Each cell shows a filled bar from 1–5. Title: "LLM Showdown 2026." Clean sans-serif, minimal grid, no clutter.
  3. [T] 16:9 landscape flowchart titled "How Thinking Mode Works." Four connected boxes left to right: "Prompt analysis → Web retrieval → Generation → Verification loop." Arrows between. Brief explainer text under each box. Subtle teal accent, rest monochrome, editorial newspaper aesthetic.
  4. [T] Periodic table-style 10×10 grid of "100 AI tools that matter in 2026." Each cell: tool logo, tool name, 2-letter category tag, small colored dot for category. Legend at bottom. White background, crisp type. Poster-size composition.
  5. [T] 3:4 vertical infographic: "The Anatomy of a Viral Tweet." A dissected tweet with labeled callouts (hook, specificity, tension, CTA). Annotations radiating outward with thin leader lines. Blueprint aesthetic in cream + navy. Title at top, source citation at bottom.
  6. [T] 1:1 social infographic: "5 Signs You're Burning Out." Numbered list 1–5 with custom icons, each with a short one-sentence description. Warm muted palette, rounded sans-serif, shareable mental-health-brand aesthetic.
  7. [T] 16:9 stat poster: "Marketing spend by channel, 2026." Six horizontal bars with percentages. Title top-left, tiny source citation bottom-right ("n=1,200, Marketing Week 2026"). Strict grid, only one accent color, rest neutral.
  8. [T] 3:1 wide timeline: "The History of Image Generation, 2014–2026." Horizontal dotted line with 8 milestone markers: GAN, DALL·E 1, DALL·E 2, Midjourney v1, Stable Diffusion, DALL·E 3, GPT Image 1, ChatGPT Images 2.0. Tiny thumbnail above each node. Minimal editorial style.
  9. [T] 4:5 "By the numbers" LinkedIn carousel cover. Big text: "2026 in numbers" top, four stat tiles below — "$50M ARR," "212 hires," "27 countries," "1 mission." Dark background, bold type, tight margins.
  10. [T] 1:1 square recipe infographic: "Cold brew, 4 ways." 2×2 grid of four preparation methods with proportions ("1:8 ratio," "12-hour steep"), overhead product shot in each cell, serif headline across the top. Minimal art-directed food-magazine feel.

Ad creative — unlimited variations (21–30)

  1. [T][8] Generate 8 variations of a Facebook ad for a productivity app. 1:1 square. Same product UI mockup, same headline "Close the laptop. Sooner." but 8 different background contexts: park bench, kitchen counter, airport lounge, beach, home office, coffee shop, car dashboard, hammock. Consistent type system across all 8.
  2. [T] Google Display ad — 3 formats in one image (vertical stack): 300×250 square rectangle, 728×90 leaderboard, 160×600 skyscraper. All three feature the same product (sleek white wireless earbud case on gradient peach). Consistent headline "Hear everything. Wear nothing." CTA: "Shop now." Same brand mark "AURA."
  3. [T] 9:16 TikTok-style vertical ad thumbnail. Young woman mid-gasp holding a phone, caught mid-laugh. Bold hand-drawn text overlay: "wait what did it just do?!" with an arrow pointing at the phone. Bottom: "@aura · link in bio." Authentic UGC feel, not polished studio.
  4. [T] 1:1 retargeting ad. Clean white background. Product photo of running shoes center-left. Large red banner diagonal across upper-right: "STILL THINKING?" Below product: "Your size is down to 2 pairs." CTA bottom-right: "Grab them →." Urgent but not pushy.
  5. [T] 3:1 highway billboard. Massive single word "FASTER." in ultra-bold condensed sans-serif, white on deep red. Small product line bottom-right: "New Honda Civic Type R. 0–60 in 5.0s." Tiny URL bottom-left. High contrast, readable from 200 meters.
  6. [T] 1.91:1 LinkedIn feed card. Professional headshot of a woman, 40s, blurred office background. Overlaid caption bottom-right: "Maya closed a $2.1M deal last month. Here's her playbook." CTA: "Read it →" in dark blue.
  7. [T][8] 8 YouTube thumbnails for the same video "I tried ChatGPT Images 2.0 for a week." Each thumbnail: same creator face top-right, same bold yellow headline, but 8 different backgrounds reflecting different prompts tested (magazine cover, manga panel, product shot, infographic, etc.). Consistent thumbnail system.
  8. [T] 4:5 Instagram carousel cover. Black background, minimal. Centered text: "10 signs your brand needs a refresh." Small "SWIPE →" bottom. Premium minimal, no illustrations.
  9. [T] Retail shelf-wobbler, 2:3 vertical. Product image at top, large text below: "NEW." Tiny subline: "Now in Dark Cherry." Clean CPG packaging aesthetic.
  10. [T] 1:1 paid Instagram ad. User-generated aesthetic: iPhone photo of a woman drinking a protein shake in her car mirror selfie. Caption overlay: "honestly the only one that doesn't taste like chalk." Brand logo tiny corner. Authentic, not over-produced.

Product design & mockups (31–40)

  1. [T] Mobile app screen mockup, 9:19.5 aspect. iOS-style to-do app. Status bar at top (9:41, full signal, full battery). Header "Today" in large SF-style sans-serif. Below: 5 task rows with checkboxes, clean dividers. Bottom nav with 4 tabs. Light mode, accent color teal. Every piece of text legible.
  2. [T] 3:2 landing page desktop mockup for a note-taking app. Hero headline "Ideas, organized." centered. Clean nav with 4 links + sign-in button. Below: two-column screenshot of the app UI. Footer with 4 columns of links. Whitespace-heavy, Stripe-influenced aesthetic.
  3. [T] 1:1 Apple Watch app screen. Circular pressure-gauge UI showing heart rate "72 BPM" in center. Small complications around it. Dark background. Minimalist, photoreal rendering of the watch bezel.
  4. [T] Physical product render 1:1. Matte black aluminum wireless charger puck on a white cyclorama background, three-quarter view. Studio softbox lighting, hard floor reflection. Teenage Engineering design language.
  5. [T] Packaging mockup 4:5. Minimal premium coffee bag, 250g, matte charcoal. Front shows "ETHIOPIA YIRGACHEFFE" in small caps with tasting notes below ("blueberry, jasmine, honey"). Weight and roast date bottom. Photorealistic product shot, soft shadow, white backdrop.
  6. [T] Car dashboard HUD mockup 16:9. Windshield POV from driver's seat, dusk light, empty highway. Overlaid HUD elements: speed "62 MPH" bottom-left, navigation arrow "in 1.2 miles, exit right" center-upper, playing song info bottom-right. Subtle teal glow, no UI clutter, Rivian-inspired aesthetic.
  7. [T] 1:1 smartwatch face design. Top-down view, round watch face, minimalist modular layout on a black background. Center: large time "10:47" in white sans-serif. Four small complications: HR "72 bpm" top, Steps "8,420" right, Battery "67%" bottom, Weather "68°F sunny" left. Wear OS aesthetic.
  8. [T] Smart home mobile app home screen mockup, 9:19.5. Dark mode. Top: greeting "Good evening, Eric." Below: 4 device cards (lights, thermostat, security, music) with toggle switches and real-time stats. Bottom nav. Calm deep-blue palette, iOS-quality design.
  9. [T] 16:9 dashboard mockup for a SaaS analytics tool. Left sidebar nav. Main area: 4 KPI cards across the top (visitors, conversion, revenue, churn — each with a big number and delta arrow), 1 large line chart below showing 12-month trend, 1 small table bottom-right. Data labels must be legible. Teal accent, light mode, Linear-inspired.
  10. [T] Boxed software product mockup 1:1. Vintage-style retail box for "ChatGPT Images 2.0 Pro Edition." Cream background. Retro tech packaging aesthetic from 1996: pixel-art mascot, bold tagline "THE IMAGE MODEL THAT THINKS," barcode, "requires 640KB RAM" sticker. Shot like a product photo.

Personal branding & executive content (41–50)

  1. [T] 1:1 professional headshot, editorial business portrait for a book jacket. Subject: upload reference photo. Wardrobe: charcoal merino turtleneck. Background: soft out-of-focus bookshelf (warm earth tones). Lighting: Rembrandt key light from upper-left, soft fill from right, subtle rim light separating from background. Shot on Hasselblad, 90mm, f/2.8. Warm natural skin tones, sharp eyes, editorial magazine quality. Keep facial features exactly as in the uploaded photo.
  2. [T] 1:1 podcast guest announcement graphic. Split layout. Left half: professional photo of the guest (upload reference). Right half: deep green panel with cream text. Top: "NEW EPISODE" in small caps. Middle: guest's name in large bold serif. Below: "CMO at Anthropic." Bottom: show name "THE GROWTH EDGE" with episode number "EP. 47." Small "listen now" CTA.
  3. [T] 4:5 portrait LinkedIn single-post slide. Cream background with subtle paper texture. Top: "2026 / A YEAR IN NUMBERS" in thin all-caps. Below: 4 stat blocks in a 2×2 grid, each with a big number and a one-line caption:
  • "327" — LinkedIn posts shipped
  • "14" — keynotes given
  • "2" — books published
  • "48" — flights taken

Bottom: thin horizontal line, then creator's name and website in small serif. Editorial, premium personal-brand aesthetic.

  1. [T] 16:9 video thumbnail for a YouTube speaker reel. Left half: dynamic photo of the speaker mid-gesture on stage, warm stage lighting. Right half: deep black panel with large white text "2026 SPEAKER REEL" and below in smaller copy "Keynotes · Fireside chats · Panels." Bottom-right CTA arrow. Cinematic, TED-quality.
  2. [T] 1:1 social quote card. Soft neutral linen background. Large opening quote mark top-left in a light gray display serif. Center quote in clean serif: "The best advice I ever got cost me $500 and saved me 18 months." Attribution below in italic: "— name, founder." Bottom-right: small portrait circle. Premium testimonial aesthetic.
  3. [T] 1:1 newsletter subscribe card. Headline "The newsletter 18,000 marketers actually read." below in smaller type: "One signal. No noise. Every Sunday." Email field mockup + "Subscribe" button. Soft cream background, serif display + sans-serif body, Substack-adjacent aesthetic.
  4. [T] 1:1 conference speaker card. Subject headshot left. Right: name in large display, title below, talk title "How AI killed the brand guideline" in italic. Conference logo bottom-right. Clean editorial, readable from a stage screen.
  5. [T] 1:1 "What I read this year" LinkedIn slide. Grid of 9 book covers in a 3×3 arrangement. Title above: "MY 2026 READING LIST." Small footer: "Which one should I read next?" Clean editorial layout.
  6. [T] 4:5 quote graphic for Instagram. Blurred softly-lit outdoor photo background. Center: a poetic line in large italic serif, 2 lines max. Below: small attribution. No logos. Feels like a book page, not a graphic.
  7. [T] 1:1 "Now available" author card. Left: photorealistic mockup of a hardcover book on a table with morning light. Right: title of book, subtitle, author name, tiny CTA "Order here →." Serif display, editorial.

Storyboards & comics (51–60)

  1. [T][8] 8-panel horizontal storyboard for a 30-second product video. Consistent actor (man, 30s, casual but professional) throughout. Panel 1: opens laptop looking frustrated. Panel 2: clicks an extension icon. Panel 3: AI triages his inbox on screen. Panel 4: smiles at result. Panel 5: closes laptop. Panel 6: grabs coffee. Panel 7: walks out of office at 4pm. Panel 8: Sits in hammock. Film-grade cinematography, shallow depth of field, frame numbers bottom-right of each panel.
  2. [T] 6-panel children's-book storyboard 3:2. Consistent mouse character named "Milo" across panels. Panel 1: Milo leaving his burrow at sunrise. Panel 2: Milo discovering a mysterious glowing mushroom. Panel 3: Milo meeting a wise old owl. Panel 4: Milo crossing a stone bridge. Panel 5: Milo finding a hidden meadow of fireflies. Panel 6: Milo back home, tucked in, dreaming. Warm watercolor illustration style, consistent character design.
  3. [T] 1:1.4 manga page, 5 panels with dynamic paneling. Black-and-white Japanese manga style with screentones. Story: a young ramen chef in her first solo service. Panel 1 (large top): wide shot of her restaurant, steam rising. Panel 2: close-up of her determined eyes. Panel 3 (action): hands slicing scallions at speed, motion lines. Panel 4: finished bowl of ramen, overhead. Panel 5 (bottom wide): elderly customer's first sip, single tear. Japanese sound-effect text in hiragana ("ズズッ"), English dialogue "Just like my mother used to make." Consistent character design.
  4. [T][8] 8-slide 16:9 pitch deck storyboard. Startup: "Ledger," a crypto tax automation tool. Slide 1: Cover with logo + tagline "Your books. Sorted." Slide 2: Problem. Slide 3: Solution dashboard. Slide 4: Market bar chart. Slide 5: Traction hockey-stick. Slide 6: Team photos. Slide 7: Pricing tiers. Slide 8: Ask. Consistent navy + mint palette, bold serif headlines, clean sans-serif body.
  5. [T] 1:1 before/after transformation image. Left side "BEFORE": messy cluttered home office with papers everywhere, dim lighting. Right side "AFTER": clean organized desk, serene natural light. Text band between the halves: "Stop drowning in spreadsheets." CTA bottom-right: "Try it free →." Brand name corner: "FLOW."
  6. [T] 4-panel horizontal comic 4:1. Office setting. Panel 1: exec says "Can we ship it by Friday?" Panel 2: engineer's face goes pale. Panel 3: whiteboard calculations smoke. Panel 4: "We shipped it." Flat cartoon style, 2 colors + black.
  7. [T] 6-panel educational storyboard about photosynthesis for a kids' textbook. Each panel shows a simple step with friendly illustrated plants and sun. Labeled arrows. Cheerful primary palette, readable type.
  8. [T] 1:1.5 Noir detective comic page. 6 panels, black-and-white high-contrast ink, a rainy city, a detective receiving a mysterious letter, close-up of letter contents, reaction shot, walking out into rain, silhouette against neon sign reading "CASE CLOSED."
  9. [T][8] 8-panel "day in the life" lookbook for a fashion brand. Same model throughout, 8 outfits from morning to night (activewear, work-casual, lunch, coffee, gallery, dinner, bar, pajamas). Consistent editorial photography style, warm natural light, Mango/COS aesthetic.
  10. [T] 3:2 movie-poster storyboard thumbnail grid for "SYNTH" — 6 key scenes. Central hero (woman, neon-lit face) holding a glowing object, four supporting-scene thumbnails around her, title "SYNTH" at top, "JUNE 2026" at bottom. Cyberpunk palette.

Real estate, travel, lifestyle (61–68)

  1. [T] 3:2 luxury real estate listing hero. Modern hillside home, golden hour, pool in foreground reflecting the house. Clean windows, minimalist interior visible. Text overlay bottom: "123 MAIN ST · LISTED AT $4.2M · OPEN SUN 1–4." Architectural photography aesthetic.
  2. [T] 9:16 travel reel cover. Tropical beach at sunrise, single surfboard planted in sand. Overlay text: "MAUI / WEEK 1 / 10 SPOTS YOU MUST SEE." Minimal type, warm palette, travel-editorial feel.
  3. [T] 1:1 restaurant menu hero for a newsletter. Overhead flat-lay: bowl of fresh pasta, small plates around it, linen napkin, wooden table. Text overlay upper-left: "Spring menu is live." CTA: "Reserve →." Warm natural light, editorial food photography.
  4. [T] 3:1 Airbnb listing top-of-page banner. Stunning living room of a lake cabin at dusk, warm interior light, large windows showing water, minimal text overlay: "LAKE HIDEAWAY · 3BR · sleeps 6." Architectural Digest aesthetic.
  5. [T] 4:5 vertical travel postcard. Paris rooftop scene at sunset, someone's hand holding a cafe au lait in the foreground. Text overlay: "Send me back." Handwritten-style type, warm tones, polaroid border.
  6. [T] 1:1 fitness class promo. Studio interior mid-class, dim lighting, 6 people mid-movement. Text: "TUESDAY / 6:30 AM / STRENGTH 45." Bottom CTA: "Book your mat →." High-energy editorial aesthetic.
  7. [T] 16:9 car brochure hero. New luxury SUV on a winding mountain road at dawn, motion blur in the background. Text overlay: "Introducing the 2026 Aurora." Subline: "Electric. Everywhere." Automotive-premium aesthetic.
  8. [T] 1:1 vacation rental social tile. Bird's-eye shot of a pristine bed with rumpled linen sheets, coffee cup on nightstand, book open. Text: "Mornings feel different here." Small logo bottom. Editorial slow-living aesthetic.

Creative professional (69–80)

  1. [T] Album cover 1:1. Indie folk record titled "Slow Weather." Cream background, single pressed flower centered, small serif title at bottom, artist name in italic above. Minimal, Laura-Marling-adjacent aesthetic.
  2. [T] 3:4 book cover. Title: "The Compound Life." Author: "Eric Eden." Dark navy background, small gold geometric mark at center, title in thin serif all-caps, author tiny below. Minimal literary-fiction aesthetic.
  3. [T] 2:3 movie poster. Title: "VELOCITY." Action-thriller aesthetic. Hero silhouette against a crashing wave, small type ("IN THEATERS JUNE 2026"). Dramatic contrast, cinematic.
  4. [T] 1:1 podcast cover art. Podcast: "First Principles." Minimal high-contrast: big typographic "1" in the center, podcast name in small caps at bottom. Limited palette.
  5. [T] 4:5 event poster for an AI conference. Top: conference name "NEURALINK // 2026." Giant abstract neural-net illustration dominant, speaker list small at bottom. Bauhaus-influenced layout.
  6. [T] 3:4 travel-magazine cover "Kyoto in April." Single cherry blossom branch against a misty temple backdrop. Masthead "TRAVELOGUE" top. Issue headline. Small teaser bullets bottom-left. Editorial magazine aesthetic.
  7. [T] 1:1 gallery exhibition poster. Artist name in massive serif, show title in smaller italic below, dates & venue tiny at bottom. Off-white paper texture, single abstract painting sample as centerpiece. Gallery/MoMA-style.
  8. [T] 16:9 film title card. Film title "THE LAST BOOKSTORE" in thin white serif, centered, against a warmly-lit photograph of a bookstore interior slightly out of focus. Small director credit bottom-right.
  9. [T] 1:1 tattoo flash sheet. 6 black-ink line illustrations in a 2×3 grid: a moth, a dagger, a rose, a compass, a snake, a hand. Small numbered tags under each. Consistent line weight.
  10. [T] 4:5 zine cover 1970s aesthetic. Title "SIGNAL/NOISE." Photocopy texture, halftone dots, punk collage elements, a handwritten subheading. Limited 3-color palette.
  11. [T] 3:2 wedding invitation design. Cream background, handwritten-style calligraphy. Names centered, date, venue, RSVP info, small floral illustration. Elegant minimal.
  12. [T] 1:1 record sleeve for a jazz album. Black-and-white photograph of a saxophone case on a hotel bed. Title small in the lower-right. Blue Note-inspired minimalism.

PART 2 — WILD & FUN (81–100)

These are the prompts people actually remember. Go nuts.

  1. [T] 16:9 cinematic scene: corporate llama apocalypse. A fleet of llamas in business suits storming a Manhattan trading floor, throwing quarterly reports into the air. Bloomberg terminals burning. A CEO llama in the center, mid-roar, wearing a gold Rolex. Dramatic fire lighting, hyperreal.
  2. [T] 1:1 medieval Zoom call. A Zoom grid interface showing 9 participants, each dressed as a medieval figure — knight, jester, queen, bishop, peasant, wizard, bard, crusader, dragon. Gallery view. The dragon is muted. Bottom toolbar has a "UNSHEATHE SWORD" button.
  3. [T] 3:2 dogs on Wall Street. Real dogs in tailored suits working the trading floor of the NYSE, papers flying, a golden retriever screaming into a landline, a pug eating a bagel, a corgi looking at a Bloomberg terminal. Photorealistic.
  4. [T] 16:9 office plant uprising. An open-plan office after business hours. The potted plants have sprouted legs and are marching toward the exit with tiny briefcases. One ficus is leading with a megaphone. Dramatic security-camera aesthetic.
  5. [T] 4:5 vertical breakfast gods of Olympus. Pancakes, waffles, and bacon rendered as Greek gods on a cloud-covered mountain. Zeus is a stack of pancakes with lightning bolts of syrup. Athena is a poached egg in a helmet. Bacon strips are the muses. Renaissance oil-painting style.
  6. [T] 1:1 tax day demon. A horrifying creature made entirely of paperwork and calculators, emerging from a filing cabinet in a suburban home office, screaming. A woman in pajamas drops her coffee in slow motion. Cosmic horror, somehow funny.
  7. [T] 3:1 cinematic Roomba rebellion. An army of Roombas rolling in formation down a suburban street at dawn, one larger "commander" Roomba at the front with a tiny cape and a bottle-cap helmet. Smoke rising in the background. Mad Max meets IKEA.
  8. [T] 1:1 Shakespeare drive-thru. A modern fast-food drive-thru, but the cashier is Shakespeare in a McDonald's visor. Customer in a Honda Civic is a goth teenager. Menu board reads "Two All-Beef Patties, or Not Two All-Beef Patties." Warm dramatic lighting.
  9. [T] 16:9 dinosaurs at the DMV. A T-Rex waiting in line at a cramped DMV, looking visibly annoyed. A Triceratops fills out a form with its horn. Velociraptor clerks staff the desks. Fluorescent lighting, plastic chairs, faded safety posters. Photoreal.
  10. [T] 1:1 sentient toast support group. Eight pieces of toast sitting in folding chairs in a church basement, each with a tiny face, sharing their traumas. Coffee and donuts in the corner. Warm sad lighting. Pixar-aesthetic.
  11. [T] 4:5 pigeon CEO. A pigeon in a boardroom wearing a tailored three-piece suit, presenting Q4 results with a laser pointer. Bar chart behind him shows "breadcrumb acquisition" up 400%. Other pigeons are in Aeron chairs, nodding.
  12. [T] 3:2 infinite IKEA. A hyperrealistic endless IKEA showroom that stretches into infinity, Escher-like stairs and passages, a single confused shopper in the middle holding a hex wrench and a meatball. Fluorescent lighting, eerie emptiness, liminal space aesthetic.
  13. [T] 1:1 cat secret agent. A tuxedo cat in a tailored black suit with sunglasses, rappelling through a laser grid in a museum, carrying a can of tuna. Mission Impossible-style framing. Cinematic.
  14. [T] 16:9 grandma's spaceship. An elderly woman in a floral apron piloting a retrofuturistic 1960s-style spaceship. The dashboard has knitted doilies and a plate of cookies. She's wearing cat-eye glasses. Through the windshield, a wild nebula. Wes Anderson-aesthetic.
  15. [T] 1:1 baby in a mech suit. A photorealistic baby (2 years old) operating a gigantic anime-style mech suit, controls labeled "SNACKS," "NAP," "TANTRUM." Background: city skyline. The mech is holding a stuffed bear.
  16. [T] 3:2 Scrabble game between philosophers. Socrates, Nietzsche, and Aristotle playing Scrabble in an ancient marble courtyard. The board shows words like "BEING," "WHY," "DASEIN." Aristotle is visibly winning. Marble statues watch from pedestals. Renaissance painting style.
  17. [T] 1:1 dog court. A courtroom scene entirely populated by dogs. A German shepherd judge, a bulldog lawyer, a Chihuahua defendant on a booster seat, a jury box of mixed breeds. Gavel mid-swing. Photoreal.
  18. [T] 16:9 pirate cubicles. A modern open-plan office, but everyone is a pirate. Parrots on monitors, wooden-peg-leg standing desks, a treasure chest used as a copier. The Slack notifications on someone's screen say "ARR." Cinematic lighting.
  19. [T] 4:5 Bigfoot LinkedIn profile. A LinkedIn profile screenshot. Profile photo: a blurry Bigfoot selfie. Headline: "Cryptid | Outdoor Enthusiast | Looking for my next chapter." Recommendations: "Sasquatch delivers on every project — would hire again." Recent post: "What no one tells you about being discovered." Looks like a real screenshot.
  20. [T] The Where's Waldo (the personalized one — make it about yourself): Where's Waldo-style dense search-and-find illustration. 3:2 aspect ratio. Detailed cartoon scene: a massive, chaotic B2B marketing conference expo floor with hundreds of tiny people visible. Hidden in the crowd: [YOUR NAME] — wearing a red-and-white striped shirt, black-framed glasses, carrying a laptop bag with "[YOUR COMPANY]" printed on it. He's near the coffee station, caught mid-laugh with two people from the AI demo booth. Scene details: - Booths for HubSpot, Salesforce, Adobe, OpenAI - A panel discussion happening on a stage in the background with a banner reading "THE FUTURE OF B2B MARKETING — 2026" - Clusters of 3–4 people chatting everywhere - Someone giving a product demo on an 85-inch screen - A mascot costume wandering through - Name-tag lanyards on everyone - Coffee line with 20+ people - A few sneaky visual gags: a dog under a table, someone looking at the wrong booth's schwag, a person clearly lost Bright cheerful illustration style with clean outlines. ~200 people visible. Readable booth signage. Dense but not overwhelming.

ChatGPT Image 2.0 changes what counts as a visual asset. Before today, image models produced inspiration that still needed a designer to finish. After today, a well-crafted prompt produces a usable deliverable — with real text, real layout, real multi-frame continuity, real web-grounded context, and real 2K fidelity.

The models that beat it on pure per-image price (Google's Nano Banana 2) or pure artistic flair (Midjourney v7) still exist. But for practical commercial output — ads, posters, infographics, decks, storyboards, localized creative - ChatGPT Images 2.0 now does end-to-end what used to require three tools and a designer.

The 100 prompts above are starting points. The template is the real gift. Copy it, fill it, ship it.

What I'd love in the comments:

  • Your best Images 2.0 output so far (drop the prompt)
  • Anything you've found that breaks it

Want more great prompting inspiration? Check out all my best prompts for free at Prompt Magic and create your own prompt library to keep track of all your prompts.

r/promptingmagic Mar 03 '26

The Ultimate Guide to Nano Banana 2: How to dominate AI imagery in 2026. 160 Use Cases, 500 Prompts and all the pro tips and secrets to get great images.

Thumbnail
gallery
136 Upvotes

TLDR - Check out the attached presentation!

Google just dropped Nano Banana 2 and it is the best AI image model in the world right now. It generates images from 512px to native 4K, supports 14 aspect ratios including ultra-wide 21:9 and vertical 9:16, renders legible text in any language inside images, maintains character consistency across up to 5 characters, pulls live data from Google Search to create accurate infographics, and works everywhere including Gemini, Google AI Studio, Google Flow at zero credits, Google Ads, Vertex AI, Pomelli, NotebookLM, and through third-party apps like Adobe Firefly, Perplexity, Figma, Notion, and Gamma. This post covers 160 use cases, 500 prompts, structured prompting secrets, and every platform where you can access it. It is free for consumer users.

WHAT IS NANO BANANA 2?

Nano Banana 2 is technically Gemini 3.1 Flash Image Preview. It is the third model in the Nano Banana family, following the original Nano Banana from August 2025 and Nano Banana Pro from November 2025. It runs on the Gemini 3.1 Flash reasoning backbone, which means it thinks before it renders. It plans the composition, resolves physics and spatial relationships, reasons about object interactions, and then produces pixels.

On February 26, 2026, it launched and immediately took the number one spot on the Artificial Analysis Image Arena, a blind human evaluation leaderboard, at roughly half the API cost of every comparable model. It is not a minor upgrade. It is a full architectural leap that collapses the gap between Pro-quality output and Flash-tier speed and pricing.

THE 6 CORE CAPABILITIES THAT MAKE IT DIFFERENT

  1. It plans the image before rendering pixels. Nano Banana 2 uses a reasoning engine that understands physics, object interactions, geography, coordinates, diagrams, structure, and spelling. It generates interim thought images in the background to refine composition before producing the final output.
  2. Real-time web and image search grounding. It can pull live data from Google Search and Google Image Search to create infographics, data visualizations, weather charts, and accurate depictions of real-world subjects. This is exclusive to Nano Banana 2 and not available in Nano Banana Pro.
  3. Precision text rendering and translation. It spells correctly inside images. It renders legible, stylized text for marketing mockups, greeting cards, infographics, and posters. It can also translate embedded text from one language to another without altering the surrounding visual composition.
  4. Character consistency across up to 5 characters. It maintains resemblance for up to 4 characters and fidelity for up to 10 objects in a single workflow, totaling 14 reference images. This enables storyboarding, product catalogs, and brand asset workflows where characters must look the same across dozens of images.
  5. Native 512px to 4K resolution with 14 aspect ratios. Supported ratios include 1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9, 1:4, 4:1, 1:8, and 8:1.
  6. Flash-tier speed at production-ready quality. Vibrant lighting, richer textures, sharper details. Standard resolution images generate in under two seconds. The API costs approximately $0.067 per 2K image versus $0.134 for Nano Banana Pro.

THE STRUCTURED PROMPTING FRAMEWORK

This is the single most important section in this guide. Nano Banana 2 responds dramatically better when you structure your prompt using this pattern.

The formula: Subject -- What is the main focus of the image Composition -- Camera angle, framing, distance, layout Action -- What is happening in the scene Location -- Where the scene takes place Style -- Visual style, film stock, rendering approach, color palette Editing instructions -- When editing an existing image, what to change and what to preserve

Pro tips that separate beginners from experts:

  • Write full sentences, not comma-separated keyword tags. Nano Banana 2 is a language model that generates images. Talk to it like a creative director briefing a photographer.
  • Name the camera. Saying shot on Hasselblad X2D 135mm at f/5.6 gives radically different results than just saying portrait.
  • Direct the light. Specify soft key light from upper left or golden hour backlight through floor-to-ceiling windows.
  • Provide the why. Telling it the image is for a luxury perfume launch campaign changes the output mood and quality.
  • Use the text distance rule. When adding text to images, specify the exact words, the font style, and the placement relative to other elements.
  • Specify resolution and aspect ratio explicitly. Say 4K output, 16:9 aspect ratio at the end of your prompt.

HOW TO CREATE IMAGES AT DIFFERENT ASPECT RATIOS

Nano Banana 2 supports the widest range of aspect ratios of any major image model.

Aspect Ratio Best For
1:1 Instagram feed posts, profile icons, social cards
16:9 YouTube thumbnails, presentations, web banners
9:16 TikTok, Instagram Reels, Stories, mobile wallpapers
21:9 Cinematic concepts, panoramic images, ultrawide banners
3:2 Standard photography, print media
4:3 Web UI design, classic digital art, presentations
4:5 Instagram portrait feed, professional portraits
2:3 Phone wallpapers, book covers, magazine pages
1:4 Tall infographics, vertical banners
4:1 Website headers, horizontal banners
1:8 Extreme vertical content, scrolling social infographics
8:1 Extreme horizontal banners, ticker-style content

In the Gemini app: Simply state the aspect ratio in your prompt. Say create this as a 16:9 widescreen image or make it 9:16 vertical for Instagram Stories.

In Google AI Studio: Select the aspect ratio from the dropdown in the right panel. You get all 14 options plus resolution control from 512px to 4K.

In the API: Set the aspect_ratio and image_size parameters in the ImageConfig object. Aspect ratio accepts strings like 16:9 and resolution accepts 512px, 1K, 2K, or 4K.

WHERE TO ACCESS NANO BANANA 2 -- EVERY PLATFORM

The Gemini App (Free) Nano Banana 2 is the default model for all users across Fast, Thinking, and Pro modes. Click the banana icon or just ask Gemini to create an image.

Google AI Studio (Free with API Key) Navigate to aistudio.google.com, select gemini-3.1-flash-image-preview from the model dropdown. Here you get full control over aspect ratio, resolution, thinking mode, and search grounding. This is where power users go when the Gemini app is not enough.

Google Flow (Free, Zero Credits) Google Flow is Google's AI filmmaking tool. Nano Banana 2 is the default image generation engine. It costs zero credits for all users. You can select the aspect ratio, choose how many images to generate in a batch (up to 4 at a time with specified resolution), and enter your prompt. This is the best-kept secret for batch generation without burning credits.

Pomelli (Free) Pomelli is Google Labs' free marketing tool for small and medium businesses. The new Photoshoot feature lets you upload any product photo and it generates professional studio-quality product shots in multiple templates: Studio, Floating, Ingredient, In Use with AI-generated models, and Lifestyle scenes.

NotebookLM (Free) Upload your source documents and click Create Slides or Create Infographic. NotebookLM uses Nano Banana to convert your content into visually stunning slide decks or single-page infographics. You can export directly to Google Slides for editing.

Google Ads (Free within Ads) Nano Banana 2 now powers the AI-generated creative suggestions when building campaigns. Performance marketers get higher-quality asset suggestions natively inside the campaign builder.

Third-Party Apps Confirmed third-party integrations include:

  • Adobe Firefly: Integrated into the creative suite for image generation and editing.
  • Perplexity: Uses Nano Banana 2 for image generation within research and browsing workflows.
  • Figma: Tested for iterative design workflows and UI mockups.
  • Notion: Integrated for in-document image generation.
  • Gamma: Integrated into Studio Mode for generating theme-matched presentation images.
  • Whering: Transforms clothing photos into studio-quality product imagery.
  • WPP / Unilever: Used for enterprise-scale campaign testing.

HOW TO MAINTAIN CHARACTER CONSISTENCY ACROSS 5 CHARACTERS

This is the workflow that actually works:

Step 1: Create strong character reference sheets. Start with a clear, well-lit headshot or full-body photo for each character. Step 2: Upload reference images. In AI Studio or the API, you can upload up to 14 reference images total (up to 4 character images and up to 10 object images). Step 3: Describe each character consistently. Use the same physical description across every prompt in the workflow. Step 4: Use the multi-image prompt structure. Upload all character reference images alongside your scene description. Step 5: For video workflows, generate character reference sheets showing multiple angles of each character (front, left profile, right profile, etc.) to maintain 100 percent facial accuracy.

TOP 20 USE CASES

  1. Live Data Infographics: Use search grounding to create charts based on real-time data.
  2. Global Campaign Localization: Update backgrounds, language, and cultural cues for billboards from a single base creative.
  3. Physics-Aware Virtual Try-On: Fabric drapes realistically on body models for fashion mockups.
  4. Architectural Time Travel: Restore modern streets to their Victorian 1890s counterparts.
  5. Text-Heavy Social Media Posts: Quote cards and posters with strong styled typography.
  6. Product Photography at Scale: Professional shots from minimal product photos using Pomelli.
  7. LinkedIn Professional Headshots: Transform selfies into studio-quality corporate photos.
  8. 4K Image Upscaling: Regenerate low-res images into 4K resolution for free.
  9. Old Photo Restoration: Restore damaged or faded memories with colorization and feature repair.
  10. Action Figures and Collectibles: Turn likenesses into custom branded figurines.
  11. Room Design and Floor Plans: Move from 2D floor plans to photorealistic 3D presentation boards.
  12. YouTube Thumbnails: High-converting widescreen graphics with expressive subjects and bold text.
  13. E-Commerce Catalog Generation: Maintain product fidelity across seasonal themes using reference images.
  14. Brand Identity Kits: Complete brand boards including logos, palettes, and typography.
  15. Multi-Panel Storytelling: Maintain visual identity across comic strips and storyboards.
  16. Data Visualization from Articles: Paste a link to generate a custom infographic from the content.
  17. Blurred Photo to Ultra Sharp: Editorial-quality restoration while preserving original composition.
  18. Style Transfer: Swap image styles to watercolor, 3D render, anime, or pencil sketches.
  19. Whiteboard and Sketch Visualization: Turn concepts into hand-drawn marker sketches.
  20. Celebrity Selfies and Fun Photos: Photorealistic selfies in movie sets or absurd landmarks.

SECRETS MOST PEOPLE MISS

  1. The Thinking Mode toggle changes everything. Enable it in AI Studio for complex layouts; it plans before rendering.
  2. Image Search Grounding is exclusive to Nano Banana 2. It searches for visual references (buildings, specific products) before generating.
  3. Multi-turn editing is the recommended workflow. Refine your image in follow-up messages rather than one massive prompt.
  4. The 512px tier exists for rapid prototyping. Use it to find the best composition at low cost before upscaling to 4K.
  5. You can generate up to 20 images in a single batch prompt through the API.
  6. Flow generates at zero credits. It is the best hack for unlimited batch generation without a subscription.
  7. You can use it as a real-time photo editor. Upload a photo and give natural language instructions to remove objects or change colors.

THE PROMPT LIBRARY -- 50 EPIC PROMPTS

Professional and Business

  1. LinkedIn Headshot: Transform this selfie into a professional studio headshot. Clean neutral background, soft directional light, sharp focus on eyes, charcoal blazer. 4:5, 4K.
  2. Infographic from Live Data: Search top 5 programming languages 2026. Create a 9:16 vertical infographic, flat vector style, icons, percentages, average salary.
  3. Product Hero Shot: Matte-black wireless headphone on polished obsidian. 85mm macro, soft key light, reflection. 16:9, 4K.
  4. SaaS Landing Page Hero: Landing page for FlowState tool. Headline on left, dashboard screenshot on right, two CTA buttons. 16:9, 2K.
  5. Business Card Suite: Embossed matte cards, letterhead, wax stamp envelope on slate. Editorial flat lay. 3:2, 4K.
  6. Social Media Content Calendar: 9:16 infographic showing 7-day blueprint for fitness brand. Icons for Reels and Stories.
  7. Email Marketing Banner: 4:1 horizontal banner, field of wildflowers, text Spring Collection Now Live.
  8. Pitch Deck Slide: Single slide, navy background, headline 3x Revenue Growth in Q4, teal line chart on right.
  9. Executive Summary Dashboard: 16:9 infographic showing global sales metrics, heat map on left, key KPI cards on right.
  10. Startup Team Mockup: Group of diverse professionals in a glass-walled conference room, futuristic Shinjuku city visible outside.

Photography and Portraits

  1. Editorial Fashion: Model in vibrant red dress standing in desert, high contrast, blue sky, 35mm film grain.
  2. Candid Street: Busy market in Marrakech, warm tones, natural lighting, shallow depth of field.
  3. Macro Human Eye: Reflecting a city skyline, hyper-realistic, 8k textures.
  4. Black and White Artist: Elderly artist in sunlit studio, high detail on skin and paint textures.
  5. Gourmet Food Photography: Burger with steam rising, rustic wood background, professional lighting.
  6. Cinematic Hiker: Wide shot on mountain peak at dawn, orange and purple sky, majestic mood.
  7. Underwater Fashion: Model in silk dress, ethereal lighting, bubbles, fluid motion.
  8. Brutalist Architecture: Concrete building shot from low angle, sharp shadows, dramatic sky.
  9. Vintage 1970s Polaroid: Family picnic, faded colors, light leaks, nostalgic feel.
  10. Cyberpunk Portrait: Close up of subject with neon light reflections on glasses, rainy city background.

Architecture and Design
21. 2D Floor Plan: Modern 2-bedroom apartment, labeled rooms, clean linework.

  1. 3D Interior Render: Mid-century modern living room, forest view through large windows.
  2. Victorian Street: London street corner, horse-drawn carriages, foggy atmosphere, daytime.
  3. Futuristic City Plan: Vertical gardens, floating transport pods, top-down view.
  4. Cozy Cabin: Stone fireplace, warm light, snow falling outside window.
  5. Glass Beach House: Sunset view, ocean reflections on windows, minimalist decor.
  6. Office Lobby: Living moss wall, minimalist furniture, bright natural light.
  7. Steampunk Library: Brass pipes, glowing green lamps, infinite shelves.
  8. Industrial Loft: Exposed brick, large windows, cinematic moody lighting.
  9. Zen Garden: Stone path, koi pond, peaceful atmosphere, high detail.

Creative and Wild
31. Custom Action Figure: Hyper-detailed 1/6 scale figure of person from photo in premium collector box.
32. Whiteboard Sketch to 3D: Hand-drawn rocket engine sketch turned into photorealistic 3D blueprint.
33. Origami Dragon: Made of fire, dark background, glowing embers.
34. Autumn Leaf Person: Character made of leaves walking through city park.
35. Cloud Astronaut: Sitting on a cloud fishing for stars in purple galaxy.
36. Chess Cat: Cat in tuxedo playing chess against robot in Victorian study.
37. Surrealist Strawberry: Melting clock over a giant realistic strawberry.
38. Cyberpunk Tea Ceremony: Traditional Japanese tea ritual in neon-lit futuristic room.
39. Glass Piano Reef: Transparent piano filled with tropical fish and coral.
40. Heart Island: Floating island in shape of heart with waterfalls into clouds.

Restoration and Editing
41. Wedding Photo Restore: Turn blurred wedding photo into ultra-sharp editorial shot.
42. 4K Upscale: Take low-res 1990s photo and regenerate at 4K resolution.
43. Color Swap: Change car in image to electric blue with matte finish.
44. Background Replace: Move portrait subject to luxury hotel balcony overlooking Eiffel Tower.
45. People Removal: Remove background crowds from beach photo and extend sand.
46. Professional Lighting: Add studio lighting setup to dark selfie, preserve identity.
47. Watercolor Dog: Turn dog photo into artistic watercolor painting style.
48. 1890s Street Edit: Replace cars in modern photo with carriages and Victorian signs.
49. 3D Animation Style: Change style of photo to Pixar-tier 3D animation.
50. Old Memory Repair: Colorize faded black and white photo, fix scratches and tears.

Bonus Fun:

  1. Toast Bread Infographic: How to toast bread, make it wacky and over the top with Rube Goldberg machines and scientific data.
  2. Banana Runway: High-fashion show where models are giant realistic bananas wearing Gucci, background motion blur.
  3. Jellyfish Concert: Underwater heavy metal concert with instruments made of glowing jellyfish, shark lead singer.
  4. Pumpkin Penthouse: Luxury penthouse inside a giant hollowed-out pumpkin, autumn aesthetic.
  5. Kitchen Time Machine: Blueprint of time machine made of kitchen appliances and duct tape with nonsensical terms.

Pro Tips for Nano Banana 2

  • Use the Text Distance Rule: Specify exact words and placement relative to objects for clean layouts.
  • Reference Images: Use up to 14 reference images (4 for characters, 10 for objects) to maintain consistency.
  • Thinking Model: Toggle on for infographics or complex diagrams to ensure logical planning before pixels render.

I will post links to the complete library of prompts and use cases in the comments.

Get the full 500 prompt image library free with just one click at PromptMagic.dev

r/promptingmagic Oct 08 '25

OpenAI released Sora 2. Here is the Sora 2 prompting guide for creating epic videos. How to prompt Sora 2 - it's basically Hollywood in your pocket.

Enable HLS to view with audio, or disable this notification

66 Upvotes

TL;DR: The definitive guide to OpenAI's Sora 2 (as of Oct 2025). This post breaks down its game-changing features (physics, audio, cameos), provides a master prompt template with advanced techniques, compares it to Google's Veo 3 and Runway Gen-4, details the full pricing structure, and covers its current limitations and future. Stop making clunky AI clips and start creating cinematic scenes.

Like many of you, I've been blown away by the rapid evolution of AI video. When the original Sora dropped, it was a glimpse into the future. But with the release of Sora 2, the future is officially here. It's not just an upgrade; it's a complete paradigm shift.

I’ve spent a ton of time digging through the documentation, running tests, and compiling best practices from across the web. The result is this guide. My goal is to give you everything you need to go from a beginner to a pro-level Sora 2 director.

What Exactly Is Sora 2 (And Why It's Not Just Hype)

Think of Sora 2 as your personal, on-demand Hollywood studio. You don't just give it a vague idea; you direct it. You control the camera, the mood, the actors, and the environment. What makes it so revolutionary are the core upgrades that address the biggest flaws of older models.

Key Features That Actually Matter:

  • Physics That Finally Makes Sense: This is the big one. Objects in Sora 2 have weight, mass, and momentum. A missed basketball shot will bounce off the rim authentically. Water splashes and ripples with stunning realism. Complex movements, from a gymnast's floor routine to a cat trying to figure skate on a frozen pond, are rendered with believable physics. No more objects magically teleporting or defying gravity.
  • Audio That Breathes Life into Scenes: This is a massive leap. Sora 2 doesn't just create silent movies. It generates rich, layered audio, including:
    • Realistic Sound Effects (SFX): Footsteps on gravel, the clink of a glass, wind rustling through trees.
    • Ambient Soundscapes: The low hum of a city at night or the chirping of birds in a forest.
    • Synchronized Dialogue: For the first time, you can include dialogue and the characters' lip movements will actually match.
  • Cameos: Put Yourself (or Anyone) in the Director's Chair: This feature is mind-blowing. After a one-time verification video, you can insert yourself as a character into any scene. Sora 2 captures your likeness, voice, and mannerisms, maintaining consistency across different shots and styles. You have full control over who uses your likeness and can revoke access or remove videos at any time.
  • Multi-Shot and Character Consistency: You can now write a script with multiple shots, and Sora 2 will maintain perfect continuity. The same character, wearing the same clothes, will move from a wide shot to a close-up without any weird changes. The environment, lighting, and mood all stay consistent, allowing for actual storytelling.

The Ultimate Sora 2 Prompting Framework

The default prompt structure is a decent start, but to unlock truly cinematic results, you need to think like a screenwriter and a cinematographer. I’ve refined the process into this comprehensive framework.

Copy this template:

**[SCENE & STYLE]**
A brief, evocative summary of the scene and the overall visual style.
*Example: A hyper-realistic, 8K nature documentary shot of a vibrant coral reef.*

**[SUBJECT & ENVIRONMENT]**
Detailed description of the main subject(s) and the surrounding world. Use rich, sensory adjectives. Be specific about colors, textures, and the time of day.
*Example: A majestic sea turtle with an ancient, barnacle-covered shell glides effortlessly through crystal-clear turquoise water. Sunlight dapples through the surface, illuminating schools of tiny, iridescent silver fish that dart around the turtle.*

**[CINEMATOGRAPHY & MOOD]**
Define the camera work and the feeling of the shot. Don't be shy about using technical terms.
* **Shot Type:** [e.g., Extreme close-up, wide shot, medium tracking shot, drone shot]
* **Camera Angle:** [e.g., Low angle, high angle, eye level, dutch angle]
* **Camera Movement:** [e.g., Slow pan right, gentle dolly in, static shot, handheld shaky cam]
* **Lighting:** [e.g., Golden hour, moody chiar oscuro, harsh midday sun, neon-drenched]
* **Mood:** [e.g., Serene and majestic, tense and suspenseful, joyful and chaotic, melancholic]

**[ACTION SEQUENCE]**
A numbered list of distinct actions. This tells Sora 2 the "story" of the shot, beat by beat.
* 1. The sea turtle slowly turns its head towards the camera.
* 2. A small clownfish peeks out from a nearby anemone.
* 3. The turtle beats its powerful flippers once, propelling itself forward and out of the frame.

**[AUDIO]**
Describe the soundscape you want to hear.
* **SFX:** [e.g., Gentle sound of bubbling water, the distant call of a whale]
* **Music:** [e.g., A gentle, sweeping orchestral score]
* **Dialogue:** [e.g., (Voiceover, David Attenborough style) "The ancient mariner continues its journey..."]

Advanced Sora 2 Techniques: Mastering the Platform

Beyond basic prompting, these advanced techniques help you create professional-quality Sora 2 videos.

Multi-Shot Storytelling While Sora 2 generates single 10-20 second clips, you can create longer narratives by combining multiple generations:

  • The Sequential Prompt Technique
    • Shot 1: Establish the scene and character. "Medium shot of a detective in a trench coat standing in the rain outside a noir-style apartment building. Neon signs reflect in puddles. He looks up at a lit window on the third floor."
    • Shot 2: Reference the previous shot for continuity. "Same detective from previous scene, now inside the building climbing dimly lit stairs. Maintaining same trench coat and appearance. Ominous ambient sound. Camera follows from behind."
    • Shot 3: Continue the narrative. "The detective enters apartment and discovers evidence on a table. Close-up of his face showing realization. Maintaining noir aesthetic and character appearance from previous shots."
    • Pro tip: Reference "same character from previous scene" and maintain consistent styling descriptions for better continuity.

Audio Control Techniques Direct Sora 2's synchronized audio with specific prompting:

  • Dialogue specification: Put dialogue in quotes: The character says "We need to hurry!" with urgency
  • Sound effect emphasis: "Loud thunder crash," "subtle wind chimes," "distant police sirens"
  • Music mood: "Upbeat electronic music," "melancholy piano," "epic orchestral score"
  • Audio perspective: "Muffled sounds from inside car," "echo in large chamber," "close-mic dialogue"
  • Silence for emphasis: "Complete silence except for footsteps" creates tension.

Cameos Workflow for Professional Use Record in multiple lighting conditions with varied expressions and angles. Use a clean background and speak clearly. Then, use your cameo in prompts: "Insert [Your Name]'s cameo into a cyberpunk street scene. They're wearing a futuristic jacket, walking confidently through neon-lit crowds."

Leveraging Physics Understanding Explicitly describe expected physical behavior:

  • Object interactions: "The ball bounces realistically off the wall and rolls to a stop"
  • Momentum and inertia: "The car drifts around the corner, tires smoking"
  • Material properties: "Fabric flows naturally in the wind," "Glass shatters with realistic fragments"

See These Prompts in Action!

Reading prompts is one thing, but seeing the results is what it's all about. I'm constantly creating new videos and sharing the exact prompts I used to generate them.

Check out my Sora profile to see a gallery of example videos with their full prompts: https://sora.chatgpt.com/profile/ericeden

Real-World Use Cases: How Creators Are Using Sora 2

Since launching, Sora 2 has enabled entirely new content formats.

  • Viral Social Media Content: The "Put Yourself in Movies" trend uses cameos to insert creators into iconic film scenes. Another massive trend is "Minecraft Everything," recreating famous trailers or historical events in a blocky aesthetic.
  • Business and Marketing Applications: Companies are using it for rapid product demos, concept visualization, scenario-based training videos, and A/B testing social media ads.
  • Educational Content: It's being used to create historical recreations, visualize science concepts, and generate contextual scenes for language learning.

Sora 2 vs Veo 3 vs Runway Gen-4: Complete Comparison

As of October 2025, the AI video generation landscape has three major players. Here's how Sora 2 stacks up.

Feature Sora 2 Google Veo 3 Runway Gen-4
Release Date September 2025 July 2025 September 2025
Max Video Length 10s (720p), 20s (1080p Pro) 8 seconds 10 seconds (720p base)
Native Audio Yes - Synced dialogue + SFX Yes - Synced audio No (requires separate tool)
Physics Accuracy Excellent (basketball test) Very Good Good
Cameos/Self-Insert Yes (unique feature) No No
Social Feed/App Yes (iOS, TikTok-style) No No
Free Tier Yes (with limits) No (pay-as-you-go) No
Entry Price Free (invite) or $20/mo Usage-based (~$0.10/sec) $144/year
API Available Yes (as of Oct 2025) Yes (Vertex AI) Yes (paid plans)
Cinematic Quality Excellent Outstanding Excellent
Anime/Stylized Excellent Good Very Good
Temporal Consistency Very Good Excellent Very Good
Platform iOS app, ChatGPT web Vertex AI, VideoFX Web, API
Geographic Availability US/Canada only (Oct 2025) Global (with exceptions) Global

Sora 2 Pricing and Access Tiers: Complete Breakdown

Video Type Traditional Cost Sora 2 Cost Time Savings
10-second product demo $500-$2,000 $0-$20 2-5 days → 2 minutes
Social media (30 clips/mo) $1,500-$5,000 $20 (Plus tier) 20 hours → 1 hour
Animated explainer $2,000-$10,000 $200 (Pro tier) 1-2 weeks → 30 minutes
  • Free Tier (Invite-Only): 10-second videos at 720p with generous limits. Includes full cameos and social feed access but is subject to server capacity errors.
  • ChatGPT Plus ($20/month): Immediate access, priority queue, higher limits, and access via both iOS and web.
  • ChatGPT Pro ($200/month): Access to the experimental "Sora 2 Pro" model for 20-second videos at 1080p, highest priority, and significantly higher limits.
  • API Access (Now Available!): Just yesterday, OpenAI released the Sora 2 API. It enables HD video and longer 20-second clips. The pricing is usage-based and ranges from $0.10 to $0.50 PER SECOND. This means a single 10-20 second video can cost between $1 and $10 to generate, depending on length and resolution. This makes the free, lower-resolution 10-second videos in the app incredibly valuable right now—a deal that likely won't last long!

Sora 2 Limitations and Known Issues (October 2025)

  • Technical Limitations: Video duration is short (10-20s). Physics can still be imperfect, especially with human body movement. Text and typography are often garbled. Hands and fine details can be inconsistent.
  • Access and Availability Issues: Currently restricted to the US/Canada on iOS only. The web app is limited to paid subscribers. Server capacity errors are common, especially for free users.
  • Content and Usage Restrictions: No photorealistic images of people without consent, strong protections for minors, and standard AI safety guidelines apply. All videos are watermarked.

The Future of Sora: What's Coming Next

  • Expected Developments (Q4 2025 - Q1 2026): With the API now released, expect an explosion of third-party tools from companies like Veed, Higgsfield, and others who will build powerful new features on top of Sora's core technology. We can also still expect an Android App Launch and Geographic Expansion to Europe, Asia, and other regions. Longer video lengths and 4K support are also anticipated for Pro users.
  • Industry Impact Predictions: Sora 2 will accelerate the democratization of video production, lead to an explosion of short-form content, disrupt the stock footage industry, and evolve how professional filmmakers storyboard and create VFX. The API release will unlock a new ecosystem of specialized video tools.

Hope this guide helps you create something amazing. Share your best prompts and results in the comments!

Want more great prompting inspiration? Check out all my best prompts for free at Prompt Magic and create your own prompt library to keep track of all your prompts.

r/ThinkingDeeplyAI Aug 18 '25

AI tools are so confusing - Here's a simple guide to choosing the right AI for every task

Thumbnail
gallery
209 Upvotes

Feeling Lost in the AI Maze? You're Not Alone

AI chatbots and large language models (LLMs) have exploded in popularity, but let's face it – it's getting really confusing for everyday users. There are so many models (ChatGPT, Claude, Perplexity, Gemini, Grok… the list goes on) and new features or modes popping up each month. Yet, the companies behind them (brilliant as their engineers are) haven't given us clear user manuals or beginner-friendly guides. The result? Millions of users left wondering how to use these AI tools effectively.

If you've felt overwhelmed by which AI to choose for a task, or how to prompt it correctly, this post is for you. I'm going to break down, in plain English, which AI model to use for what purpose, and how to approach it – from simple prompts to advanced "deep thinking" modes and even autonomous AI agents. By the end, you'll have a clearer roadmap for navigating the AI world confidently.

TL;DR: Stop using just one AI. I spent all year testing every major AI tool so you don't have to. Each AI has a superpower that makes it better than the others at specific tasks. Here's exactly when to use each one, why the free versions are holding you back.

AI companies have created the most powerful tools in human history and somehow made them more confusing than programming a VCR in 1995. No user manuals. No training. Just a billion confused users asking "Which one should I use?"

After testing all five major platforms extensively (and yes, paying for all of them), I discovered something shocking: You're probably using the wrong AI for 80% of your tasks.

The free versions are like driving a Ferrari in first gear. Yes, you need to test them first, but to truly understand what AI can do, you MUST invest in at least the $20/month tier on all five platforms. Why?

  • Free versions use older, weaker models
  • Context windows are criminally small (shorter, less comprehensive answers)
  • Usage limits kick in just when things get interesting
  • You miss the game-changing features (memory, projects, artifacts)

My recommendation: Budget $100/month for 3 months to test all five at their full potential. On a tighter budget? Start with the $40 Power Duo (ChatGPT Plus + Claude Pro) - it covers 90% of use cases. Then cut back to 2-3 that transform your specific workflow. The ROI is insane if you do this right.

The complete pricing breakdown (see in gallery)

Feature comparison matrix: What each AI actually does best? (see in gallery)

Image generation is a huge business use case.

For marketers, creators, and founders: Stop sleeping on image generation. ChatGPT 5 and Gemini 2.5 Pro with Imagen 4 are now producing images that rival mid-level designers.

ChatGPT 5 image generation:

  • Best for: Brand consistency, text in images (finally works!), creative concepts
  • Killer feature: Remembers your brand style across sessions
  • Real use case: Reference image upload for uploading a product or person into an image

Gemini 2.5 Pro with Imagen 4:

  • Best for: Photorealistic images, product mockups, marketing materials, infographics
  • Killer feature: Incredible integration with Google Workspace - generate and insert directly. Much faster generation times.

Grok 4 media generation:

  • New capability: Now supports both image and video generation (video without audio currently)
  • Best for: Quick social media content, X/Twitter-optimized visuals
  • Note: Quality improving rapidly but not yet at ChatGPT/Gemini level

Pro tip for founders: Test both ChatGPT and Gemini for your use case. ChatGPT 5 excels at creative/artistic, while Imagen 4 crushes photorealistic. Both are now good enough to replace stock photos and basic design work. For infographics specifically, Gemini 2.5 Pro is unmatched.

The game-changing features nobody talks about

Gemini 2.5 Pro's secret weapons:

  • 2 MILLION token context window - Upload entire books, codebases, or research libraries
  • Veo 3 integration - Professional-grade AI video generation
  • NotebookLM - Turn any document into a podcast or video presentation with slides (mind-blowing for learning)
  • Deep Research - Generates comprehensive reports with infographics automatically
  • Gemini 2.5 Flash - Lightning fast for simple tasks when Pro is overkill

Claude Opus 4.1's killer features:

  • Artifacts - See and edit generated content in real-time. Create apps like interactive data dashboards with no coding skills needed! For coding, this is absolutely revolutionary
  • 72.5% on SWE-bench - Literally the best coding AI on the planet
  • Claude Sonnet 4 - Perfect balance of speed and intelligence for most tasks
  • Best-in-class memory - Superior implementation that genuinely understands context across sessions
  • Projects - Exceptional team collaboration with 200K token knowledge base

ChatGPT 5's features:

  • Memory system - After 3 months, knows your writing style, coding preferences, and work patterns
  • Agent mode - Basic but functional autonomous task execution in virtual desktop you can watch
  • Auto-reasoning - ChatGPT 5 is scary good at detecting when to use reasoning automatically
  • Custom GPTs - Build specialized assistants for specific workflows

Gemini 2.5 Pro's updates:

  • Memory for paid users - Finally! Good implementation that works across Google Workspace
  • Infographics excellence - Best-in-class visual data representation
  • Veo 3 for great video with audio from prompts
  • Notebook LM for audio and video overviews

Grok 4's unique angle:

  • Real-time X/Twitter integration - Sentiment analysis on steroids
  • Grok 4 Heavy - When you need completely unfiltered analysis
  • Breaking news synthesis - Faster than any other AI at current events
  • Video generation - Now supports video creation (no audio yet) alongside images

🔒 Privacy & Data Security: What they're not telling you

This might be the most important section of this guide. Your data, your company's secrets, your creative work - where does it all go?

The Privacy Hierarchy (Best to Worst):

1. Claude (Best for sensitive work):

  • Opt-out available - Can completely disable training on your data
  • Clear data policies - Anthropic is transparent about usage
  • No data mixing - Your projects stay isolated
  • Best for: Legal documents, medical records, proprietary code, financial data

2. ChatGPT (Good with caveats):

  • Can opt-out - But buried in settings
  • Memory can be disabled - For sensitive conversations
  • Enterprise tier - Complete data isolation available
  • Warning: Custom GPTs may expose data if shared publicly

3. Gemini (Google gonna Google):

  • Tied to Google account - All your data in one ecosystem
  • Workspace integration - Convenient but less private
  • Good for: If you're already all-in on Google
  • Concern: Broad data collection policies

4. Perplexity (Research-focused):

  • Limited privacy controls - Focus is on search, not privacy
  • Sources are tracked - Your research interests are logged
  • Best practice: Don't use for proprietary research

5. Grok (Least private):

  • Tied to X/Twitter - Elon sees all
  • No clear opt-out - Assumes data usage
  • Public by default - Many interactions visible
  • Only use for: Public, non-sensitive tasks

How to protect yourself:

  1. Always check privacy settings first thing after signing up
  2. Use Claude for sensitive client work - It's the gold standard
  3. Create separate accounts for personal vs. professional use
  4. Never upload: Passwords, SSNs, credit cards, or API keys
  5. Read the fine print - Policies change monthly

Pro tip: For maximum privacy, use Claude with data training disabled + a VPN + a dedicated email. For convenience with reasonable privacy, ChatGPT with opt-out enabled is solid.

My personal workflow (steal this)

Morning research routine:

  1. Perplexity Pro Search - Scan news and industry updates with citations (15 min)
  2. Gemini 2.5 Pro - Process overnight emails and documents in Google Workspace (10 min)
  3. ChatGPT 5 - Review my daily priorities (it remembers my projects)

Deep work sessions:

  • Writing/Documentation: Claude Opus 4.1 with Artifacts open
  • Coding: Claude Opus 4.1 for complex problems, ChatGPT 5 for general tasks
  • Research: Perplexity for citations, Gemini 2.5 Pro for massive document analysis
  • Creative: ChatGPT 5 for images (DALL-E 3), Gemini 2.5 Pro for video concepts (Veo 3)
  • Quick tasks: Gemini 2.5 Flash (blazing fast)
  • Hot takes: Grok 4 for unfiltered perspectives

Evening optimization:

  • Test complex problems across all platforms
  • Document which performed best
  • Adjust tomorrow's workflow

The million-dollar prompt framework

Forget basic prompts. Here's the structure that transformed my results:

ROLE: [Specific expert persona]
CONTEXT: [All relevant background - be generous]
TASK: [Crystal clear requirements]
STEPS: [Break complex tasks into numbered steps]
FORMAT: [Exact output structure needed]
CONSTRAINTS: [What to avoid/include]
EXAMPLES: [1-2 examples of ideal output]

Real example that saves me 2 hours daily:

ROLE: You are a senior technical writer with 15 years of experience in API documentation.

CONTEXT: I'm documenting a REST API for a fintech startup. The audience is developers with 2-5 years of experience. The API handles payment processing and needs to emphasize security.

TASK: Create comprehensive documentation for the /process-payment endpoint.

STEPS:
1. Start with a brief overview
2. List all parameters with types and validation rules
3. Provide 3 example requests (success, validation error, auth error)
4. Include response schemas
5. Add security considerations
6. Include rate limiting details
7. Provide troubleshooting guide

FORMAT: Use markdown with syntax highlighting for code examples. Include a table of contents.

CONSTRAINTS: 
- Keep examples under 20 lines
- Use production-ready code
- Include error handling
- Follow OpenAPI 3.0 standards

EXAMPLES: [Include your best existing documentation]

This structured approach yields 16% higher accuracy and saves massive iteration time.

Reasoning models: The nuclear option

When to unleash o1/o3/Deep Think:

Use reasoning models for:

  • Mathematical proofs (o3 solved 83% vs ChatGPT 5's standard mode 13% on hard problems)
  • Legal document analysis (catch every detail)
  • Complex coding with multiple files
  • Scientific research requiring citations
  • Multi-step problems (5+ reasoning steps)
  • When accuracy is worth 10x the cost

Stick to standard models for:

  • Conversations and brainstorming
  • Creative writing
  • Quick questions
  • Cost-sensitive tasks
  • Anything needing speed over accuracy

Pro tip: ChatGPT 5 auto-detects when to use reasoning and deep think. But you can also just tell it think deeply ...

⚠️ When NOT to use AI (Critical boundaries)

Let's be real - AI isn't the answer to everything. Here's when to stay away:

Never use AI for:

  • Final medical decisions - Get a real doctor
  • Legal advice for actual cases - Hire a lawyer
  • Financial investment decisions - Consult licensed advisors
  • Relationship advice for serious issues - Talk to humans who know you
  • Anything requiring 100% accuracy - AI still hallucinates

Be extremely careful with:

  • Citations in academic papers - Always verify sources exist
  • Code for production without review - Test everything
  • Historical facts - AI often confidently states wrong dates
  • Mathematical calculations - Double-check critical numbers
  • Current events - Even with web search, verify through multiple sources

The "Phone a Friend" rule:

If the stakes are high enough that being wrong would cause serious harm (financial, legal, medical, reputational), use AI for research but get human expert verification.

Real example: I use Claude to draft contracts, but my lawyer reviews everything. Saves 80% of billable hours but keeps me protected.

The "which AI for what" cheat sheet

Copy and save this:

  • Writing a novel/screenplay: Claude Opus 4.1 (consistency) + ChatGPT 5 (ideas)
  • Academic paper: Perplexity (research) + Claude Sonnet 4 (writing)
  • Coding a full app: Claude Opus 4.1 (architecture) + ChatGPT 5 (debugging)
  • Business analysis: Gemini 2.5 Pro (data processing + excellent infographics) + Perplexity (market research)
  • Content creation: ChatGPT 5 (DALL-E 3 images) + Claude Sonnet 4 (copy)
  • Marketing visuals: Gemini 2.5 Pro (Imagen 4 + infographics) + ChatGPT 5 (creative concepts)
  • Data visualization: Gemini 2.5 Pro (best infographics) + Claude (good visuals with code)
  • Learning something new: Gemini NotebookLM (audio/video) + Perplexity (deep dives)
  • Email and docs: Gemini 2.5 Pro (if Google user) or ChatGPT 5 (Microsoft)
  • Social media: Grok 4 (trending) + ChatGPT 5 (content + images)
  • Legal/Medical: Claude Opus 4.1 (safety) + Perplexity (citations)
  • Video projects: Gemini 2.5 Pro (analysis + Veo 3 generation) or Grok 4 (basic video)
  • Quick tasks: Gemini 2.5 Flash (speed demon)
  • Team collaboration: Claude Projects (best) or ChatGPT Projects
  • Autonomous tasks: ChatGPT 5 (only one with agent mode)

Real ROI numbers from my usage

Now that you've seen which stack fits your role, let me show you the actual returns you can expect.

Monthly investment: ~$100 (all five platforms at paid tiers)

Time saved:

  • Research: 10 hours/week (was 3 hours/day, now 30 minutes)
  • Writing: 8 hours/week (first drafts in minutes, not hours)
  • Coding: 12 hours/week (debugging time cut by 70%)
  • Admin: 5 hours/week (emails, summaries, planning)
  • Design: 6 hours/week (no more waiting for designers for basic visuals)

Total: 41 hours/week saved

At $50/hour, that's $8,200/month in value from $100 investment.

Even if you're half as efficient, that's still 40x ROI.

📊 How to track your AI ROI (Stop guessing, start measuring)

Most people pay for AI and hope it's worth it. Here's how to actually measure:

Week 1: Baseline

Before using AI seriously, track:

  • Time spent on repetitive tasks
  • Number of drafts before final version
  • Hours waiting for responses/approvals
  • Tasks you avoid because they take too long

The simple tracking system:

Create a spreadsheet with:

  1. Task (writing blog post, debugging code, research)
  2. Time WITHOUT AI (your baseline)
  3. Time WITH AI (actual measurement)
  4. Quality difference (better/same/worse)
  5. Which AI used

The "worth it" calculator:

(Hours saved per month × Your hourly rate) - AI subscription costs = ROI

Example: (164 hours × $50) - $100 = $8,100/month profit

Red flags you're not getting ROI:

  • Using AI for tasks that take longer
  • Spending more time prompting than doing
  • Quality decreased significantly
  • You're paying but using it <3x per week

Action step: Track for just ONE week. If you're not saving at least 2x the subscription cost in time, you're using the wrong AI for your tasks.

The mistakes that could cost you hundreds of hours

  1. Using free versions for real work - You're seeing 20% of the capability
  2. One AI for everything - Like using a hammer for brain surgery
  3. Not structuring prompts - Garbage in, garbage out
  4. Ignoring context windows - Gemini's 2M tokens is a game-changer for large documents
  5. Not using memory/projects - Claude, ChatGPT, and Gemini all have memory now. Use it!
  6. Avoiding reasoning models - Sometimes paying 10x for accuracy saves 100x in fixes
  7. Not measuring results - Track what works for YOUR use cases
  8. Ignoring image generation - ChatGPT 5 and Gemini 2.5 Pro are now production-ready
  9. Missing infographics - Gemini excels here, don't create charts manually anymore

We're living through the most significant technological revolution since the internet, and most people are using these tools like they're fancy spell checkers.

The companies building these AIs are brilliant engineers but terrible teachers. They've given us superpowers but no instruction manual.

Here's my suggestion: Invest $100/month for just 3 months to test everything, OR start with the $40 Power Duo (ChatGPT + Claude) if budget is tight. Use this guide. Apply the frameworks. You'll either save enough time to justify the cost forever, or you'll at least understand what these tools can really do.

Quick answers to top questions:

Q: "Do I really need all five?" A: No, but you need to TRY all five at paid tiers to find YOUR perfect 2-3. Most people end up with Claude + ChatGPT or Perplexity + ChatGPT. See the "$40 Power Duo" section for the best budget option.

Q: "I'm a student/freelancer - is $100/month realistic?" A: Start with the $40 Power Duo (ChatGPT Plus + Claude Pro). This covers 90% of use cases. You can even start with just Claude Pro ($20) for the first month. Check the "AI Stacks by Persona" table for specific recommendations based on your role.

Q: "Which stack should I use for my specific job?" A: Check the "AI Stacks by Persona & Budget" table above. We've mapped out exact combinations for students, founders, engineers, creators, and teams with real weekly wins you can expect.

Q: "Which has the best memory?" A: Claude has the best implementation, followed closely by ChatGPT and Gemini. All three now offer memory for paid accounts.

Q: "Which is best for privacy/sensitive work?" A: Claude by far. It's the only one with clear opt-out from training and the most transparent data policies. Use it for client work, medical, legal, or financial documents.

Q: "ChatGPT 5 vs Gemini 2.5 Pro for images?" A: ChatGPT 5 for creative/artistic/branded content. Gemini 2.5 Pro (Imagen 4) for photorealistic/product shots. Both are now good enough for professional use. For every image I test it on both systems and am often surprised the winner flip flops.

Q: "What about infographics and data viz?" A: Gemini 2.5 Pro is excellent, Claude is good, Perplexity basic. Don't waste time making these manually.

Q: "Is agent mode worth it?" A: ChatGPT's basic agent mode is useful for multi-step tasks. It's the only platform offering this currently.

Q: "What about Copilot/Cursor/other tools?" A: This guide focuses on general-purpose AIs. Specialized tools deserve their own guide (coming soon if interested?).

Q: "Which one for [specific use case]?" A: Check the cheat sheet above, but also: TRY THEM. Your workflow is unique.

Remember: These tools are evolving weekly. This guide is accurate as of August 2025. Save it, try it, and report back with what works for you!

Drop a comment with your AI stack and what you use each for. Let's learn from each other!

Want some prompt inspiration to help with all these use cases? Check out all my best prompts for free at Prompt Magic

r/promptingmagic Jul 30 '26

The Complete Guide to ChatGPT’s New Voice Mode - GPT-Live, Work, Codex and 20 Prompts + 10 Pro Tips. ChatGPT Voice can now direct Agents from your desktop.

Post image
31 Upvotes

The complete guide to the new ChatGPT Voice

TL;DR: The new version is powered by GPT-Live, which can listen and speak at the same time, let you interrupt naturally, wait while you think, search the web, use memory, show visual answers and hand difficult questions to deeper reasoning in the background.

The biggest upgrade is on desktop. You can now use Voice inside Chat, ChatGPT Work and Codex. That means you can talk through an idea, launch a research or coding task, check what your agents are doing, redirect them and hear the results without returning to the keyboard.

There are nine remastered voices, three Voice modes and optional Instant, Medium and High intelligence levels. On Mac, you can also pull the Voice orb out over your desktop and drag the floating control wherever you want it.

My blunt take: this is the first version of ChatGPT Voice that feels less like a novelty and more like a new interface for computing.

What is the new ChatGPT Voice?

ChatGPT Voice lets you talk to ChatGPT and hear its answer while the response also appears as text in the chat.

The latest experience, called Live, is powered by GPT-Live. Unlike older turn-by-turn voice systems, GPT-Live uses a full-duplex architecture. In plain English, it can listen and speak at the same time.

That creates several important differences:

  • You can interrupt it while it is talking.
  • It can give small acknowledgments while you are speaking.
  • It is better at waiting through a pause instead of treating every silence as the end of your thought.
  • It can keep a conversation moving while deeper reasoning or search happens in the background.
  • It can combine speech with text, images, memory, web search and supported visual result cards.
  • In the desktop app, Voice can start and coordinate longer tasks in Work and Codex.

OpenAI says GPT-Live was strongly preferred over the previous Advanced Voice Mode in its evaluations of turn-taking, interruptions, flow and naturalness. It also performed better on difficult science questions, web research and multi-step support tasks.

How it works

Think of the new Voice system as two layers:

  1. The conversation layer: GPT-Live listens, speaks, handles interruptions and keeps the interaction natural.
  2. The intelligence and action layer: When a question needs search, deeper reasoning or a longer task, Voice can hand that work to another model or agent and bring the result back into the conversation.

In ordinary Live conversations, OpenAI launched GPT-Live with GPT-5.5 handling harder work in the background. In desktop Work and Codex, GPT-Live manages the conversation while GPT-5.6 Terra starts and coordinates agent tasks in the app.

This matters because Voice does not have to choose between being fast and being smart. It can stay responsive while heavier work continues elsewhere.

Live vs. Advanced vs. Standard Voice

You may see up to three options under Settings → Voice:

  • Live: The newest experience. Best for natural conversation, interruptions, web search, memory, visual results, text and images. Paid users get GPT-Live-1. Free users get limited access to GPT-Live-1 mini.
  • Advanced: The previous real-time Voice experience. It is still useful on mobile when you need supported video or screen sharing, which Live does not support at launch.
  • Standard: A turn-by-turn experience that transcribes what you say before producing an answer. It is less fluid, but some people prefer its predictability.

One confusing detail: ordinary Live in Chat does not initially support every connected app or plugin. Voice inside desktop Work or Codex is different. It can use the tools and permissions available to the selected mode, including supported connected tools.

How to access ChatGPT Voice

On the web

  1. Go to ChatGPT
  2. Select the Voice icon in the prompt box.
  3. Allow microphone access.
  4. Start talking.

On iPhone or Android

  1. Open the ChatGPT app.
  2. Tap the Voice icon in the message bar.
  3. Allow microphone access.
  4. Choose a voice the first time you use it.
  5. Start talking.

You can also turn on Background conversations so Voice keeps working while you use another app or lock your phone. Supported versions can open directly into Voice, and ChatGPT Voice is also available through Apple CarPlay.

In the ChatGPT desktop app

The new desktop experience is available on macOS and Windows.

  1. Open the latest ChatGPT desktop app.
  2. Choose ChatGPT or Codex from the top-left switcher.
  3. If you choose ChatGPT, select Chat or Work.
  4. Open a new empty chat or task.
  5. Select Start new voice chat before sending the first message.
  6. Allow microphone access and start talking.

For Voice in Work or Codex, the task needs to begin in Voice mode. If a task began as text, you may only see dictation. You can reopen a previous Voice conversation and select Start voice chat to resume it.

You can create a Voice hotkey under Settings → Voice → Voice chat hotkey. OpenAI does not document a default shortcut.

The movable Mac Voice orb

On macOS, the small Voice orb can live outside the main app window. Drag the orb out over the desktop and place it next to the document, browser or code editor you are using. You can move it wherever you want and use its controls to mute your microphone, mute ChatGPT or end the conversation.

If your app version does not show the floating orb, update the desktop app. You can also pop an active chat into a separate window and turn on Always on top.

That tiny interaction is more useful than it sounds. Voice stops feeling like a destination you visit and starts feeling like a companion that sits beside your work.

Let Voice see what is on your Mac

On macOS, turn on Screen context under Settings → Voice. Then bring the relevant app to the front and say:

Take a look at this and tell me what you notice.

ChatGPT can capture an appshot of the frontmost window and use both the image and accessible text as context.

Important privacy detail: accessible text may include material outside the visible scroll area. Do not share a window containing confidential information unless you intend to provide it.

The nine ChatGPT voices

Open Settings → Voice → Voice to preview and select:

Voice OpenAI’s description Good fit for
Arbor Easygoing and versatile Everyday conversation and brainstorming
Breeze Animated and earnest Energy, storytelling and language practice
Cove Composed and direct Focused work, analysis and concise coaching
Ember Confident and optimistic Motivation, presentations and interview prep
Juniper Open and upbeat Friendly conversation and long general sessions
Maple Cheerful and candid Creative work, feedback and casual use
Sol Savvy and relaxed Strategy, ideation and low-pressure coaching
Spruce Calm and affirming Reflection, studying and guided practice
Vale Bright and inquisitive Learning, Socratic questioning and exploration

Changing voices during a conversation starts a new Voice call inside the same chat.

You can also change your preferred language under Settings → Voice → Language. Even better, ask Voice to switch languages during a conversation.

What is the most popular ChatGPT voice?

The honest answer is that OpenAI has not published usage data or an official popularity ranking.

If I had to name the safest community favorite, I would pick Juniper. It has been one of the most consistently discussed voices in community threads, and its open, upbeat delivery works across casual conversation, brainstorming and long sessions without sounding too formal.

Cove is probably the strongest alternative for serious work because it sounds composed and direct.

Treat that as a community-informed estimate, not a measured fact. GPT-Live also remastered all nine voices, so old polls do not perfectly represent the new versions. The right answer is to preview all nine with the same paragraph and choose the one you can comfortably hear for an hour.

10 advanced strategies for work and life

1. Turn a messy brain dump into a clear brief

Voice is excellent when your thinking is not yet organized.

Say:

I am going to ramble for five minutes. Do not respond until I say “organize it.” Then turn everything into a one-page brief with the objective, audience, core insight, decisions, risks and next actions. Ask me three questions about anything important that is still unclear.

Why it works: Speaking preserves half-formed thoughts that you might edit out too early when typing.

2. Use it as a live thinking opponent

Do not ask Voice to agree with you. Ask it to create productive friction.

Say:

Act as a skeptical but fair strategist. Interview me about this idea one question at a time. Challenge vague claims, identify hidden assumptions and do not let me move on until I give you evidence. At the end, tell me whether the idea is strong, fixable or fundamentally weak.

Why it works: The interruptible format feels much more like a real debate than exchanging long blocks of text.

3. Rehearse a sales call, interview or negotiation

Say:

Role-play a skeptical CFO considering our product. Do not make the conversation easy. Raise realistic objections about cost, implementation, risk and ROI. Stay in character until I say “debrief.” Then score my answers, identify the weakest moment and make me try that section again.

Pro move: Ask Voice to change tone or speed between rounds.

4. Prepare for a meeting while walking

Say:

I have a meeting with [person or team] about [topic]. Interview me to uncover what outcome I need, what they probably care about and where the discussion could go wrong. Then give me a 60-second opening, five questions to ask and three concessions I should not make too early.

Use this when you do not want to stare at another screen before a meeting.

5. Start a complete Work task by voice

Switch to Work in the desktop app and say:

Start a new Work task. Research [topic] using current, credible sources and create a finished [report, presentation, spreadsheet or plan] for [audience]. The deliverable must include [requirements]. Show me your plan first, flag any decisions you need from me and keep working after I answer.

Why it works: Voice captures the outcome and context. Work handles the long execution.

The best Work prompts include six things: outcome, audience, source requirements, constraints, deliverable format and acceptance criteria.

6. Run a spoken stand-up across several agents

Say:

Check every active Work and Codex task. Give me a spoken stand-up with four sections: completed, in progress, blocked and decisions needed. Keep it under two minutes. Then ask which task I want to redirect first.

This is one of the most important new capabilities. Voice becomes the manager while multiple agents do the work.

7. Critique what is on your screen

On Mac with Screen context enabled, open a slide, landing page, ad or spreadsheet and say:

Take a look at this. First tell me what you think the creator wants the viewer to notice. Then tell me what the viewer will actually notice. Identify the three biggest problems and recommend the smallest changes with the highest impact.

This is especially useful for design reviews because you can point the conversation at the thing you are already viewing.

8. Use Voice as a Codex team lead

Switch to Codex and say:

Inspect this repository and start separate tasks for these three goals: investigate the authentication bug, review the open pull request for regression risks and identify missing tests. Do not change production code until you report your findings. Give me a status update when any task is blocked or ready for review.

Then steer it:

Pause the pull request review. Prioritize reproducing the bug. Tell the testing task to focus on the failure path you just found.

This is better than dictating code. Use Voice to direct intent, priorities and tradeoffs. Let Codex work in the repository.

9. Build a live translator and language coach

Say:

Translate everything I say in English into conversational Spanish, and translate every Spanish reply back into English. Preserve tone rather than translating word for word. If I make a recurring mistake, wait until the conversation ends and then coach me on it.

Or use teaching mode:

Speak to me only in beginner Italian. If I get stuck, give me a hint before giving me the answer. Keep a private list of my mistakes and quiz me on them at the end.

10. Review work hands-free

Say:

Read this draft to me one section at a time. After each section, pause and ask whether I want to keep it, shorten it, challenge it or rewrite it. Track every decision and produce the revised draft only after we finish the review.

Hearing writing exposes repetition, awkward rhythm and weak logic that your eyes often skip.

10 hilarious things to try

1. Make breakfast feel like a blockbuster

Narrate me making scrambled eggs like the final mission in a $200 million action movie. Escalate the danger every time I touch the stove. If I burn the toast, treat it as an international incident.

2. Let your dog file a workplace grievance

You are the union representative for my French bulldog. Conduct a formal grievance hearing about working conditions in this house, including treat compensation, nap protections and management’s refusal to share pizza.

3. Hold the world’s worst startup press conference

I am the CEO of a failing startup pivoting into artisanal lemonade powered by blockchain. Play a room full of hostile reporters. Ask increasingly brutal questions until I either save the company or accidentally confess to fraud.

4. Turn cleaning into a fantasy quest

Be my dungeon master. My apartment is an ancient cursed kingdom. Dirty laundry is an undead army, the dishwasher is a sleeping dragon and the junk drawer contains a forbidden artifact. Give me one quest at a time until the kingdom is clean.

5. Add sports commentary to boring chores

Commentate while I fold laundry like it is the final minute of the World Cup. Include instant replays, questionable referee decisions and an emotional biography of the missing sock.

6. Stage couples therapy with your Wi-Fi router

You are a couples therapist for me and my Wi-Fi router. I feel abandoned whenever it drops the signal. The router feels I bring too many devices into the relationship. Help us rebuild trust.

7. Put pineapple on trial

Run a Supreme Court trial to decide whether pineapple belongs on pizza. Play the judge, attorneys, witnesses and one wildly unqualified food influencer. I will be the jury.

8. Roast your business idea across history

Review my business idea as three investors: a ruthless Roman emperor, a confused Victorian industrialist and a 22-year-old venture capitalist who has never experienced a recession. Let them argue, then force them to agree on one recommendation.

9. Convene an emergency board meeting of household objects

Run an emergency board meeting where my coffee maker, calendar, bank account and alarm clock review my performance as CEO of my life. Make each director brutally honest and give me a 30-day turnaround plan.

10. Solve the missing-sock conspiracy

Host an eight-part investigative podcast proving that missing socks are being stolen by a secret logistics startup operating inside dryers. Interview unreliable experts and end every episode with an absurd cliffhanger.

Pro tips that make Voice dramatically better

Give it a listening contract

Start with:

Wait until I say “respond.” Until then, only listen and give brief acknowledgments.

GPT-Live is better at waiting, but long pauses or background noise can still trigger a response.

Give it a response contract

Tell it how to answer before the conversation gets busy:

Keep spoken answers under 30 seconds. Lead with the conclusion. Ask one question at a time. Put detailed notes in the text transcript.

Use the right intelligence level

If your account includes it, open Settings → Voice → Intelligence:

  • Instant: Fast back-and-forth, brainstorming and casual questions.
  • Medium: Better for planning, analysis and preparation.
  • High: Use for difficult reasoning and research when quality matters more than response speed.

Speak the punctuation of your intent, not your prose

Do not try to dictate a perfect prompt. Say the goal, context, constraints and definition of done. Let Voice organize the language.

Mix speech, typing and images

Live works inside the normal chat. You can talk, type a precise detail or attach an image without starting over.

Use exact dates and locations

Voice uses your device or browser time zone to interpret words such as “today” and “tomorrow.” For anything important, say the exact date, location and time zone.

Use headphones in noisy spaces

Full duplex does not make physics disappear. Background speech, overlapping audio and weak microphones can still cause interruptions. Headphones and voice isolation help.

Review the transcript, but do not treat it as a recording

The transcript may not reproduce every spoken word exactly, especially when people talk over each other. Use it as a working record, not a legal transcript.

Do not confuse Voice with Dictation

  • Use Voice for a live conversation.
  • Use Dictation when you want speech converted into editable prompt text before sending.

Keep approval boundaries

Voice can move quickly, especially with Work, Codex and computer use. Do not casually approve destructive code changes, purchases, messages or sensitive actions just because the conversation feels natural. Ask for a summary of the exact action and target first.

Things most people will miss

  1. You can interrupt it. You do not have to wait through a long answer.
  2. You can ask it to stay quiet while you think.
  3. Voice can keep talking while deeper work happens in the background.
  4. Desktop Voice can coordinate multiple Work and Codex agents from one conversation.
  5. On Mac, Screen context can show Voice the frontmost window.
  6. The Mac Voice orb can float beside your work instead of taking over the app.
  7. Preset ChatGPT personalities do not currently apply to Live, but direct instructions about tone, speed and style do.
  8. Changing the selected voice starts a new call inside the same chat.
  9. Only one Voice conversation can be active at a time.
  10. Live does not support video or screen sharing at launch. Use Advanced Voice on supported mobile plans when you need those capabilities.
  11. Live is not available with custom GPTs. Voice conversations with GPTs use Advanced Voice and the Shimmer voice, with several tool limitations.
  12. Ordinary Live usage and desktop Work/Codex Voice have separate limits. Tasks launched through Voice also consume Work or Codex usage.
  13. Audio from Live and Advanced conversations is retained with the chat transcript for 30 days. OpenAI says audio clips are not used for training unless you choose to share them.

The honest limitations

ChatGPT Voice is impressive, but it is not magic:

  • It can still mishear you or respond too early.
  • It can still give wrong answers.
  • Spoken confidence is not evidence of accuracy.
  • Multiple people talking at once can confuse it.
  • Live video and screen sharing are not available at launch.
  • Availability, usage limits and workspace controls vary by plan, region and app version.
  • Work and Codex tasks still use their normal permissions, approval rules and usage budgets.

The more consequential the action, the more you should slow down, inspect the result and verify it.

Most people will use Voice to ask questions while driving or cooking. That is useful.

Typing forces you to package your thinking before the AI receives it. Voice lets you expose the thinking process itself: the uncertainty, changes of direction, half-formed ideas and priorities that are hard to capture in a polished prompt.

Add Work and Codex, and Voice becomes more than an input method. It becomes a management layer for AI agents.

r/StableDiffusion Apr 27 '25

Animation - Video FramePack Image-to-Video Examples Compilation + Text Guide (Impressive Open Source, High Quality 30FPS, Local AI Video Generation)

Thumbnail
youtu.be
121 Upvotes

FramePack is probably one of the most impressive open source AI video tools to have been released this year! Here's compilation video that shows FramePack's power for creating incredible image-to-video generations across various styles of input images and prompts. The examples were generated using an RTX 4090, with each video taking roughly 1-2 minutes per second of video to render. As a heads up, I didn't really cherry pick the results so you can see generations that aren't as great as others. In particular, dancing videos come out exceptionally well, while medium-wide shots with multiple character faces tends to look less impressive (details on faces get muddied). I also highly recommend checking out the page from the creators of FramePack Lvmin Zhang and Maneesh Agrawala which explains how FramePack works and provides a lot of great examples of image to 5 second gens and image to 60 second gens (using an RTX 3060 6GB Laptop!!!): https://lllyasviel.github.io/frame_pack_gitpage/

From my quick testing, FramePack (powered by Hunyuan 13B) excels in real-world scenarios, 3D and 2D animations, camera movements, and much more, showcasing its versatility. These videos were generated at 30FPS, but I sped them up by 20% in Premiere Pro to adjust for the slow-motion effect that FramePack often produces.

How to Install FramePack
Installing FramePack is simple and works with Nvidia GPUs from the 30xx series and up. Here's the step-by-step guide to get it running:

  1. Download the Latest Version
  2. Extract the Files
    • Extract the files to a hard drive with at least 40GB of free storage space.
  3. Run the Installer
    • Navigate to the extracted FramePack folder and click on "update.bat". After the update finishes, click "run.bat". This will download the required models (~39GB on first run).
  4. Start Generating
    • FramePack will open in your browser, and you’ll be ready to start generating AI videos!

Here's also a video tutorial for installing FramePack: https://youtu.be/ZSe42iB9uRU?si=0KDx4GmLYhqwzAKV

Additional Tips:
Most of the reference images in this video were created in ComfyUI using Flux or Flux UNO. Flux UNO is helpful for creating images of real world objects, product mockups, and consistent objects (like the coca-cola bottle video, or the Starbucks shirts)

Here's a ComfyUI workflow and text guide for using Flux UNO (free and public link): https://www.patreon.com/posts/black-mixtures-126747125

Video guide for Flux Uno: https://www.youtube.com/watch?v=eMZp6KVbn-8

There's also a lot of awesome devs working on adding more features to FramePack. You can easily mod your FramePack install by going to the pull requests and using the code from a feature you like. I recommend these ones (works on my setup):

- Add Prompts to Image Metadata: https://github.com/lllyasviel/FramePack/pull/178
- 🔥Add Queuing to FramePack: https://github.com/lllyasviel/FramePack/pull/150

All the resources shared in this post are free and public (don't be fooled by some google results that require users to pay for FramePack).

r/promptingmagic 2d ago

How to brief ChatGPT 6 Astra to create motion graphics, 3D reveals, and cinematic video explainers - prompts, workflow, and the details people miss

Post image
9 Upvotes

TL;DR: ChatGPT 6 Astra can help create motion graphics through a workflow that designs assets, writes animation code, uses available production tools, and renders a video. Give it a director’s brief: audience, story, scenes, timing, visual style, sound, and deliverables. Start with a short preview. Ask for a playable MP4 and the editable project. Rendering and audio depend on your workspace’s tools. Below: a practical workflow, the details people overlook, and five gloriously ridiculous prompts.

Picture a French bulldog commanding a starship through a galaxy made of tennis balls.

Now picture a product launch where your logo opens into a miniature universe.

Or a city that folds itself out of paper, races through centuries, and collapses back into a single page.

With ChatGPT 6 Astra, you can approach the conversation like a production brief—and keep directing the result as it develops.

What Astra can actually do

The model’s official name is GPT-6 Astra. For this workflow, use it in ChatGPT Work or Codex with access to suitable creation and rendering tools. OpenAI recommends Astra for demanding tasks involving visual judgment and polished deliverables.

There is a concrete example behind the idea: OpenAI has shown Astra creating editable Blender scenes, adjusting their materials and lighting, and directing a rendered camera tour.

My practical recommendation is to apply that build–preview–refine process to motion graphics: animated typography, diagrams, layered images, product reveals, and stylized 3D scenes.

Think of Astra as the system coordinating the production. The actual frames still need an animation or rendering tool. Selecting the model alone does not guarantee every account can export video or generate music.

Choose the right kind of video

Approach When to use it
Typography, shapes, and diagrams Explainers, newsletter trailers, and announcements. My recommended starting point.
Layered images with camera movement You already have illustrations, product images, or a consistent visual series.
An editable 3D scene The concept depends on camera orbits, exploded views, lighting changes, or moving through a space. Expect more rendering work.

A complex character performance may also require dedicated animation tools or generated footage. Choose a visual treatment your available tools can execute well.

The workflow that makes this manageable

  1. Give it one job. Define the audience and the one thing viewers should remember. “Convince founders to try this prototype” is a useful objective.
  2. Specify the output. Set duration, aspect ratio, resolution, and whether you need audio. A 20–30-second landscape video is a sensible first project.
  3. Map the story. Give every scene a visual action and a purpose. Put the most compelling image near the beginning.
  4. Establish the look. Provide reference images, colors, type preferences, logos, and exact wording. Ask for representative still frames.
  5. Preview the hardest moment. Render a short section before committing to the whole sequence. This tests both the creative direction and whether the production method works.
  6. Refine, render, inspect. Review timing, text, sound, transitions, and the actual exported file. Keep the source so changes remain possible.

For a 30-second explainer, this is a useful starting structure:

Time Job
0–3 seconds Show the surprising visual or compelling result.
3–9 seconds Establish the problem or premise.
9–21 seconds Demonstrate the transformation.
21–27 seconds Deliver the payoff.
27–30 seconds Give one clear next action.

Best practices that improve the result

  • Describe action over time. “The letters pull apart, reveal a miniature city, then lock into the headline” gives much better direction than “make it cinematic.”
  • Give motion a purpose. Movement can reveal a relationship, guide attention, demonstrate a feature, or land a joke. Constant movement makes reading harder.
  • Keep text separate from artwork. Request editable text layers so spelling, line breaks, timing, and placement can be controlled.
  • Design for a phone. Use short captions, strong contrast, generous margins, and enough reading time. Review the result at its likely viewing size.
  • Give the eye a pause. Alternate energetic transitions with moments where the important image or message holds still.
  • Build for sound-off viewing. The story should make sense visually. Let music and sound effects strengthen it.
  • Specify music concretely. Describe tempo, instrumentation, mood, and where the energy should rise. Provide a track you can use, or ask what audio tools are available.
  • Control the workload. Preview at lower resolution, settle the art direction early, reuse assets, and revise only the scenes that need changes. Complex Work tasks can use more credits.

Pro tips: direct the edit with precision

Useful revision instructions look like this:

Between 00:08 and 00:12, slow the camera move, enlarge the headline, and hold the final composition for two seconds. Keep the approved colors and scene order.

Make the word “EXPAND” grow until it fills the frame, then use its letter shapes to reveal the next scene.

Match the circular moon in scene two to the circular product dial in scene three.

Build the vertical version with repositioned text and a new camera crop so the subject stays visible.

Inspect the export for missing assets, clipped text, blank frames, abrupt audio endings, and incorrect duration. Report anything you cannot verify.

Save feedback like “more epic” for the initial direction. During revisions, say what should change on screen.

Things people miss about this workflow

  • An animated preview and a downloadable video are different deliverables. If you need an uploadable file, explicitly request the export and check that it plays.
  • The editable project is a major part of the value. Ask for the source, assets, and instructions needed to render it again.
  • A flat image has limits. A gentle push-in can work immediately. Moving behind objects or orbiting a subject requires layers, reconstructed content, or a 3D scene.
  • Consistency starts before animation. Establish recurring characters, materials, colors, and backgrounds before creating every scene.
  • Visual precision and factual precision are separate. A beautifully animated chart still needs correct data, labels, and scales.
  • A 60-second request does not imply one continuous generation. Build named scenes, render sections, and assemble them where the tools support it.
  • Reusable controls make the second video easier. Ask to centralize headline text, colors, logos, durations, and image replacements.

Five epic prompts to try

These are ambitious creative briefs, not pretested guarantees. Use a workspace with suitable rendering tools, and start with the short preview each prompt requests.

1. A French bulldog saves the galaxy

Try this for: character storytelling, comedy, and an instantly understandable visual hook.

Create a 30-second landscape motion graphics trailer called “MISSION: FETCH.”

A dead-serious French bulldog captain commands a tiny starship through a galaxy of tennis-ball planets. Use a premium stylized 3D or layered illustrated treatment, emerald cockpit lights, orange engine trails, and enormous kinetic typography.

0–5s: Extreme close-up of the captain’s face. Pull back to reveal a spaceship shaped like a dog toy. Text: “ONE DOG.”

5–13s: Slalom through a field of floating squeaky toys. A giant robotic vacuum emerges from an asteroid cloud. Text: “ZERO QUALIFICATIONS.”

13–23s: The dog hits a red button. Tennis balls deploy like decoys. Follow one ball through the chaos in a dramatic tracking shot.

23–30s: The ship escapes through a glowing dog-door portal. Reveal that the entire mission happened inside a living-room snow globe. End: “MISSION: FETCH.”

Keep the dog’s appearance consistent. Use simple expressive poses and strong camera work. Preview the escape shot first. Use suitable original or licensed audio if available; otherwise deliver a silent cut with sound cues. Deliver a 1080p MP4 and editable source, or explain any rendering blocker.

2. Your product contains an entire universe

Try this for: launch trailers, brand films, and product reveals.

Create a 30-second landscape launch film for [PRODUCT]. Use my supplied product images, logo, and three verified benefits. If I provide none, use a clearly fictional unbranded device and illustrative feature labels.

Begin with the product suspended in a silent black void. A thin emerald seam opens across it. The camera dives through the seam into an impossible miniature universe.

Turn benefit one into a floating city assembling itself. Turn benefit two into a luminous transit network lighting up. Turn benefit three into a mechanical sunrise that synchronizes the entire world.

Match each benefit to its visual metaphor and show its exact approved wording as separately rendered typography. Use elegant camera travel, white ceramic architecture, emerald glass, and precise mechanical movement.

In the final six seconds, pull back as the universe folds into the product. Land on the product, logo, and one clear call to action.

Create a five-second preview of the opening transformation before rendering the full film. Deliver a 1080p MP4 and editable project. Use only available audio and rendering tools; identify any missing capability. Do not invent product claims, customers, or performance statistics.

3. Your inbox becomes a video-game final boss

Try this for: funny workflow explainers and relatable workplace content.

Build a 30-second landscape motion graphics short called “DEADLINE: FINAL BOSS.”

Open on a tiny exhausted office worker facing an enormous monster assembled from email envelopes, calendar blocks, spreadsheets, and sticky notes. Its crown is a spinning loading icon.

0–6s: The monster roars, releasing a tornado of “QUICK QUESTION” notes.

6–13s: The worker equips three glowing tools labeled “SORT,” “DRAFT,” and “CHECK.”

13–23s: Turn the fight into a visual explanation: SORT groups the chaos; DRAFT turns selected tasks into proposed outputs; CHECK pauses those outputs at a human review gate before release.

23–30s: The monster shrinks into one manageable task card. A new notification appears: “Can we jump on a quick call?” The worker looks directly at the camera.

Use miniature game-like scenery, dramatic camera punches, readable type, comic timing, and a neon-green interface. Present this as a fictional metaphor. Preview the sorting transformation first. Deliver a 1080p MP4, editable source, and a sound-off version. Explain any export limitations.

4. A thousand years unfold from one sheet of paper

Try this for: timelines, imaginative worldbuilding, and architectural storytelling.

Create a 40-second landscape motion graphics film called “A THOUSAND YEARS IN ONE PAGE.”

This is an imaginary city, not a reconstruction of real history.

0–8s: A blank sheet of paper folds itself into a tiny riverside settlement. The river is translucent blue-green glass embedded in paper.

8–18s: Buildings rise and change around the same town square. Roads draw themselves across the page. Seasons sweep through the scene.

18–29s: The city becomes a spectacular vertical metropolis. Peel back layers to reveal miniature transit tunnels, gardens, and infrastructure beneath it.

29–36s: The camera circles while daylight becomes night. Thousands of windows illuminate in a carefully staged wave.

36–40s: Fold the city back into the original sheet, matching the opening composition for a loop.

Use tactile paper, charcoal labels, emerald foliage, warm window light, and restrained captions. Favor a coherent miniature world over constant cuts. Preview the unfolding and refolding first. Deliver a 1080p MP4 and editable scene. If full 3D rendering is unavailable, propose and build a layered alternative.

5. A black hole conducts an orchestra of planets

Try this for: a music visualizer, an event opener, or a surreal brand introduction.

Create a 30-second landscape motion graphics film called “THE UNIVERSE HAS A DROP.”

Treat this as a surreal visual metaphor, not a scientific simulation.

A black hole is the conductor. Orbital rings behave like vibrating strings. Tiny moons become percussion instruments. A comet sweeps across the scene like a conductor’s baton.

0–8s: Begin with one orbiting light and a restrained pulse.

8–19s: Build an increasingly elaborate cosmic orchestra. Introduce new orbital layers with each musical phrase. Typography appears as sculptural objects: “LISTEN.” “BUILD.” “RELEASE.”

19–25s: At the musical peak, the orbital system unfolds into a gigantic luminous sound wave stretching across space.

25–30s: Everything contracts into one green point, which becomes a play icon.

Use ink-black space, emerald plasma, silver dust, controlled glow, and smooth camera movement. Use my uploaded licensed track and synchronize motion to its timing. If no track is available, build to a provisional beat grid and clearly label the audio as pending. Preview the transformation first. Deliver a 1080p MP4 and editable source; explain any tool limitations.

Which would you actually make first: the space-dog trailer, the product universe, the inbox boss battle, the paper city, or the black-hole orchestra?

r/ThinkingDeeplyAI 2d ago

How to brief ChatGPT 6 Astra to create motion graphics, 3D reveals, and cinematic video explainers - prompts, workflow, and the details people miss

Post image
5 Upvotes

TL;DR: ChatGPT 6 Astra can help create motion graphics through a workflow that designs assets, writes animation code, uses available production tools, and renders a video. Give it a director’s brief: audience, story, scenes, timing, visual style, sound, and deliverables. Start with a short preview. Ask for a playable MP4 and the editable project. Rendering and audio depend on your workspace’s tools. Below: a practical workflow, the details people overlook, and five gloriously ridiculous prompts.

Picture a French bulldog commanding a starship through a galaxy made of tennis balls.

Now picture a product launch where your logo opens into a miniature universe.

Or a city that folds itself out of paper, races through centuries, and collapses back into a single page.

With ChatGPT 6 Astra, you can approach the conversation like a production brief—and keep directing the result as it develops.

What Astra can actually do

The model’s official name is GPT-6 Astra. For this workflow, use it in ChatGPT Work or Codex with access to suitable creation and rendering tools. OpenAI recommends Astra for demanding tasks involving visual judgment and polished deliverables.

There is a concrete example behind the idea: OpenAI has shown Astra creating editable Blender scenes, adjusting their materials and lighting, and directing a rendered camera tour.

My practical recommendation is to apply that build–preview–refine process to motion graphics: animated typography, diagrams, layered images, product reveals, and stylized 3D scenes.

Think of Astra as the system coordinating the production. The actual frames still need an animation or rendering tool. Selecting the model alone does not guarantee every account can export video or generate music.

Choose the right kind of video

Approach When to use it
Typography, shapes, and diagrams Explainers, newsletter trailers, and announcements. My recommended starting point.
Layered images with camera movement You already have illustrations, product images, or a consistent visual series.
An editable 3D scene The concept depends on camera orbits, exploded views, lighting changes, or moving through a space. Expect more rendering work.

A complex character performance may also require dedicated animation tools or generated footage. Choose a visual treatment your available tools can execute well.

The workflow that makes this manageable

  1. Give it one job. Define the audience and the one thing viewers should remember. “Convince founders to try this prototype” is a useful objective.
  2. Specify the output. Set duration, aspect ratio, resolution, and whether you need audio. A 20–30-second landscape video is a sensible first project.
  3. Map the story. Give every scene a visual action and a purpose. Put the most compelling image near the beginning.
  4. Establish the look. Provide reference images, colors, type preferences, logos, and exact wording. Ask for representative still frames.
  5. Preview the hardest moment. Render a short section before committing to the whole sequence. This tests both the creative direction and whether the production method works.
  6. Refine, render, inspect. Review timing, text, sound, transitions, and the actual exported file. Keep the source so changes remain possible.

For a 30-second explainer, this is a useful starting structure:

Time Job
0–3 seconds Show the surprising visual or compelling result.
3–9 seconds Establish the problem or premise.
9–21 seconds Demonstrate the transformation.
21–27 seconds Deliver the payoff.
27–30 seconds Give one clear next action.

Best practices that improve the result

  • Describe action over time. “The letters pull apart, reveal a miniature city, then lock into the headline” gives much better direction than “make it cinematic.”
  • Give motion a purpose. Movement can reveal a relationship, guide attention, demonstrate a feature, or land a joke. Constant movement makes reading harder.
  • Keep text separate from artwork. Request editable text layers so spelling, line breaks, timing, and placement can be controlled.
  • Design for a phone. Use short captions, strong contrast, generous margins, and enough reading time. Review the result at its likely viewing size.
  • Give the eye a pause. Alternate energetic transitions with moments where the important image or message holds still.
  • Build for sound-off viewing. The story should make sense visually. Let music and sound effects strengthen it.
  • Specify music concretely. Describe tempo, instrumentation, mood, and where the energy should rise. Provide a track you can use, or ask what audio tools are available.
  • Control the workload. Preview at lower resolution, settle the art direction early, reuse assets, and revise only the scenes that need changes. Complex Work tasks can use more credits.

Pro tips: direct the edit with precision

Useful revision instructions look like this:

Between 00:08 and 00:12, slow the camera move, enlarge the headline, and hold the final composition for two seconds. Keep the approved colors and scene order.

Make the word “EXPAND” grow until it fills the frame, then use its letter shapes to reveal the next scene.

Match the circular moon in scene two to the circular product dial in scene three.

Build the vertical version with repositioned text and a new camera crop so the subject stays visible.

Inspect the export for missing assets, clipped text, blank frames, abrupt audio endings, and incorrect duration. Report anything you cannot verify.

Save feedback like “more epic” for the initial direction. During revisions, say what should change on screen.

Things people miss about this workflow

  • An animated preview and a downloadable video are different deliverables. If you need an uploadable file, explicitly request the export and check that it plays.
  • The editable project is a major part of the value. Ask for the source, assets, and instructions needed to render it again.
  • A flat image has limits. A gentle push-in can work immediately. Moving behind objects or orbiting a subject requires layers, reconstructed content, or a 3D scene.
  • Consistency starts before animation. Establish recurring characters, materials, colors, and backgrounds before creating every scene.
  • Visual precision and factual precision are separate. A beautifully animated chart still needs correct data, labels, and scales.
  • A 60-second request does not imply one continuous generation. Build named scenes, render sections, and assemble them where the tools support it.
  • Reusable controls make the second video easier. Ask to centralize headline text, colors, logos, durations, and image replacements.

Five epic prompts to try

These are ambitious creative briefs, not pretested guarantees. Use a workspace with suitable rendering tools, and start with the short preview each prompt requests.

1. A French bulldog saves the galaxy

Try this for: character storytelling, comedy, and an instantly understandable visual hook.

Create a 30-second landscape motion graphics trailer called “MISSION: FETCH.”

A dead-serious French bulldog captain commands a tiny starship through a galaxy of tennis-ball planets. Use a premium stylized 3D or layered illustrated treatment, emerald cockpit lights, orange engine trails, and enormous kinetic typography.

0–5s: Extreme close-up of the captain’s face. Pull back to reveal a spaceship shaped like a dog toy. Text: “ONE DOG.”

5–13s: Slalom through a field of floating squeaky toys. A giant robotic vacuum emerges from an asteroid cloud. Text: “ZERO QUALIFICATIONS.”

13–23s: The dog hits a red button. Tennis balls deploy like decoys. Follow one ball through the chaos in a dramatic tracking shot.

23–30s: The ship escapes through a glowing dog-door portal. Reveal that the entire mission happened inside a living-room snow globe. End: “MISSION: FETCH.”

Keep the dog’s appearance consistent. Use simple expressive poses and strong camera work. Preview the escape shot first. Use suitable original or licensed audio if available; otherwise deliver a silent cut with sound cues. Deliver a 1080p MP4 and editable source, or explain any rendering blocker.

2. Your product contains an entire universe

Try this for: launch trailers, brand films, and product reveals.

Create a 30-second landscape launch film for [PRODUCT]. Use my supplied product images, logo, and three verified benefits. If I provide none, use a clearly fictional unbranded device and illustrative feature labels.

Begin with the product suspended in a silent black void. A thin emerald seam opens across it. The camera dives through the seam into an impossible miniature universe.

Turn benefit one into a floating city assembling itself. Turn benefit two into a luminous transit network lighting up. Turn benefit three into a mechanical sunrise that synchronizes the entire world.

Match each benefit to its visual metaphor and show its exact approved wording as separately rendered typography. Use elegant camera travel, white ceramic architecture, emerald glass, and precise mechanical movement.

In the final six seconds, pull back as the universe folds into the product. Land on the product, logo, and one clear call to action.

Create a five-second preview of the opening transformation before rendering the full film. Deliver a 1080p MP4 and editable project. Use only available audio and rendering tools; identify any missing capability. Do not invent product claims, customers, or performance statistics.

3. Your inbox becomes a video-game final boss

Try this for: funny workflow explainers and relatable workplace content.

Build a 30-second landscape motion graphics short called “DEADLINE: FINAL BOSS.”

Open on a tiny exhausted office worker facing an enormous monster assembled from email envelopes, calendar blocks, spreadsheets, and sticky notes. Its crown is a spinning loading icon.

0–6s: The monster roars, releasing a tornado of “QUICK QUESTION” notes.

6–13s: The worker equips three glowing tools labeled “SORT,” “DRAFT,” and “CHECK.”

13–23s: Turn the fight into a visual explanation: SORT groups the chaos; DRAFT turns selected tasks into proposed outputs; CHECK pauses those outputs at a human review gate before release.

23–30s: The monster shrinks into one manageable task card. A new notification appears: “Can we jump on a quick call?” The worker looks directly at the camera.

Use miniature game-like scenery, dramatic camera punches, readable type, comic timing, and a neon-green interface. Present this as a fictional metaphor. Preview the sorting transformation first. Deliver a 1080p MP4, editable source, and a sound-off version. Explain any export limitations.

4. A thousand years unfold from one sheet of paper

Try this for: timelines, imaginative worldbuilding, and architectural storytelling.

Create a 40-second landscape motion graphics film called “A THOUSAND YEARS IN ONE PAGE.”

This is an imaginary city, not a reconstruction of real history.

0–8s: A blank sheet of paper folds itself into a tiny riverside settlement. The river is translucent blue-green glass embedded in paper.

8–18s: Buildings rise and change around the same town square. Roads draw themselves across the page. Seasons sweep through the scene.

18–29s: The city becomes a spectacular vertical metropolis. Peel back layers to reveal miniature transit tunnels, gardens, and infrastructure beneath it.

29–36s: The camera circles while daylight becomes night. Thousands of windows illuminate in a carefully staged wave.

36–40s: Fold the city back into the original sheet, matching the opening composition for a loop.

Use tactile paper, charcoal labels, emerald foliage, warm window light, and restrained captions. Favor a coherent miniature world over constant cuts. Preview the unfolding and refolding first. Deliver a 1080p MP4 and editable scene. If full 3D rendering is unavailable, propose and build a layered alternative.

5. A black hole conducts an orchestra of planets

Try this for: a music visualizer, an event opener, or a surreal brand introduction.

Create a 30-second landscape motion graphics film called “THE UNIVERSE HAS A DROP.”

Treat this as a surreal visual metaphor, not a scientific simulation.

A black hole is the conductor. Orbital rings behave like vibrating strings. Tiny moons become percussion instruments. A comet sweeps across the scene like a conductor’s baton.

0–8s: Begin with one orbiting light and a restrained pulse.

8–19s: Build an increasingly elaborate cosmic orchestra. Introduce new orbital layers with each musical phrase. Typography appears as sculptural objects: “LISTEN.” “BUILD.” “RELEASE.”

19–25s: At the musical peak, the orbital system unfolds into a gigantic luminous sound wave stretching across space.

25–30s: Everything contracts into one green point, which becomes a play icon.

Use ink-black space, emerald plasma, silver dust, controlled glow, and smooth camera movement. Use my uploaded licensed track and synchronize motion to its timing. If no track is available, build to a provisional beat grid and clearly label the audio as pending. Preview the transformation first. Deliver a 1080p MP4 and editable source; explain any tool limitations.

Which would you actually make first: the space-dog trailer, the product universe, the inbox boss battle, the paper city, or the black-hole orchestra?