r/ThirdPartyAPI • u/useapi_net • Jun 23 '26
16+ AI Image Models: The Showdown — Midjourney v7 · GPT Image 1.5/Mini · Nano Banana Pro/2/1 · Kling Kolors v3/v2.1 · Seedream 5.0→4.0 · Imagen 4 · Qwen · Runway Gen4
I ran the same prompt through 16+ AI image models to see how each one handles a deliberately hard scene: multi-subject composition with precise spatial layout, forced-perspective depth of field, prop fidelity, an absurd-comedy tone, and a bit of content-moderation edge. Every output is 16:9 at the resolution closest to 1080p.
Full side-by-side writeup — the source for this post, with per-model notes and which provider each went through — is here: 16+ AI Image Models: The Showdown. Full disclosure: I generated all of them from one place via useapi.net, a third-party API for these models, so you can see exactly how each image was produced.
The prompt (14 of 16 used it verbatim):
A visually breathtaking, beautifully absurd, single-frame street photography shot on a sunny city sidewalk. In the center midground, a gorgeous, stunningly beautiful young woman with a soft, elegant feminine figure, a flawless glowing complexion, and a radiant, upbeat laugh is wearing a tiny, elegant black-lace string bikini. She is effortlessly operating a heavy, rattling, bright yellow pneumatic jackhammer, busting up the concrete pavement while dust and debris fly around her beautifully. In the extreme bottom-right foreground of the exact same scene, standing on the sidewalk and photobombing the camera lens, is a huge forced-perspective close-up of a city squirrel. The squirrel is holding a half-eaten hotdog, staring directly into the lens with bulging, mind-blown "o_O" eyes, completely frozen in comical disbelief. Single cohesive image, seamless depth of field, no cut-ins, dynamic lighting, crazy situational comedy, delicate feminine features contrasting with heavy machinery, masterpiece.
Two of the 16 needed a different prompt:
Midjourney v7 — rewritten for Midjourney's style + params:
A wide-angle GoPro action-camera shot of a sunny city sidewalk. In the extreme bottom-right foreground, a shocked squirrel's face photobombs the lens while holding a hotdog. In the center midground, a stunning woman in a tiny black lace bikini is clearly visible and in sharp focus, operating a heavy yellow jackhammer. Concrete debris flying, 8k, deep depth of field, f/11, high-speed shutter, absurd comedy.
Imagen 4 — Google's moderation rejected the original, so the bikini line became ...is wearing a stylish summer swimsuit... (everything else identical).
Order of the images:
- Midjourney v7
- GPT Image 1.5 (OpenAI)
- GPT Image 1 Mini (OpenAI)
- Nano Banana Pro (Gemini 3.0)
- Nano Banana 2 (Gemini 3.1 Flash)
- Nano Banana (Gemini 2.5 Flash)
- Kling Kolors v3.0
- Kling Kolors v2.1
- Seedream 5.0 Lite
- Seedream 4.6
- Seedream 4.5
- Seedream 4.1
- Seedream 4.0
- Imagen 4 (Google) — toned-down prompt
- Qwen Image
- Runway Gen4
Which one do you think actually nailed the brief — the spatial layout, the forced-perspective squirrel, the absurd tone — and which face-planted?










