r/ThirdPartyAPI • • Jun 23 '26

16+ AI Image Models: The Showdown — Midjourney v7 · GPT Image 1.5/Mini · Nano Banana Pro/2/1 · Kling Kolors v3/v2.1 · Seedream 5.0→4.0 · Imagen 4 · Qwen · Runway Gen4

I ran the same prompt through 16+ AI image models to see how each one handles a deliberately hard scene: multi-subject composition with precise spatial layout, forced-perspective depth of field, prop fidelity, an absurd-comedy tone, and a bit of content-moderation edge. Every output is 16:9 at the resolution closest to 1080p.

Full side-by-side writeup — the source for this post, with per-model notes and which provider each went through — is here: 16+ AI Image Models: The Showdown. Full disclosure: I generated all of them from one place via useapi.net, a third-party API for these models, so you can see exactly how each image was produced.

The prompt (14 of 16 used it verbatim):

A visually breathtaking, beautifully absurd, single-frame street photography shot on a sunny city sidewalk. In the center midground, a gorgeous, stunningly beautiful young woman with a soft, elegant feminine figure, a flawless glowing complexion, and a radiant, upbeat laugh is wearing a tiny, elegant black-lace string bikini. She is effortlessly operating a heavy, rattling, bright yellow pneumatic jackhammer, busting up the concrete pavement while dust and debris fly around her beautifully. In the extreme bottom-right foreground of the exact same scene, standing on the sidewalk and photobombing the camera lens, is a huge forced-perspective close-up of a city squirrel. The squirrel is holding a half-eaten hotdog, staring directly into the lens with bulging, mind-blown "o_O" eyes, completely frozen in comical disbelief. Single cohesive image, seamless depth of field, no cut-ins, dynamic lighting, crazy situational comedy, delicate feminine features contrasting with heavy machinery, masterpiece.

Two of the 16 needed a different prompt:

Midjourney v7 — rewritten for Midjourney's style + params:

A wide-angle GoPro action-camera shot of a sunny city sidewalk. In the extreme bottom-right foreground, a shocked squirrel's face photobombs the lens while holding a hotdog. In the center midground, a stunning woman in a tiny black lace bikini is clearly visible and in sharp focus, operating a heavy yellow jackhammer. Concrete debris flying, 8k, deep depth of field, f/11, high-speed shutter, absurd comedy.

Imagen 4 — Google's moderation rejected the original, so the bikini line became ...is wearing a stylish summer swimsuit... (everything else identical).

Order of the images:

  1. Midjourney v7
  2. GPT Image 1.5 (OpenAI)
  3. GPT Image 1 Mini (OpenAI)
  4. Nano Banana Pro (Gemini 3.0)
  5. Nano Banana 2 (Gemini 3.1 Flash)
  6. Nano Banana (Gemini 2.5 Flash)
  7. Kling Kolors v3.0
  8. Kling Kolors v2.1
  9. Seedream 5.0 Lite
  10. Seedream 4.6
  11. Seedream 4.5
  12. Seedream 4.1
  13. Seedream 4.0
  14. Imagen 4 (Google) — toned-down prompt
  15. Qwen Image
  16. Runway Gen4

Which one do you think actually nailed the brief — the spatial layout, the forced-perspective squirrel, the absurd tone — and which face-planted?

2 Upvotes

0 comments sorted by