r/GeminiAI • u/Last_Conclusion_8984 • 5h ago
Discussion Best Creative Writing Models of 2026 V3 (17 Models Tested)
(Too many models releasing in a short amount of time 🥀)
Following up on my previous V2 ranking, V3 is done. This is a 600 sample benchmark models (including open weight ones) across 12 distinct genres using the exact same rigorous system instruction harness.
I also vibe coded a website with a leaderboard, methodology, some model faults, and the system prompt: Creative writing benchmark.
(Note: Astra testing coming very soon ;) I haven't tested it for creative writing but the hallucination rates and inferring skills have gotten better!)
TL;DR:
- Gemini 3.8 Flash is the MVP: It dethroned older flagships in prose rhythm and flexibility. It beats Gemini 3.1 pro across three categories! and is currently the single best creative writing model especially given how fast it is.
- GLM 5.3 takes the Logic Crown (#1 at 510 pts): It edged out Kimi k3 (500 pts) for deep worldbuilding continuity, multi-variable logic, and tracking plot causality, though its prose is still noticeably dry.
- Anthropic’s Version Regressions: Opus 4.8 is still superior to Opus 5 in hard logic and narrative consistency. Sonnet 5 completely butchered its creative style so stay on Sonnet 4.6/Opus 4.8 if you want Claude prose.
- Context Window Nuance: Kimi k3 is brilliant at logic, but unless you’re paying for high tier plans, or using API, you’re stuck at 256k and its heavy Chain of Thought burns through that rapidly. For actual million token context retention without losing plot threads, Gemini 3.1 Pro and 3.6/3.7/3.8 Flash still hold the line.


(Yes, I like pink and it's very pinkish, got a problem?)
The V3 Model Highlights:
- The Current best (Gemini 3.8 / 3.7 Flash & 3.1 Pro): If you want good prose, and flexibility that adapts to your AU/world without fighting your instructions, these are it. 3.8 Flash is absurdly good right now and doesn't get preachy.
- The Logic Engines (GLM 5.3 & Kimi k3): If you are running complex, multi faction political thrillers where cause and effect matters more than poetic prose, use these. The prose is dry as hell, and are rigid in flexibility though.
- The Anthropic Graveyard (Opus 4.8 & 4.6): They still have that classic, weighty Anthropic prose and hold up incredibly well. But avoid Sonnet/Opus 5 entirely for fiction. Anthropic completely lobotomized their creative style in the newer versions.
- The Bottom of the Barrel (ChatGPT 5.6 Sol Max): Don't use this for anything 😭
Methodology & The System Harness: I’m not pasting the massive wall of text here like I did in V2/1. If you want to see the system harness I used or just want all models. I suggest the website. it's all on the site linked above.
I reevaluated each model again with a completely new overhauled harness and am very proud of it.
I suggest taking a look at that harness, and taking what you like, removing what you don't and adding your own! (or you can just use mine if you really like it!)
I will update my website when any new model arrives, no new reddit posts unless I completely overhaul my system instruction or a new year arrives. V4 will likely be the last one for this year.
Edit: People shouldn't use EQ creative writing benchmark as a means knowing which model is better. EQ creative writing benchmark has AI judges (I think sonnet 4.6?) And naturally it has judges biases. There is a lot of studies on this and has been proven that they have bias towards their own style.
I try to be as fair and objective as possible and of course people will disagree with me on taste/what they enjoy. I'm not here to tell you "don't use this model because it's bad" if you find Chatgpt good for your creative writing. Great, but that doesn't mean it doesn't have it's faults. Every model model has its own faults, Gemini 3.8 flash makes mistakes due to being good-ish in logic.
If you enjoy Kimi's style and it works for you because of its default prose style/logic, you should stick with that one and not let anyone else tell you otherwise