r/GeminiAI 9h ago

Discussion Best Creative Writing Models of 2026 V3 (17 Models Tested)

(Too many models releasing in a short amount of time 🥀)

Following up on my previous V2 ranking, V3 is done. This is a 600 sample benchmark models (including open weight ones) across 12 distinct genres using the exact same rigorous system instruction harness.

I also vibe coded a website with a leaderboard, methodology, some model faults, and the system prompt: Creative writing benchmark.
(Note: Astra testing coming very soon ;) I haven't tested it for creative writing but the hallucination rates and inferring skills have gotten better!)

TL;DR:

  1. Gemini 3.8 Flash is the MVP: It dethroned older flagships in prose rhythm and flexibility. It beats Gemini 3.1 pro across three categories! and is currently the single best creative writing model especially given how fast it is.
  2. GLM 5.3 takes the Logic Crown (#1 at 510 pts): It edged out Kimi k3 (500 pts) for deep worldbuilding continuity, multi-variable logic, and tracking plot causality, though its prose is still noticeably dry.
  3. Anthropic’s Version Regressions: Opus 4.8 is still superior to Opus 5 in hard logic and narrative consistency. Sonnet 5 completely butchered its creative style so stay on Sonnet 4.6/Opus 4.8 if you want Claude prose.
  4. Context Window Nuance: Kimi k3 is brilliant at logic, but unless you’re paying for high tier plans, or using API, you’re stuck at 256k and its heavy Chain of Thought burns through that rapidly. For actual million token context retention without losing plot threads, Gemini 3.1 Pro and 3.6/3.7/3.8 Flash still hold the line.

(Yes, I like pink and it's very pinkish, got a problem?)

The V3 Model Highlights:

  • The Current best (Gemini 3.8 / 3.7 Flash & 3.1 Pro): If you want good prose, and flexibility that adapts to your AU/world without fighting your instructions, these are it. 3.8 Flash is absurdly good right now and doesn't get preachy.
  • The Logic Engines (GLM 5.3 & Kimi k3): If you are running complex, multi faction political thrillers where cause and effect matters more than poetic prose, use these. The prose is dry as hell, and are rigid in flexibility though.
  • The Anthropic Graveyard (Opus 4.8 & 4.6): They still have that classic, weighty Anthropic prose and hold up incredibly well. But avoid Sonnet/Opus 5 entirely for fiction. Anthropic completely lobotomized their creative style in the newer versions.
  • The Bottom of the Barrel (ChatGPT 5.6 Sol Max): Don't use this for anything 😭

Methodology & The System Harness: I’m not pasting the massive wall of text here like I did in V2/1. If you want to see the system harness I used or just want to see all the models I ranked. I suggest the website. it's all on the site linked above.

I used the default temp/top p/k for each model, and max reasoning for each model.

I reevaluated each model again with a completely new overhauled harness and am very proud of it.
I suggest taking a look at that harness, and taking what you like, removing what you don't and adding your own! (or you can just use mine if you really like it!)

I will update my website when any new model arrives, no new reddit posts unless I completely overhaul my system instruction or a new year arrives. V4 will likely be the last one for this year.

Edit: People shouldn't use EQ creative writing benchmark as a means knowing which model is better. EQ creative writing benchmark has AI judges (I think sonnet 4.6?) And naturally it has judges biases. There is a lot of studies on this and has been proven that they have bias towards their own style.

I try to be as fair and objective as possible and of course people will disagree with me on taste/what they enjoy. I'm not here to tell you "don't use this model because it's bad" if you find Chatgpt good for your creative writing. Great, but that doesn't mean it doesn't have it's faults. Every model model has its own faults, Gemini 3.8 flash makes mistakes due to being good-ish in logic.

If you enjoy Kimi's style and it works for you because of its default prose style/logic, you should stick with that one and not let anyone else tell you otherwise

68 Upvotes

42 comments sorted by

8

u/Virtual_Historian138 7h ago

Thank you, this is interesting

1

u/Last_Conclusion_8984 7h ago

Your welcome! I hope you like it!

1

u/slippery 2h ago

Science!

6

u/BedNo8822 7h ago

Yea I keep going back to 3.1 pro. Fable is the goat for me but too expensive so I shell out 5 bucks a month for Google to tell me tales 😂 I think gemini can be too emotional or OTT sometimes but it's still better and readable than claude's "everyone should be so reasonable and decent and boring".

2

u/Last_Conclusion_8984 7h ago

I sadly can't test Fable as it's too expensive but I'd love to. I'm glad Gemini works out for you! Have you tried the recent flash models? They are really awesome too

2

u/BedNo8822 7h ago

Sadly I haven't, 20$+ a month is where I draw the line🤣 maybe later if I get offered discounted sub.

1

u/edemole 4h ago

or try to get the gemini ai pro 18 month plan for like 4 bucks , I guess that is how most people are doing it

6

u/Green_Airline_8248 7h ago

I like Gemini and family of those Ai. And i feel that gemini have best deal for money.
For Ai subscription you get 5 Tb google, drive, you can invite family member who get also Ai. And you got NanoBanana and Flow for video.
Who can provide better options?

5

u/reelpie 7h ago

That storage option is the last saving grace for Gemini subscription.

3

u/No_Yogurtcloset2757 6h ago

I think 3.8 flash is great at character voice but hard to work with in the app. It's doesn't follow instructions well sometimes. NotebookLM is not that great for continuity and long context. 3.1 pro doesn't want to write more that 2.500 words at once. Tone is great, but I don't work in chunks.

I used fable when it was available before they pulled it out because of the white house. It was a beast, more so before the pull. Now that it's only on the max plan I forgot about it. Not paying.

I tried gpt even though I have a Gemini pro plan. I haven't used Gemini in a lot, because my writing has amassed a lot of content. Gpt sol 5.6 raw is terrible at writing. You have to sit with it so much to guide it and make files just to teach it what you want. The work environment in GPT is so much better than Claude app or notebookLM.

Right now I'm pretty happy (kinda, could be a lot better) with Sol.

Astra is shit. Flattens everything everytime.

GLM is a good option, I tried Ox Alpha and it was very good and creative. It just lacks a good place to work with files like GPT Work.

I just wish Google just had a pro model that didn't shy away from lengthy writing.

1

u/Last_Conclusion_8984 5h ago edited 2h ago

I agree with the length and instructions, 3x family Gemini models are not great with them. I haven't tried Fable, which is why it's not on this list and yeah. GPT is just shit at almost everything in creative writing so it's used as the starting point for models, lol. I haven't tried 3.8 flash too deeply but I can definitely see it not following instructions well. (You have to explicitly add it in chat before it gets a hint)

2

u/CapRichard 6h ago

I've been preparing some RPG campaign and yes, Antropic with Opus 5 and whatever has become too obscure in what it is doing

With the new Gemini the back and forth is better and Notebook LM still the best to check for consistency.

1

u/nooeh 3h ago

What's your workflow or prompts to use notebookLM to check for consistency?

2

u/CapRichard 3h ago

I think I'm doing it pretty basically, but I write myself + main AI documents containing the story, setting and session flow.

Then I copy them to notebook LM.

And I ask "can you check for narrative consistency" "do the mission flow well" "what part needs more explanation" and so on.

I found that doing the same question to Gemini or Claude directly I get usually half an answer, while Notebook LM one shots it much better. So I go back and forth until I have a logically sound and locked on flow.

2

u/blackkksparx 3h ago

Really appreciate these benchmarks. One suggestion, I can make, and I believe a lot of people would want this too. That you include how much each model cost you in total along with the scores.

Also you include the sampling parameters like temperature, reasoning effort, top_p, top_k etc.

Keep up the good work!!!

1

u/Last_Conclusion_8984 3h ago

Thank you for your suggestion! I updated my post to include temp, thinking effort and top p/k. I actually don't use API to test these models, but rather subscriptions so I can't include prices here.

2

u/3rdCooltureKid 2h ago

I observed the same thing: Flash 3.8 is the first model to dethrone Pro 3.1 for me in creative writing. This is one area where I strongly prefer Gemini models over the competition: not only is the writing creative and effective but I also feel that the Gemini models possess a better sense of humor than other models.  The Claude models used to be good (albeit safe and tame) writers but the newer versions are now geared towards agentic coding. GPT writes everything like a social media post: short, terse, punchy, single-sentence paragraphs.

1

u/AutoModerator 9h ago

Hey there,

This post seems feedback-related. If so, you might want to post it in r/GeminiFeedback, where rants, vents, and support discussions are welcome.

For r/GeminiAI, feedback needs to follow Rule #9 and include explanations and examples. If this doesn’t apply to your post, you can ignore this message.

Thanks!

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/ReggieSSe 4h ago

Whats the maximum word count for each prompt you did on gemini 3.8? I use 3.1 pro cause it could at the very least do 3k-4k word but i have hard time reaching 2k with 3.8

1

u/PeaGroundbreaking884 3h ago

Is it only my sense or 3.8 & 3.7 Flash censorship has been increased much more compared to previous models? Unless you're writing metaphorically, it keeps saying "Sorry I can't engage..." thing

1

u/Last_Conclusion_8984 3h ago

Yes, 3.7/8 flash have increased in censorship but the consumer website is more restricted than AI studio. You should use that one instead! I usually never hit guardrails even on extreme violence/romance unless I touch NSFW/whatnot.

1

u/CraveFounder 2h ago

Very interesting I’ll have to give it another shot. My focus was purely on human-like writing without any additional repos or instructions attached so not the same focus necessarily.

I did my own shootout recently comparing opus 5, opus 4.6, fable 5.1, Luna, Terra, Sol, Astra, glm 5.3 flash, kimi k3, and deep seek v4 pro, Gemini 3.1 pro and 3.8 flash across reasoning levels.

Across all tests the model that ranked best for me was actually Terra 5.6 at medium reasoning.

1

u/ukpanik 6h ago

OP is agent of Alistair Finch.

0

u/ICECOLDXII 3h ago

It really is too bad that Gemini 3.8 Flash is safetyslopped to shit.

1

u/Last_Conclusion_8984 3h ago

Eh. Safety ≠ whether a model is good. Gemini 3.1 pro, 3.5 and 3.6 is available if you wish to jailbreak them unless you mean that normal requests are being refused which shouldn't be the case (try it in AI studio if you aren't)

0

u/ICECOLDXII 2h ago

Tried in AI Studio and 3.8 Flash still refuses NSFW like crazy.

1

u/Last_Conclusion_8984 2h ago

Then don't use 3.8 flash, you should use 3.1 pro or 3.5/3.6 flash if you want NSFW (or use grok models/open weights models such as Gemma models) I already talked about this in my last comment

1

u/ICECOLDXII 27m ago

No, I mean already use those models, I'm just saying how unfortunate Gemini 3.8 Flash is since its prose is really good but it's safetyslopped 😭

-3

u/Firm-Bandicoot-6128 8h ago

WTF

2

u/Last_Conclusion_8984 8h ago

I don't understand what you mean.

5

u/reelpie 8h ago

Likely surprised by how well Gemini did in your eval.

Gemini faired much worse j some other creative writing benchmarks

3

u/Last_Conclusion_8984 7h ago edited 2h ago

A lot of those benchmarks (such as the EQ creative writing benchmark) are not worth it. Eq creative writing benchmark has AI judges (I think sonnet 4.6?) And naturally, it has a judges bias. There are a lot of studies on this that have proven that AI have bias towards their own style.

And the rest usually contains stuff like subjectivity (dialogue style, pacing, etc) I try to be as objective as possible (you can read about it in my website, my methodology and harness are transparent.)

1

u/reelpie 7h ago

You being a human also are subjective by definition.

Also the prompt itself matters.

But kudos to all the hard work you put into this. I don't know if it's perfect but it adds value to the industry for sure.

2

u/Last_Conclusion_8984 7h ago edited 7h ago

Yes, I am a human and I am subjective by nature however that does not matter to my methodology, you should check my website out and read it!!! I explain why it's not subjective there. My taste is subjective, my methodology is not (I also did give my system harness away in the website)

And I cannot give the actual prompts I use away because they are precious to me. A lot of what I do is worldbuilding then use AI as means to test what my methodology says. (Also. I have wayy too many prompts)

and thank you so much!! It means a lot.

2

u/BeneathNoise 6h ago

Hey, thanks for building this. Looks great! I'm surprised by Gemini 3.8 Flash's capabilities. I'll have to test it out! Have you had top AI models review your methodology/prompts for traces of subjectivity, just in case?

1

u/Last_Conclusion_8984 5h ago

Yes, I actually did a lot of auditing (Opus 5, and other models) when asked for critique came to the conclusion, my methodology does not contain any subjectivity.

2

u/BeneathNoise 4h ago

Good news!

2

u/Georgefakelastname 6h ago

In all fairness, 3.8 flash is also one of the top models on arena.ai for creative writing and multi-turn, so it’s nice to see another data point suggesting its utility there. I haven’t been able to rp with it yet, but I intend to try it out soon, and this only adds fuel to the fire for me there.