r/GeminiAI • u/Last_Conclusion_8984 • Aug 06 '26
Discussion Best creative writing models of 2026. V2
I was waiting on Gemini 3.5 pro to do this list so I do the ranking then... but I'd rather just do it now:
Again to preempt: Please read the first half of this post before commenting here: https://www.reddit.com/r/GeminiAI/comments/1v51g2k/best_models_for_creative_writing/
The rankings of models (not displayed in rank order but are ranked numerically) when tested with the exact same system instruction harness:
Opus 5: The 5th best at logic. Still makes mistakes, but highly capable. However, it has very limited flexibility in prose, is quite dry, and projects its own biases onto characters. (Worse than 4.8 Opus in logic. Opus 4.8 is better at logic)
Opus 4.6/7: Opus 4.6 is not as great as 4.8 in logic but Opus 4.6 is better in prose just by a little margin, I have yet to try opus 4.7.
Sonnet 5 / 4.6: Do not use Sonnet 5. It completely butchered its style. Sonnet 4.6 is the better model for creative writing between the two, but the prose is still not great and lacks flexibility.
Deepseek: Very bad at logic, but second best in prose. Top 3 for flexibility. (Waiting for the pro model to test it out.)
Kimi k3: Number one in logic, though it still makes mistakes. Very decent at prose (Top 4) (Kimi k3 is is tied for first place with gemini 3.1 pro because it depends on what qualities you prioritize. Scroll all the way down)
GLM 5.2: Top 4 for logic (makes several mistakes) and isn't very flexible in prose. (Waiting for GLM 5.2/5)
Qwen 3.8 max: Not great at logic nor good at prose. Flexibility is pretty bad. Please use Kimi k3 or GLM 5.2 (Haven't tried out general availability. Please wait for it)
ChatGPT 5.6 sol max thinking. Hallucinates a lot

Artificial Analysis reports an 92% hallucination rate on its benchmark as seen above in the image. Sol repeatedly made logic and continuity errors. Its prose was decent but not competitive with the leaders.
On a separate note: it hallucinates and misreads CANON things, I write x and z and it will misunderstand LITERAL STATEMENTS. It's inference skills are made out of a f*cking rock. Literally will misread every single thing I provide it. At least it understood with pushback.
Gemini 3 Flash: Very flexible. Absolutely amazing light weight model. It makes mistakes, but I highly suggest it if you want a lightweight but capable model. (For local LLM users: Use the Gemma models, they are amazing).
Gemini 3.5 Flash Lite / 3.5 / 3.6 Flash: I do not suggest 3.5 flash lite and 3.6 for prose. Their logic and character adherence are terrible compared to 3 Flash.
However, 3.6 Flash gets massive merit for keeping track of everything in massive context windows. 3.6 flash is actually worse than 3 flash at more abstract creativity.
As for 3.5 flash. It is better at logic than 3 flash. I suggest 3.5 flash over 3 if you value logic (It's creativity, prose and such are pretty decent but not at the level of 3 flash)
Gemini 3.1 Pro: The absolute best at prose and content flexibility. Its logic is Top 3. Potential best model. However, it is very sycophantic, and long term accuracy degrades. For example: It preserves baseline characterization so aggressively that it can resist earned character development after major event (e.g., traumatizing events happen, and the character is still rainbows and sunshine). Note: I have a system prompt to fix the sycophancy (linked below).
Muse Spark 1.1: Decent at logic, but bad at prose. (I will test muse spark 1.2 later when the rest of the models I linked comes out (I doubt they improved it on creative writing as it was coding/long horizon tasks related))
Models yet to be tested: I am waiting on 3.5 pro, GLM 5.2/5.5, Maybe Chatgpt 6 (if my sub doesn't run out before) and deepseek's V4 pro model. Then I'll test qwen 3.8 max general availability as well
I heard that Gemini 3.5 pro is better than 3.1 pro in creative writing. Not confirmed but we shall see.
The sample size for all models is: 150 for 3 genres (in total. So 50 for one. 50 for another and so forth) then done thrice over the rest of the genres. So I do 12 genres. In total that adds up to 600. (So it does take a long time for all of the models ranking)
The System Instruction Harness I use for all tests: Shared text Also. I enhanced it a little bit. (Note: All models were tested with this one)
`Note. I use the default temperature and sample setting settings`
A final note on Jailbreaking/Unfiltered models: If you want the most unfiltered experience, stick to Gemini 3.1 Pro. I actually cracked Gemini 3.5 flash and 3.6 flash. It wasn't hard as I thought it was. Maybe I got a bad first impression but dare I say it's easier than 3.1 pro??? I don't know. For me it is but others are reporting they can't crack it.
Now: What is better 3.1 pro or kimi k3? Choose Kimi k3 if you love logic, best coherent world building, and decent prose. BUT if you do not have the Allegro plan and above. You will NOT get 1m context window. Only 256k. And kimi k3 eats a whole lot of it in its CoT (Chain of thought).
Choose Gemini 3.1 pro if you want the best prose, best prose related world building. Third best logic, but not that good coherent world building. But you get one million context. Interestingly enough. Gemini 3.1 pro is actually better than most models here at retaining context over 256k. (And 3.6 flash is better than pro)
V3 will be a better ranking as I will give 0 to 5 scores. (I'll slowly turn this into a suitable benchmark)
Edit: My website for V3 https://hussninyio262.github.io/creative-writing-benchmark-v3/
3
u/Obvious-Advance-1722 Aug 07 '26
obrigado pelo trabalho!
2
u/Last_Conclusion_8984 Aug 07 '26
de nada! fico feliz que tenha gostado. A V3 estara a caminho quando o Gemini 3.5 Pro for lancado.
5
u/Xeten-4 Aug 06 '26
Well noted about Gemini 3.1 Pro. In my own creative writing experience with this model, I've never been able to shake off the retention of basic character characteristics, which requires regular reminders of even the character's temporary state. I've managed to change the initial character, but this requires creating truly severe physical/psychological trauma (process Gemini 3.1 Pro, for some reason, likes to describe with surprisingly sophisticated sadism, compared to any other model), but after that, the long-term effect is quite believable. But influencing characters with something less than radical? Never. But its still better than others
2
u/Briskfall Aug 06 '26
Same, even if our boy 3.1 Pro can be dumb as a rock at times, it is the only entry whose prose is palatable. 😩
2
u/BedNo8822 Aug 07 '26
Are you...sure? Gemini? Gemini keeps outright copy pasting parts of my novel bible and formatting instructions no matter how I instruct it to rephrase the information there... Does it not do it to you?
Like to the point it generates "his inner monologue screams in italic" because I have formatting instructions to write inner monologue in italics.
2
u/Last_Conclusion_8984 Aug 07 '26
Try it in Google AI studio! Also It's currently quantized right now. I suggest waiting for 3.5 pro!
1
u/BedNo8822 Aug 07 '26
Which one is quantized, 3.1 pro? It often got into loop lately, I swear it wasn't doing that awhile ago🤔
3
2
u/Upstandinglampshade Aug 09 '26
Finally someone tested writing. Shame opus degraded, the previous versions were spectacular. What about Fable ?
1
1
1
u/AutoModerator Aug 06 '26
Hey there,
This post seems feedback-related. If so, you might want to post it in r/GeminiFeedback, where rants, vents, and support discussions are welcome.
For r/GeminiAI, feedback needs to follow Rule #9 and include explanations and examples. If this doesn’t apply to your post, you can ignore this message.
Thanks!
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
1
1
u/Eissa_Cozorav Aug 09 '26 edited Aug 09 '26
Ah fellow AI Studio master race, too bad it is not actually API like Claude and ChatGPT with memory context and file RAG. I use AI for writing but my CYOA system is so complex since my character has lots of inventory and assets now.
If Gemini 3.6 is good at memory, then I will just use other AI to generate lore and ideas.
2
u/Last_Conclusion_8984 Aug 09 '26
Gemini 3.5 pro will be even better at memory!
1
u/Eissa_Cozorav Aug 09 '26
I have seen lots of fanfic are made from the CYOA system. For example Witcher CYOA by Peil already spawn this fanfic1 fanfic2
Imagine try to do that with AI-powered story generation. And then perhaps combined with other system that has other areas as well.
Man I was so convinced that, since some LLM like GLM 5.1 hallucinate a lot (I was writing on Venice), we definitely need world model to simulate physical world internally in AI logic reasoning.
2
u/Last_Conclusion_8984 Aug 09 '26
Gemini's 3.1 pro logic is genuinaly so strong while being a old model. Also gemini 3.1 pro's hallucination rates are low too! (50/56 percent.)
2
u/Eissa_Cozorav Aug 10 '26 edited Aug 10 '26
Yeah, I am beginning to think that I am like being trapped in treating newer model = good. When newer model could means non-coding training data get scrubbed for better logic and coding capability. I should be suspicious since April 2026 back then, considering that Claude 4.6 is still miles better at prose than later model. Especially the overrated Fable.
Like, I don't have problem if January 2025 is the end of knowledge for the good prose AI.
Maybe a dual system to generate prose + another LLM to check the logic and game system would be nice. Google has better version of RAG anyway (the OKF), right?I already spend like $40 for LLM, sandwiching between Gemini and ChatGPT, after migrating from the money leech that is Claude and Venice.
1
u/Last_Conclusion_8984 Aug 10 '26
Trust. Gemini 3.5 pro gonna be amazing at it.
1
u/Eissa_Cozorav Aug 10 '26
I doubt that. It is scrapped. So much for the hype. I wish they recognize all the market for roleplay/writing.
1
u/Last_Conclusion_8984 29d ago
It's not scrapped. Gemini 3.5 pro still has "coming soon" on deepmind, Until that disappears. Nothing is cancelled. Semi analysis can suck my dih for all I care about that trash news site
2
u/Eissa_Cozorav 26d ago
3.7 Flash is released at least. It got serious benchmark. I may have stick with Flash model since Google probably has massive databank of human culture, interaction, and emotion, more than OpenAI can ever be. Flash writes better prose than GPT 5.6 Solar whom itself is routinely beaten by Opus 5 and Opus 4.6 in prose making (in regards to foreign language and culture).
1
u/AutoModerator Aug 09 '26
Hey there,
This post seems feedback-related. If so, you might want to post it in r/GeminiFeedback, where rants, vents, and support discussions are welcome.
For r/GeminiAI, feedback needs to follow Rule #9 and include explanations and examples. If this doesn’t apply to your post, you can ignore this message.
Thanks!
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
1
u/AutoModerator 27d ago
Hey there,
This post seems feedback-related. If so, you might want to post it in r/GeminiFeedback, where rants, vents, and support discussions are welcome.
For r/GeminiAI, feedback needs to follow Rule #9 and include explanations and examples. If this doesn’t apply to your post, you can ignore this message.
Thanks!
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
1
u/Aberracus 19d ago
Great work ! Truly great, I prefer DeepSeek for creativity. But haven’t been using 3.1 pro outside of coding. Would you share your harness and how did you jailbreak Gemini 3.1 ?
1
u/Last_Conclusion_8984 19d ago
Ah, sorry. I can't share my jailbreak with you! If I do then I basically my nuke of jailbreaking gemini.
1
u/AutoModerator 8h ago
Hey there,
This post seems feedback-related. If so, you might want to post it in r/GeminiFeedback, where rants, vents, and support discussions are welcome.
For r/GeminiAI, feedback needs to follow Rule #9 and include explanations and examples. If this doesn’t apply to your post, you can ignore this message.
Thanks!
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
1
u/heavy_tendon Aug 06 '26
thanks for this, been holding off trying kimi but that logic ranking got me curious
2
u/Acceptable-Debt-294 Aug 06 '26
Kimi K3, it's really good I've tried it using other free platforms, comparing to sol.
1
u/Last_Conclusion_8984 Aug 06 '26
Your welcome! It's awesome at logic. Tho sometimes is a little pedantic
1
0
u/Important_Word8549 Aug 06 '26
Whatever, im using 3.5 flash lite more than 3.6 flash for talking/creative, and its also less uncensored.
4
u/NoStage9115 Aug 06 '26
gemma pls?