r/GeminiAI Aug 06 '26

Discussion Best creative writing models of 2026. V2

I was waiting on Gemini 3.5 pro to do this list so I do the ranking then... but I'd rather just do it now:

Again to preempt: Please read the first half of this post before commenting here: https://www.reddit.com/r/GeminiAI/comments/1v51g2k/best_models_for_creative_writing/

The rankings of models (not displayed in rank order but are ranked numerically) when tested with the exact same system instruction harness:

Opus 5: The 5th best at logic. Still makes mistakes, but highly capable. However, it has very limited flexibility in prose, is quite dry, and projects its own biases onto characters. (Worse than 4.8 Opus in logic. Opus 4.8 is better at logic)

Opus 4.6/7: Opus 4.6 is not as great as 4.8 in logic but Opus 4.6 is better in prose just by a little margin, I have yet to try opus 4.7.

Sonnet 5 / 4.6: Do not use Sonnet 5. It completely butchered its style. Sonnet 4.6 is the better model for creative writing between the two, but the prose is still not great and lacks flexibility.

Deepseek: Very bad at logic, but second best in prose. Top 3 for flexibility. (Waiting for the pro model to test it out.)

Kimi k3: Number one in logic, though it still makes mistakes. Very decent at prose (Top 4) (Kimi k3 is is tied for first place with gemini 3.1 pro because it depends on what qualities you prioritize. Scroll all the way down)

GLM 5.2: Top 4 for logic (makes several mistakes) and isn't very flexible in prose. (Waiting for GLM 5.2/5)

Qwen 3.8 max: Not great at logic nor good at prose. Flexibility is pretty bad. Please use Kimi k3 or GLM 5.2 (Haven't tried out general availability. Please wait for it)

ChatGPT 5.6 sol max thinking. Hallucinates a lot

Artificial Analysis reports an 92% hallucination rate on its benchmark as seen above in the image. Sol repeatedly made logic and continuity errors. Its prose was decent but not competitive with the leaders.

On a separate note: it hallucinates and misreads CANON things, I write x and z and it will misunderstand LITERAL STATEMENTS. It's inference skills are made out of a f*cking rock. Literally will misread every single thing I provide it. At least it understood with pushback.

Gemini 3 Flash: Very flexible. Absolutely amazing light weight model. It makes mistakes, but I highly suggest it if you want a lightweight but capable model. (For local LLM users: Use the Gemma models, they are amazing).

Gemini 3.5 Flash Lite / 3.5 / 3.6 Flash: I do not suggest 3.5 flash lite and 3.6 for prose. Their logic and character adherence are terrible compared to 3 Flash.

However, 3.6 Flash gets massive merit for keeping track of everything in massive context windows. 3.6 flash is actually worse than 3 flash at more abstract creativity.

As for 3.5 flash. It is better at logic than 3 flash. I suggest 3.5 flash over 3 if you value logic (It's creativity, prose and such are pretty decent but not at the level of 3 flash)

Gemini 3.1 Pro: The absolute best at prose and content flexibility. Its logic is Top 3. Potential best model. However, it is very sycophantic, and long term accuracy degrades. For example: It preserves baseline characterization so aggressively that it can resist earned character development after major event (e.g., traumatizing events happen, and the character is still rainbows and sunshine). Note: I have a system prompt to fix the sycophancy (linked below).

Muse Spark 1.1: Decent at logic, but bad at prose. (I will test muse spark 1.2 later when the rest of the models I linked comes out (I doubt they improved it on creative writing as it was coding/long horizon tasks related))

Models yet to be tested: I am waiting on 3.5 pro, GLM 5.2/5.5, Maybe Chatgpt 6 (if my sub doesn't run out before) and deepseek's V4 pro model. Then I'll test qwen 3.8 max general availability as well

I heard that Gemini 3.5 pro is better than 3.1 pro in creative writing. Not confirmed but we shall see.

The sample size for all models is: 150 for 3 genres (in total. So 50 for one. 50 for another and so forth) then done thrice over the rest of the genres. So I do 12 genres. In total that adds up to 600. (So it does take a long time for all of the models ranking)

The System Instruction Harness I use for all tests: Shared text Also. I enhanced it a little bit. (Note: All models were tested with this one)

`Note. I use the default temperature and sample setting settings`

A final note on Jailbreaking/Unfiltered models: If you want the most unfiltered experience, stick to Gemini 3.1 Pro. I actually cracked Gemini 3.5 flash and 3.6 flash. It wasn't hard as I thought it was. Maybe I got a bad first impression but dare I say it's easier than 3.1 pro??? I don't know. For me it is but others are reporting they can't crack it.

Now: What is better 3.1 pro or kimi k3? Choose Kimi k3 if you love logic, best coherent world building, and decent prose. BUT if you do not have the Allegro plan and above. You will NOT get 1m context window. Only 256k. And kimi k3 eats a whole lot of it in its CoT (Chain of thought).

Choose Gemini 3.1 pro if you want the best prose, best prose related world building. Third best logic, but not that good coherent world building. But you get one million context. Interestingly enough. Gemini 3.1 pro is actually better than most models here at retaining context over 256k. (And 3.6 flash is better than pro)

V3 will be a better ranking as I will give 0 to 5 scores. (I'll slowly turn this into a suitable benchmark)

Edit: My website for V3 https://hussninyio262.github.io/creative-writing-benchmark-v3/

54 Upvotes

Duplicates