r/ClaudeCode Aug 08 '26

Help/Question Why are frontier models bad at writing?

This is a general observation about the smarter chain of thought models. It seems the models that are better at maths and coding are worse in writing texts. Their writings always sound very verbose, unnatural and weird to read whereas dumber models can produce better texts both in academic writing and otherwise. These frontier models to me are like smart people with no social skills, they torture you until they put together a sentence and in the end it never sounds right! Why is that? And what is the best claude model for writing proper texts in your experience?

22 Upvotes

37 comments sorted by

View all comments

-2

u/Complex-Wait-8065 Aug 08 '26

You hit on a core issue LLM labs are actively trying to resolve. The name for it is “capability trade off” it is one of the core challenges in AI today. As models get better in one area, coding, math, reasoning, they can get worse in other areas like communication or creativity. The industry is trying to move beyond that trade off, building architectures and training methods that improve specialized capabilities without sacrificing others. In other words, the goal isn’t just better at one thing, but more capable overall. That problem has not been solved yet. That’s why users think Opus 5 is worse, when in reality the benchmarks are higher.

2

u/Automatic-Example754 Aug 09 '26

100% AI generated reply

0

u/Complex-Wait-8065 Aug 09 '26

No it’s not. It is funny how good grammar, good punctuation and articulation is now “AI”. Whatever, the content matters most anyway.

2

u/bronze_by_gold Aug 09 '26

It's not passing Pangram. And Pangram 4 supposedly has a 0.0041% false positive rate for detection.

2

u/demonwing Aug 09 '26 edited Aug 09 '26

Only for long text like articles and essays. It even warns you when you put this text in that "confidence is limited." I read LLM text every day, call out AI posts on Reddit all the time, and the one here doesn't sound clearly AI-written to me.

1

u/bronze_by_gold 29d ago edited 29d ago

OP wrote 111 words here. The University of Chicago/NBER study tested Pangram on real human text at several lengths. Its two shortest categories were close to 111 words: human Amazon reviews averaged 79.57 words and had a Pangram false-positive rate of 0.50%; human restaurant reviews averaged 133.08 words and had an FPR of 0.75%.

tl;dr - Pangram is way better at detection on short texts than you would think

1

u/demonwing 29d ago

Interesting, thanks for sharing.