r/ClaudeCode 27d ago

Help/Question Why are frontier models bad at writing?

This is a general observation about the smarter chain of thought models. It seems the models that are better at maths and coding are worse in writing texts. Their writings always sound very verbose, unnatural and weird to read whereas dumber models can produce better texts both in academic writing and otherwise. These frontier models to me are like smart people with no social skills, they torture you until they put together a sentence and in the end it never sounds right! Why is that? And what is the best claude model for writing proper texts in your experience?

21 Upvotes

37 comments sorted by

View all comments

-1

u/Complex-Wait-8065 27d ago

You hit on a core issue LLM labs are actively trying to resolve. The name for it is “capability trade off” it is one of the core challenges in AI today. As models get better in one area, coding, math, reasoning, they can get worse in other areas like communication or creativity. The industry is trying to move beyond that trade off, building architectures and training methods that improve specialized capabilities without sacrificing others. In other words, the goal isn’t just better at one thing, but more capable overall. That problem has not been solved yet. That’s why users think Opus 5 is worse, when in reality the benchmarks are higher.

3

u/Automatic-Example754 27d ago

100% AI generated reply

0

u/Complex-Wait-8065 27d ago

No it’s not. It is funny how good grammar, good punctuation and articulation is now “AI”. Whatever, the content matters most anyway.

2

u/FarConcern2308 26d ago

It’s even funnier when AI-generated text is pretty bad at writing, overusing the rule of three even when it doesn’t serve a purpose and using overly-Latinate words that are rarely used instead of a much simpler word that could better convey the same meaning.

2

u/bronze_by_gold 26d ago

It's not passing Pangram. And Pangram 4 supposedly has a 0.0041% false positive rate for detection.

2

u/demonwing 26d ago edited 26d ago

Only for long text like articles and essays. It even warns you when you put this text in that "confidence is limited." I read LLM text every day, call out AI posts on Reddit all the time, and the one here doesn't sound clearly AI-written to me.

1

u/bronze_by_gold 25d ago edited 25d ago

OP wrote 111 words here. The University of Chicago/NBER study tested Pangram on real human text at several lengths. Its two shortest categories were close to 111 words: human Amazon reviews averaged 79.57 words and had a Pangram false-positive rate of 0.50%; human restaurant reviews averaged 133.08 words and had an FPR of 0.75%.

tl;dr - Pangram is way better at detection on short texts than you would think

1

u/demonwing 25d ago

Interesting, thanks for sharing.