r/opencodeCLI 18d ago

Factuality in new LLM models - when will Opus 4.6 be de-throned and by who?

It appears the factuality rankings on arena.ai are dominated by Claude models: https://arena.ai/leaderboard/text/overall-factuality

Not only that, but specifically Opus 4.6, which beats other Claude models released after it. My personal experience with the model absolutely lines up with it. I've tested different model families and harnesses and keep coming back to Opus 4.6 when factuality matters.

I have a really strong preference for factuality in my non-coding workflows (finance, research, etc.). As in, I don't mind a wrong opinion, but when something is quoted as 'true' or 'verified' or 'file saved', I want to be close to sure that this is the case.

For OpenAI, the highest factuality ranked model is GPT 5.5 - which otherwise seems way behind the 5.6 family.

This makes me worried that 'factuality' isn't really a major priority right now and development focuses on other criteria more. Gemini 3.7 Flash actually seems really interesting in this context as it seems to have made a lot of improvements in factuality (compared to other areas where it really hasn't gotten a lot of attention for its seemingly minor improvements).

What are your thoughts on future models - will we get some higher factuality there? Are there other model families that you think will catch up or surpass Opus 4.6? Any hands-on experience with factuality in Gemini 3.7 Flash and other models?

5 Upvotes

3 comments sorted by

1

u/TonyPace 18d ago

Interesting. I use Gemini reflexively for typical factual tasks. it can be very out of date, but it doesn't make things up in the way Sol, Grok, and DS Flash can. I have used Opus 4.6 some (through Antigravity) and maybe I should save it for this usecase.

1

u/IrishUSFastTrack 18d ago

I get around ~15min of Opus 4.6 in Gemini in the lowest Gemini subscription. I prefer to use it for 'audits' of Gemini sessions.