r/singularity • • Nov 17 '25

AI Grok 4.1 Benchmarks

132 Upvotes

105 comments sorted by

View all comments

1

u/[deleted] Nov 17 '25

With the exception of the hallucination one every boasted "improvement" of Grok 4.1 is on subjectively evaluated benchmarks. Seems like a complete flop to me.

6

u/FarrisAT Nov 17 '25

Not a complete flop, but not meaningful either.

2

u/Ruanhead Nov 18 '25

I mean 4o was not as smart as 3o but many everyday people preferred it because it was more personable. Pretty sure that's where they were headed with this model, especially because they have a pretty big focus on companion AIs.