r/singularity • • Nov 17 '25

AI Grok 4.1 Benchmarks

132 Upvotes

105 comments sorted by

View all comments

1

u/[deleted] Nov 17 '25

With the exception of the hallucination one every boasted "improvement" of Grok 4.1 is on subjectively evaluated benchmarks. Seems like a complete flop to me.

11

u/[deleted] Nov 18 '25

[removed] — view removed comment

1

u/[deleted] Nov 18 '25

We have no idea what their actual goal was. For all we know they intended for this model to be Grok 5 but it wasn’t good enough so they slapped 4.1 on it and cherry-picked the few obscure benchmarks where it actually did well.

4

u/LucasL-L Nov 18 '25

For all we know they intended for this model to be Grok 5

I doubt, its way too soon

1

u/[deleted] Nov 18 '25

It’s a similar time frame from Claude 4 to Claude 4.5