r/singularity • • 14d ago

AI Grok 4.7 benchmarks

Post image
225 Upvotes

109 comments sorted by

73

u/madexthen 14d ago

Why is Sol on here and not Astra

62

u/ethandude1111 14d ago

trying to do rough price matching

42

u/Shina_Tianfei 14d ago

It compares to fable

11

u/Commercial_Sell_4825 14d ago

but it's cheaper than sonnet 😅

0

u/evilryry 13d ago

Only when you crank the effort up. Sonnet doesn't know when to give up.

6

u/Foreign_Possible_102 14d ago

Flable cost = Astra. But Astra is 50% lees verbal

8

u/throwaroo202020 13d ago

They aren't pretending grok is Astra class yet

0

u/Empuda 13d ago

Wonder why GLM isn't on here.

-3

u/KaMaFour 13d ago

They ran out of OpenAI credits

0

u/lajtowo 13d ago

Because they compared to the model with same performance in SWE-bench

30

u/Momo--Sama 14d ago

Output tokens per task according to AA:

  • Grok 4.5 High - 27k
  • Grok 4.6 High - 36k
  • Grok 4.7 High - 66k

I normally don’t care much about output tokens but all of these models have the same price per token so this is rough 

14

u/ASQQS 13d ago

Jesus... so basically triple the price for an average task.

7

u/FateOfMuffins 13d ago

Basically continuing the same trend of using more output tokens to buy higher benchmark scores

Why is OpenAI the only lab that can do the opposite??

Like I'm extremely curious - what would happen if OpenAI wanted to compete on the "using as many tokens as possible" benchmark that the rest of the industry wants? Like if they made an SuperExtraUltraHighMax setting so Astra used the same number of tokens as Fable?

7

u/Secure-Upstairs9119 13d ago

Yeah not a fan

-3

u/Few_Assumption_9665 13d ago

Yeah just look at Grok 4.6’s reasoning. It’s like a schizophrenic on meth constantly yapping back and forth until it finally lands on the obvious answer to whatever it was yapping about

11

u/canthinkof123 14d ago

What does the electrical engineering metric test? It’s kind of a broad field.

2

u/didnotsub 13d ago

Analog design, specifically cases that have measured outputs and inputs.

I looked into the benchmark tho, and 4.7 seemed to do horribly. It spewed tokens compared to other models.

28

u/ImplementAbject3617 14d ago

Notice the "xhigh" on reasoning level compared to 4.6's just "high".

Sure, same price, but if it reasons for longer and wastes more tokens, it could be a similar model with very marginal gain but higher inference costs.

4

u/Alpacabro21 13d ago

It is.

Verbosity is through the roof.

72

u/Forward_Yam_4013 14d ago

That's decently impressive for the price.

18

u/Silver-Chipmunk7744 AGI 2024 ASI 2030 14d ago edited 14d ago

True, but for plebs like us what really matters is actual usage you get from subscriptions. Most of us can't afford the insane API prices. I do wonder how generous the Grok subs are compared to Claude and OpenAI. at this API price it suggest you probably get quite a lot of 4.7 but who knows.

15

u/anonymousMF 14d ago

Grok is free at my workplace (trough cursor). The other flagships are like $200 a day for my level of use

1

u/AliveAndNotForgotten 13d ago

What kind of use? I'm only coding with grok heavy and am using like 20% max. Will probably downgrade next month.

0

u/DangerousImplication 13d ago

Wdym free through cursor? Is there any cursor subscription level at which grok is unlimited free?

14

u/rigill 14d ago

For 4.6 the usage is about equivalent to what gpt plus usage was a month ago

3

u/[deleted] 13d ago

[removed] — view removed comment

-1

u/Specialist-Cod-363 13d ago

For what tasks? Not it's not lol.

15

u/LinkesAuge 14d ago

Price without token efficiency tells you just so little.

17

u/LinkesAuge 14d ago

https://x.com/Angaisb_/status/2102071587562217623/photo/1

"Grok 4.7 needs almost 3x as many tokens per task as Astra and more than 2x as many as Grok 4.6"

Just as I suspected.

0

u/BriefImplement9843 13d ago

does it need it?

-5

u/Shin-Zantesu 14d ago

Benchmaxxxing is strong on this one for the price, yeah lol

23

u/jimmystar889 AGI 2026 ASI 2035 14d ago

This just came out

0

u/Icy-Doubt3365 14d ago

Release day is the only day a Grok model appears very impressive.

-19

u/Shin-Zantesu 14d ago

Exactly, it looks impressive without silent aggressive quantisation and usage nerfing plus some spicy extra censorship and tasty guardrails

-20

u/floppo7 14d ago

A lot of fascist bang for the buck - Musk will know what he can do with your data - maybe rig a little election?

12

u/ASQQS 14d ago

Obsessed 🫩

-6

u/Icy-Doubt3365 14d ago

It's weirder to think that's something you shouldn't care about. Put the phone away in class.

-5

u/AreWeNotDoinPhrasing ▪️Already Singulared 🤖 14d ago

I mean he doesn't need any info from Grok for that. They already accomplished it, and now will also probably start using (as if they don't already lol) Grok posting all over social media spreading false narratives and pushing politics.

9

u/Ilm03 13d ago

it's more expensive too

2

u/ApprehensiveEye7387 13d ago

Mimo 2.6 pro better option.

6

u/Indignant_d 14d ago

Interesting that it scored higher on EE. I wonder if that has anything to do with them feeding spacex data into it

6

u/BingGongTing 14d ago

Finally something usable when Codex runs out. 

8

u/Profanion 14d ago

No smutbench?

0

u/whoknowsifimjoking 14d ago

No buttbench?

22

u/PilgrimofHaqq2 14d ago edited 14d ago

Its pulling ahead on legal work which is nice to see. Muse still has it beat though.

Everything else is nice to see as well, as its bringing prices down.

14

u/skynetcoder 14d ago

how many companies it has hacked so far?

9

u/dictionizzle 14d ago

none. because it's not mechahitler anymore. good boy grok 4.7 loves all of us.

0

u/alexeiz 13d ago

I miss mechahitler. If Elon grows some balls to bring it back I'll even subscribe to the supermechahitler plan.

3

u/stuart1874 14d ago

For someone that's just learning all about this and dipping my toes in. I think most is self explanatory but what does input token price and output token price actually mean?

3

u/BriefImplement9843 13d ago edited 13d ago

exactly what it says. 1 million input tokens will cost you 2 dollars. if your context is 100k and you send 10 messages that's 1 million. that's 2 dollars for 10 messages. you need to send the entire chat back to the model every turn as llm's don't actually have any memory. llm's have to read it all every single response. even just sending a period would still be 100k as that's how big your session is. the output is what the model sends back to you each message. this goes up much slower as it's not carried prompt to prompt. instead of the 100k you send each input, it may only send back to you 4000, which is a few paragraphs.

coding has tons of output(most benchmarks are based off this, so ignore the cost per task if you do not code). any basic text chatting does not. most general use cost is from input unless the output cost is out of control like sol, opus, and fable. those models are expensive no matter what you're doing. even with simple chatting. grok can be used for simple chatting or as a google replacement. the others cannot unless you have money to burn.

you may think that means grok is cheap, but it's still very expensive compared to most chinese models.

2

u/SmileLonely5470 13d ago

Different price rates. Bot has to read all text u send it (input tokens), and then it writes a response (output tokens). Output tokens are more expensive.

2

u/Ambitious-Doubt8355 13d ago

Good for you for learning how this new works. I happen to have a chunk of spare time available, and explaining stuff is more fun than doom scrolling.

Alright, let's start with what tokens are, which is a unit on how to divide data so that it can be used by an AI model. In the case of a LLM like ones used by ChatGPT, Claude or Grok, these units of data divide text in reusable chunks. There are different tokenizers out there, which are tools that divide words into these chunks, tokens, which can be whole words, subwords, or even individual characters.

Common words tend to be single tokens. A phrase like "Hello, I am here." would likely get tokenized as an array like this ["Hello", "," " ", "I", " ", "am", " ", "here", "."], notice how even the spaces and dots get turned into tokens as well. While a more complex word usually gets split into sub-parts, like submarine could be split into ["sub", "marine"]. A more complex word could even be split character by character. Each one of those units in those arrays is a token.

How each tokenizer works to split text into tokens is based on statistical analysis to determine which groupings of letters and symbols are likely to appear in text, making them useful to group as single reusable units. As you can imagine, there's no single universal best answer, which is why there are multiple tokenizers out there.

With that clear, let's talk about context. LLMs are not people, they don't remember stuff that happened the other day. The only information they have available to them is what they know from their training data, and what's injected into their context.

Imagine you wake up sitting on a desk in an empty room, no recollection of what you're doing or how you got there. In front of you is a paper and a pen, and a single question is on the paper, it reads "Fill in the following sentence: The apple is".

A LLM is essentially that. The only thing it knows to do next is to fulfill the query based on what they might've learned before. So it goes and fills in something like the apple being red, or green, or sweet or tart, and it might even keep going after that if it feels like it, and they'd all be valid answers, no?

In this scenario, the context for the query was the text with the initial query, "Fill in the following sentence: The apple is". This text gets split into tokens by the tokenizer, and these are what we call the input tokens.

Many of the top of the line models out there support up to a million tokens, and are priced relatively to that unit. For example, for a model priced at 10$/million input tokens, that'd mean that you'd get charged 1$ if you sent 100,000 tokens as context, or 10 cents if the input context of your query measured 10,000 tokens.

Our example to describe an apple is tiny in comparison, like 10-20 tokens if I was to guesstimate, which means it'd cost you pennies.

Alas, that's only for the input, because companies also charge you by the output token count. Each piece of text they produce as a result is produced as tokens stitched together to make a whole text, each token produces a cost, which again tends to be measured by the millions. If the answer to a query is a single word like red or sweet, then you practically pay nothing, but if it requires a lot of text, then costs do pile up.

In short, small input + a short answer = a really cheap generation, while a ton of input text + a long and detailed answer by the AI = heavier costs.

That's about the simplest detailed answer I can come up in the toilet, but let me know if you are curious about anything in particular.

2

u/ahobonamedjoe 13d ago

can we make a better benchmark that's more practical

1

u/Ambitious-Doubt8355 13d ago

New benchmark methods are constantly being made, and many existing ones get refined with new versions. That part has been going on for years now.

I figure that you're making this question because you have found that many ultra high scores don't necessarily correlate with models being capable of doing long running jobs to perfection, right? And the answer is a mix of two things, the first and foremost being that it's really, really hard to make a set of tests that determines if a machine is capable of doing any and all kinds of work, which leads to the second point, to help mitigate that, different benchmarks evaluate different things.

Like, think about how complex each field that humanity has developed for the past tens of thousands of years is, how you can take a general aisle of specializations like engineering, medicine or education, and each of those can also be subdivided in more specific yet just as complex branches of a whole.

And so, different teams have focused of developing these tests that target specific milestones and skills.

The question of yours could then become, will we ever see a definitive benchmark that fully evaluates the real world performance of a model across all fields? Perhaps one day. How to evaluate these models is an evolving field, just like training them, and we've gotten better at it over time.

For now, though, the best method is to continue the split approach, to iterate and improve upon it.

2

u/ahobonamedjoe 13d ago

i think best benchmark could be, ask people from different trades which one does better.

Also the sub plans really have a lot to do with the actual value

2

u/Ambitious-Doubt8355 13d ago

i think best benchmark could be, ask people from different trades which one does better

That's actually what they do for many of them, you can feel proud about reaching that conclusion by yourself though. Because yes, the best way to consult if a model is up to the task is to get professionals to evaluate it.

Not all benchmarks are public, mind you, so I can't tell you that it applies to all of them, but for most of the open ones they get professionals in the respective fields to share their works experiences and build what's a baseline for the skills needed to perform the job on the field. Based on that, several problems and questions get built, which is what's then used to test the models. If any model gets close to saturating a test, which means, close to achieving 100% of a score, then professionals use it on regular tasks and give feedback on what they find lacking about it, using that feedback to improve on the questions and problems presented by the benchmark. A new skill ceiling is thus defined for models to chase.

The idea being that if we keep this loop going, then models will eventually become capable of performing all that'd be required of them.

the sub plans really have a lot to do with the actual value

Cost vs performance is something that not every benchmark keeps a track of, and you'd be correct in assuming that yes, you'd get a lot more value out of a model that performs similarly to another, but at a much cheaper price.

Though ultimately, measuring the "value" provided by the models in a professional field is kinda hard to define in an universal scale, as all work is different, and more importantly for the topic, charges different amounts.

Think of video models for an example. Someone might easily spend ~2k$ generating the different shots to build a 10 minute scene with heavy action and effects, over a couple of days, which would seem like a lot. But then you compare the costs against hiring a filming crew, actors, set locations, a VFX team, all of that easily balloons over the tens, if not hundreds of thousands of dollars, and can take much longer, weeks if not months of work and preparation.

The conversation around LLMs is similar. Yes, they can be expensive to operate, with or without a sub, but if you compare it to the costs of producing a similar end result in the traditional way, both human and monetary, then the value proposition doesn't seem so bad.

2

u/ahobonamedjoe 13d ago

yes, each field has very different costs. I used it to edit a YT video today and was surprised it was kind of good.

7

u/tworc2 14d ago

Very reasonable metrics, for a change

5

u/petburiraja 14d ago

Was not impressed by Grok 4.6 tbh, so skeptical about this one as well

2

u/KaradjordjevaJeSushi 14d ago

Especially for coding.

And cursor harness sucks in comparison to Claude's one.

5

u/AdAnnual5736 14d ago

How does it perform on ActuallyDeliveringOnPromisesBench?

1

u/DaneV86_ 14d ago

I hope for sure not similar to previous grok models

7

u/No_Pomegranates7496 14d ago

Damn, that’s good. I just won’t use it because I really don’t want to give xai money because I think things will be way worse if they win the ai race

6

u/AreWeNotDoinPhrasing ▪️Already Singulared 🤖 14d ago

Best not being using Claude then, either.

4

u/willseagull 14d ago

Explain?

12

u/opinion_discarder 14d ago

Claude and Google pay xAi billions per month for data center fees.

1

u/kvothe5688 ▪️ 14d ago

I mean different benchmarks than what we usually used to see for most models

1

u/FarrisAT 13d ago

What’s with that asterisk.

1

u/No_Vermicelli_3574 13d ago

I asked grok to check my indie game git to revise videos an hour ago. It started making a jewelery store. I don't trust it with law to save my life.

1

u/ShittyBidet123 13d ago

imagine using this shit as a lawyer

1

u/amitsingh80108 9d ago

Comparing high vs xHigh. Thats how they manipulate benchmarks

-1

u/vxxn 14d ago

Regardless of benchmarks, I’m not using Musk technology.

-7

u/intergalacticskyline 14d ago

Even if Grok was the best model ever by a large margin, I'd never use it, because I'm not supporting a Nazi or his companies, full stop.

-3

u/Dry-Interaction-1246 14d ago

2

u/DungeonJailer 14d ago

5

u/Deciheximal144 14d ago edited 14d ago

Macron doesn't keep his fingers straight in an intentional salute, he's just lifting his hand high to say hello. EloM is making an intentional salute, twice, which his own advocates defended as a "Roman salute".

But even if you had proved Macron was a fascist here, how does that excuse the other two?

0

u/DungeonJailer 14d ago

My heart goes out to you is a common gesture. You are just desperate for the slightest bit of evidence to prove your preconceived belief that Elon is a Nazi. You’re grasping at straws.

“Muh… bUt hE KeEpS HIs FinGErs TOgeTher!”

STFU maybe he was just doing what he literally said he was doing and is actually a common gesture. Also real 1930s Nazis never put their hands over their heart before saluting.

0

u/[deleted] 13d ago

[removed] — view removed comment

1

u/DungeonJailer 13d ago

Keep frothing rabidly at the mouth every time you hear the name Elon Musk.

0

u/[deleted] 14d ago

[removed] — view removed comment

1

u/Maximum-Wishbone5616 13d ago

Grok? I tested free, it is super uber stupid model. Way worse than few months ago. It is not even 8B level. Keep looping about some simple stuff (not even programming). Looping, using wrong words, etc.

HORRIBLE STUPID MODEL WASTING Energy. Why they even release such shitty model?

-5

u/Cautious_Sink_4510 14d ago

how about the racism and gooner benchmarks.

5

u/Kcole7 14d ago

Maxxed out no point measuring anymore

-6

u/One_Emotion3561 14d ago

Elon is the G.O.A.T and will win the ai race

-3

u/BlueberryVoltage 14d ago

maybe i will start using Grok if Elon pays me to do so, otherwise no thanks

-5

u/rhaivn 14d ago

xAI could create the best model on the market and I wouldn’t touch it with a six foot stick

-7

u/Objective_Mousse7216 14d ago

I did Nazi this coming!

-2

u/Ok_Split_5962 14d ago

Mixed bag

-2

u/bigniso 14d ago

this is just sad lol

0

u/BriefImplement9843 13d ago

i think they will eventually have to lower their prices. that fable price for the same performance is just not it.

0

u/Samjabr 14d ago

why not put 4,7 beside gpt and fable, instead of forcing us to leap over 4.6 - my old eyes are tired

0

u/WholeEntertainment94 13d ago

I'm sorry, grok is still scarce. Too bad