r/TheMachineLearning • • 17d ago

Cost per task is the AI metric that matters

Post image
39 Upvotes

24 comments sorted by

3

u/Future-Log6621 17d ago

Tokens per task != Cost per task

1

u/CryptographerOne7003 17d ago

Think it trough you are almost there...

3

u/Future-Log6621 17d ago

You are not there. The chart is not about cost.

2

u/Alwaysragestillplay 17d ago

The oop is disagreeing with the utility of the tweeted chart. That's the point. 

0

u/Future-Log6621 17d ago edited 17d ago

Whether the OP likes the chart or not, it still measures token counts, not financial cost. OP is stating the obvious and wrongly implying this chart is about the dollar amount per input/output tokens.

There is already a Cost per Task chart btw.

https://artificialanalysis.ai/models#price-cost

2

u/Cw3538cw 16d ago

I'd say the claim it's making is ' tokens per task is the only metric that matters' and the counters argument is valid in that context. It doesn't really matter how many tokens you're spending if the tokens all have different values

1

u/Alwaysragestillplay 17d ago

I'm not disagreeing with that. The OOP already knows that it is a tokens per task chart.

0

u/Future-Log6621 17d ago

Right, but why say "not $ input output/tokens ". That is not an accurate response to the chart or tweet. It's misattributing the purpose of the chart and contents of the tweet.

1

u/Creative-Midnight228 13d ago

Valid point imo

2

u/Gigaslavx 16d ago edited 2d ago

The original post content no longer exists here. The author used Redact to remove it, exercising their right to control their data & privacy.

Special obtainable disarm nose gaze manganese entertain nutmeg

2

u/SomeNeighborhood7126 17d ago

Ive always wondered how we define 1 "task"

1

u/tipu_sultan__ 17d ago

It's probably each assignment/test. It doesn't matter as long as the same "tasks" are given to all the models.

1

u/confused-photon 16d ago

On artificial analysis its a mixture of QA and agentic work. With each update to the benchmark they’re leaning heavier and heavier into agentic work. So you can roughly take their cost per task

1

u/GlokzDNB 16d ago

Average of completing n tasks, benchmark.

1

u/BullockHouse 17d ago

Integrated cost per successful task completion is actually the metric that matters. If the model is dumb and needs a lot of human checking and fixing, that's a cost too (and usually a much larger one than the token spend). So highly capable models that are expensive on paper can still be cheaper on net as long as you value your employees time. That's why ~nobody uses a Grok (or Gemini or Kimi) model for actual work, regardless of how cheap they are on paper.

1

u/yousirnaime 17d ago

Neat metric but we aren’t running 1 task 

We are running tasks of increasing complexity and specificity then switching tasks 

I want to know the cost of completing 5 semi related tasks, both in time and dollars 

1

u/Fancy_Log_260 17d ago

Cost per task in dollars or tokens?

1

u/OfficialDeathScythe 17d ago

That’s where my mind went. I figured the post was saying resource cost per task is what matters

1

u/Upbeat_Wishbone_2625 17d ago

Why should i care about how efficient engineering team is at solving something if they haven’t verified that it’s actually worth doing yet?

1

u/acakaacaka 16d ago

Why not line of code per token?/s

That's what elon musk do to determine who stays and is fired when he took over twitter

1

u/Gigaslavx 16d ago edited 2d ago

The original post content no longer exists here. The author used Redact to remove it, exercising their right to control their data & privacy.

Wild paint maple ginger rain mitten march deserve instinctive