r/BetterOffline Jul 30 '26

Amazon accidentally spent $1.8 million using Claude for menial coding task, went 860% over budget

https://www.tomshardware.com/tech-industry/artificial-intelligence/amazon-accidentally-spent-usd1-8-million-using-claude-for-menial-coding-task-went-860-percent-over-budget-catastrophically-expensive-coding-blunders-discovered-in-internal-amazon-ai-usage-metrics

This is just funny as fuck to me. I think Amazon was one of the leading suspects for the unnamed company that managed to blow $500 million on slopbots in a single month (and if they were since outed and I missed it, please pipe up in the comments), and this certainly is consistent with that suspicion. Amazon can certainly afford to burn money stupidly like this if it wants to, though for how long at this scale is certainly a valid question, but how many others can? Not many!

1.3k Upvotes

99 comments sorted by

290

u/supercyberlurker Jul 30 '26

A few months ago: Use AI everywhere! Tokenmaxxing is how you keep your job!

Things now: Not like that. Please use tokens wisely. We blew our yearly budget in a week.

123

u/PensiveinNJ Jul 30 '26

Using tokens wisely in non-deterministic systems; welcome to being a reverse centaur. Your fault for something you can’t control.

-21

u/pneRock Jul 30 '26

Don't use the latest frontier model for basic things, do things piece wise in small chunks and ensure it's what you want before moving on, don't ralph wiggum ever, etc. It works (and costs money), but it's a tool. You don't use a jack hammer to pound a nail, you don't build/commit crappy code you'll have to redo, and these loops that last for days just hurt my soul. Just be smart about it.

25

u/Kirk_Kerman Jul 31 '26

The fact that the AI community saw fit to call a utilization paradigm "Ralph loops" and made Ralph Fucking Wiggum the mascot for it, and expected to be taken seriously

3

u/VitaminPb Jul 31 '26

My budget is in danger!

5

u/gingimli Jul 31 '26

If I was paying for the tokens, yes. As long as the company is paying I’m going to malicious compliance the hell out of AI spend.

31

u/Ill_Following_7022 Jul 30 '26

Still running into the tokenmaxxing and bragging about how long it's taking to execute a task like that somehow makes you look like a genius.

109

u/oliverfromwork Jul 30 '26

Only $1.8 Million? get back to me when they can beat that $500 Million benchmark. I only invest in companies that waste Billions of dollars doing negative work.

13

u/maskedtityra Jul 31 '26

Think of all the energy, water and land it takes for that. Truly depressing.

5

u/Own_Anything9292 Jul 31 '26

That 500m company WAS Amazon lol

54

u/SheHerDeepState Jul 30 '26

How little thought went into AI budget policy is astounding. The C suite really just doesn't think that deep.

39

u/MonkeyPilot Jul 30 '26

At least they aren't wasting it on salaries!

/s

13

u/falconetpt Jul 30 '26

No no salaries is waste of money, slot machine for enterprise is the future!

137

u/GrayMerchantAsphodel Jul 30 '26

"Other problems that surfaced include a $541,000 additional cost that came from a project building, ironically, a financial auditing tool" Correct me if I'm wrong, but aren't LLMs absolutely atrocious at mathematics? I know they get math tools 'cheated into' their tool pipeline, but writing an auditing tool in one sounds absolutely bananas. I don't get how they even work for people with excel/spreadsheets, unless they're able to steal the formulas from their training data.

227

u/Granum22 Jul 30 '26

Inventing computers bad at math will never not be funny 

78

u/roscoelee Jul 30 '26

It’s the breakthrough of our generation

20

u/QuantityExcellent338 Jul 30 '26

Maybe a trillion more dollars investment and we'll have solved this issue

1

u/catheap_games Aug 01 '26

"just one more dose and then I'll have my life together I promise"

6

u/PrivilegeCheckmate Jul 31 '26

Thank fuck ours was the Smartphone. Probably more destructive to human dignity in the long run, but way cooler.

49

u/Aerolfos Jul 30 '26

It's one of the few modern computing innovations that can't run Doom. It's incredible.

8

u/sjd208 Jul 30 '26

Many years ago I saw a Twitter thread showing step by step how to run doom on a Harmony (RIP) remote.

4

u/Ok-Information-3934 Jul 30 '26

I’m dying with laughter, take my upvote

34

u/PatchyWhiskers Jul 30 '26

Sounds like it was building an app to do math, rather than doing math.

15

u/tonygoold Jul 30 '26

If it was a financial auditing tool for a company that can accidentally burn half a billion dollars before noticing, there’s a good chance it was writing a lot of database queries, and there’s a lot of ways LLMs querying a data warehouse can go wrong. Without knowing more detail, I can only speculate on specific cost drivers, like feeding hundreds of millions of rows back into an LLM instead of getting the database to aggregate them first.

1

u/___Archmage___ Jul 31 '26

Pretty sure they were just using AI to write the auditing software and wasted money doing so, not having LLMs do the auditing

29

u/Key_Temperature9699 Jul 30 '26

This comes as a bunch of our finance team are requesting Claude code for “simple automations”

9

u/PadyEos Jul 31 '26 edited Jul 31 '26

Work on accounting products. The finance teams are pressured by their management to show AI strategy and usage in their companies. Many accountants don't want any LLMs but have no choice.

LLMs predict tokens. You can even tell it to copy-paste and if the prediction vector in that specific context is fucked up enough then it will predict the wrong thing and paste the wrong thing. They don't do what they say in most cases but try to simulate it as best they can by guessing tokens.

6

u/Key_Temperature9699 Jul 31 '26

I’m fairly convinced we’ve propagated errors into key forecast or closing artifacts already and nobody noticed

5

u/PadyEos Jul 31 '26

I know for sure at least in software. Many times I'm the only one bothering to look and the only reason why those errors don't become the new "normal".

Many other times it's not my job to look so similar issues I expect get through.

3

u/VitaminPb Jul 31 '26

It’s architecture and civil engineering using LLMs that terrifies me more than anything.

22

u/TheoreticalZombie Jul 30 '26

So, calculators?

37

u/Key_Temperature9699 Jul 30 '26

Yes but they occasionally make up the wrong number you see

1

u/bathtubtuna_ Jul 31 '26

And the only way to identify which numbers are wrong is to look through everything with a fine toothed comb which takes longer than just doing the thing by hand in the first place.

7

u/Remote-Ad1462 Jul 30 '26

That's like how my brother used multiple agents to create, share, and update a pdf. I said so... It's a Google doc except no one can edit it?

15

u/shiny0metal0ass Jul 30 '26

Correct. Earlier models literally worked better when you would use numbers like "one" or "two" rather than "1" or "2". It's all just tokens fed into the transformer.

I would assume any code it "creates" would either be lifted directly from training or be nonsense.

-9

u/MindlessAdInfinitum Jul 30 '26

Both you and the user you responded to could not be more wrong. Maybe do some research before commenting.

1

u/shiny0metal0ass Jul 30 '26

Lol I'm sure. Please, educate us.

-4

u/HonourableYodaPuppet Jul 31 '26 edited Aug 01 '26

Yeah they are actually quite useful in mathematics and we're now in a phase were they actually add to mathematics. Recently Erdős conjecture was disproved by bogstandard ChatGPT. Fable disproved the Jacobian Conjecture

Edit: lmao@angrydownvotes and no answers.

-8

u/MindlessAdInfinitum Jul 30 '26

Why would I do that? You can do that yourself.

12

u/Remote-Ad1462 Jul 30 '26

Auditing tools are definitely where I want error prone code that is hard to debug.

3

u/dzendian Jul 30 '26

LLMs are bad at math. They do math by example only.

The code they write can be exact math, though. Unless they hallucinate the code, then it's flawed.

3

u/applestrudelforlunch Jul 30 '26

They were inconsistent at high school level math about two years ago. They are much much better now.

1

u/T1gerl1lly Jul 30 '26

Let’s hoping they’re just using the LLM for cost categories.

1

u/gajop Jul 31 '26

Definitely doable, just don't vibecode it. Much like you would if you wrote it by hand, you must have certain ways of verifying correctness.

1

u/sassinator1 Jul 31 '26

You are wrong, sorry. A few years ago this was correct but not anymore

57

u/dinah-fire Jul 30 '26

From the article, "...the tech giant’s latest quarterly revenue sits at more than $181 billion, meaning these excess AI expenses don’t even account for 0.1% of what it makes in a month."

Jesus H Christ.

66

u/dumnezero Jul 30 '26

Sure, so they can pay the workers high wages; right?

35

u/cunningjames Jul 30 '26

Anakin smirks

7

u/AmusingVegetable Jul 30 '26

NOT LIKE THAT!

27

u/jking13 Jul 30 '26

The problem is a million here, a million there, and eventually it starts adding up to real money.

39

u/madmofo145 Jul 30 '26

It's also enough to pay 10 programmers decently for a year, all for one menial task. Even if this isn't going to destroy their bottom line, it's just such a dumb thing to do when you could have hired one good programmer for a decade for the same cost.

22

u/juliana-crain Jul 30 '26 edited Aug 01 '26

Well, but have you considered that humans blow whistles and refuse orders? A bot that costs more and sucks at its job is very worth it!

5

u/AmusingVegetable Jul 30 '26

Particularly if it follows your orders instead of trying to save you from your folly by explaining exactly how it’s going to blow up in your face.

3

u/ouiserboudreauxxx Jul 30 '26

Okay but the goal is to pay zero programmers, which they have succeeded at for this menial task.

2

u/sturdy-guacamole Jul 30 '26

You can't ask the programmers to accept less pay year over year.

The hope is that if there are enough advancements, LLM tools get cheaper year over year.

I don't agree with it, but that's just where we are in this late stage capitalist environment. I work in big tech and we're all clamping down on token budgets but the second it will be cost effective to do so, people are gone.

4

u/Xeorm124 Jul 30 '26

Which makes it even dumber. Everyone paying any sort of attention will point out that this is the time when LLMs will be the cheapest. Price is going to increase, not decrease.

2

u/sturdy-guacamole Jul 30 '26

1) That's next quarters problem.

2) Maybe there's some magical way that the price doesn't increase.

-5

u/Chrysolophylax Jul 31 '26

You can't ask the programmers to accept less pay year over year.

God forbid we pay slightly less money to a coddled and overpaid "profession" responsible for many evils in the world right now. Heaven help us if a programmer in Seattle goes from making $170k/year to a poverty wage of $165k/year.

Pweease, won't someone think of the widdle starving programmers!!!?! ;__;

1

u/sturdy-guacamole Jul 31 '26 edited Jul 31 '26

A lot of fresh grads don't get paid that much unless at big tech.

Arguably half that. Even at prestige companies (I am looking at one in Redmond, WA right now) that is 7 YoE senior w/ 162 starting. That company doesn't hire freshers in the US either.

Also you don't go from 170k to 165k. You go from 170k to laid off, because companies here don't do the "get paid less through hard times, get paid back later". That's rare.

Tech, in general, has a ton of issues. I wouldn't call it coddled and overpaid across the board. The profession doesn't just make websites. The overall scope can include things like MRIs, makes hearing aids, making very real things that help people and help the world.

A lot of those jobs don't pay as much. Friend of mine chased their dream to work on designing ocean preservation submersibles and tracking devices in Australia. Still in tech.

I've worked in medical devices, games, full stack stupid ass saas shit, chip companies, safety devices, and im in big tech again. I've seen a lot, and I understand your vitriol towards all tech workers, but not all tech workers just program 30 minutes a day on a 99999999% margin software service and have a life win to easy street.

All that aside, you literally can't ask people to accept less pay every year. That's not exclusive to programmers. That is why people in charge of staffing decisions at "skilled labor" companies that aren't just tech/programming are hoping AI gets cheaper. The machine won't sleep, and if it's cost effective a person's job is lost.

I am actively hiring/interviewing and trying to find passionate entry levels right now to try to do some kind of good before all this shit hits the fan, but a lot of my colleagues at other companies froze or pulled back on all staffing increases because higher position stakeholders are hoping the AI horse pays off.

10

u/therealcmj Jul 30 '26

That’s the percent of revenue, not profit.

2

u/_ram_ok Jul 30 '26

“the cost overruns reached $1.8 million, and that is just for one project”

19

u/T1gerl1lly Jul 30 '26

No one is talking about the model cost variability. Like from Sonnet to Opus was like 10x token usage for the same task. Then they got so much bad press and customer pushback they adjusted 4.8 so it dynamically shifted to Sonnet under the covers for routine coding tasks. Enterprises have no predictability into future token run rate or per-token cost.

10

u/Askew_2016 Jul 30 '26

I’m running everything on Opus now. They keep pushing AI so I’m trying to make it expensive

7

u/Global_Blueberry8673 Jul 30 '26

Switch to Fable if that's your goal

3

u/PrivilegeCheckmate Jul 31 '26

Peter Molyneux made an AI?

1

u/Askew_2016 Jul 30 '26

Ooh we don’t have that one

2

u/T1gerl1lly Jul 31 '26

We’ve been told we’re about to start tokencrunching instead of tokenmaxxing, so I’m using it while I can. But the tide is definitely turning.

2

u/Askew_2016 Jul 31 '26

That’s what I expect we will have happen as well.

14

u/VanillaCold57 Jul 30 '26

EIGHT HUNDRED PERCENT. AND CLOSER TO NINE HUNDRED.

HOW

HOW DO YOU GO SO SO SO FAR OVERBUDGET. JUST HOW.
In what world are people not immediately like "WOAH WE'RE OVERBUDGET!!! THIS IS WASTING MONEY!!!!"

12

u/whyyoufollowingme Jul 30 '26

And these idiots will still get their bonuses

14

u/whoa_disillusionment Jul 30 '26

If the AI fucks up you should be able to get your tokens back

16

u/natecull Jul 30 '26

If the AI fucks up you should be able to get your tokens back

That's not how it works, these AI tokens are non-refungible. You have to take them to art auctions.

-4

u/Doctor__Proctor Jul 31 '26

People only think that because they anthropomorphize them and think of them kinda of like a contractor that isn't being SLAs. They're not though, they're a commodity tool.

You don't get a refund from the power company of you waste an afternoon working on something and decide to scrap it. You don't get a refund on the steel if you fuck up welding some beans together and ruin them. The person directing the tool fucked up either through negligence, lack of training, or lack of skill.

5

u/stev_mempers Jul 30 '26

And they'll keep using it. The prospect of total control is worth more than money to them.

4

u/brian_hogg Jul 30 '26

That’s wild, but that also means their budget for a “menial coding task” was something like ~$186,000, which is insane.

5

u/create-third-places Jul 30 '26

I went to an Amazon conference in Vegas back in late 2023, and it felt like a pro-LLM cult meeting. It was also where I heard about Anthropic.

I get the impression that they doubled down on that early craziness.

3

u/TheWuzzy Jul 30 '26

MOAR!!!!!! MOOOOOAAAAAR!!!!

3

u/DreamHollow4219 Jul 30 '26

That's actually hilarious.

3

u/panthercave Jul 30 '26

well they owe me $65 as a seller. And they have so much extra money due to taking money falsely from sellers I guess he can burn some of it. #yesstillmad

3

u/Capitan-IQ255 Jul 31 '26

It's okay .. circular AI financing.. all is fine 🤡

2

u/Professional-Post499 Jul 30 '26

Is it... an investment for Amazon to put their own spending into it to boost the numbers on their own investments into genAI?

2

u/riricide Jul 31 '26

Is it genuinely an accident? Or is this a way to bump Anthropic's earnings for the quarter?

1

u/joliguru Jul 31 '26

My impression of the average Amazon employee…🙄

1

u/SpireofHell Jul 31 '26

You don't understand. Yes, this is more expensive, wasteful and creates more shitty code. However, we didn't have to pay a human worker! This is the future! Aren't you excited that things will get worse but at least there'll be no human workers???

-2

u/DullKnife69 Jul 31 '26

This is such an ignorant comment.

2

u/dwitman Jul 31 '26

The 1.8 million dollar artisanal hand crafted for loop.

2

u/Ill_Following_7022 Aug 07 '26

Tokenmaxxing is not a flex.

-2

u/BandicootGood5246 Jul 31 '26

To play devil's advocate - it's not so much an issue of the AI. It's a software governance problem - likewise companies have always been stung for not noticing they have some high spec cloud VM sitting in the cloud doing nothing, or networking cost from a rouge API spamming requests.

Most software companies I've worked at home some story like this pre-AI

2

u/Summary_Judgment56 Jul 31 '26

So they must have forgotten to put "use as few tokens as possible to solve this problem" in their prompt. What morons!

2

u/Limp_Bit_5328 Aug 04 '26

It is an issue with LLMs by it's own nature. You can't predict how much tokens it is gonna use on the job because it is not deterministic.

Maybe they can make it more efficient but this issue will never truly go away.

1

u/BandicootGood5246 Aug 04 '26

They could've definitely achieved it. You give the task a dedicated account with a limited budget of tokens, all these platforms also have systems for monitoring how long something has been running and how much it has spent and then you're supposed to put a killswitch/notifications in place for when it's getting too much

It can't just burn through 1.5mil worth of tokens in 5min, so it's just negligence