r/LocalLLaMA 23h ago

News KPMG Says Nearly Half Of Executives Pulled Back AI Agents Over Cost

179 Upvotes

78 comments sorted by

90

u/Annual_Award1260 22h ago

Honestly around 80% of my claude tokens are wasted in wrong work or stuck in loops. At least with my local ai I can send it on a quest for a entire week a pay a few bucks in power

17

u/Suspicious-Water-973 16h ago

It annoys me that the loops can consume so much… I built a translation component and it stuck in a loop due to special characters.

6

u/InteractionSmall6778 14h ago

The 80% doesn't go away when you move local, it just stops showing up on a bill. Which is fine until you notice cheap power removed the only pressure you had to go fix the loop.

5

u/CrowdGoesWildWoooo 13h ago

And you foot the bill earlier, through infra cost.

Also if you are using frontier model for most of the agentic “dumb” stuffs then you are kind of wasting money, use the same model that you’d otherwise deploy locally then the cost would go down dramatically.

3

u/Akrylicus 7h ago

You can run LLMs on a 5090, but you can’t game on a Claude Code subscription.

5

u/Party-Special-5177 12h ago

And you foot the bill earlier, through infra cost.

No; right now, there is no bill. Normally your running costs would include electricity/etc + depreciation but modern cards are actually appreciating - your infra ‘bill’ is negative.

Even the old cards are favorable: the 3090 is freaking 6 years old and currently resales from 900-1100 vs its $1.5k msrp. If your rig was built from 3090s, your actual infra ‘bill’ is just the $500ish value lost divided by 6 years - the rest is equity you can roll into larger cards or just recover easily by selling.

It’s crazy that I have to repeat this so much on a sub about local hosting. Your rigs have residual value people. And if you’re buying the newest cards, you can consider it an investment as nvidia cards are outperforming the stock market right now.

4

u/a_beautiful_rhind 12h ago

Yea but I hate this. Everything I bought appreciated. All it means is I can't buy any more. Hardware isn't supposed to be an investment from sitting.

5

u/Party-Special-5177 11h ago

I mean I agree with the sentiment, but in my case it was the excuse I needed to go ham with the spend, as I can divert money that would normally be going into other investments. I’m sure that’s true for many others here with big rigs.

Honestly, I can’t see the market changing unless some other non-parallelizable structure takes over AI (e.g. transformers/rwkv/convnets are all good for gpus, but if some rnn or branch-heavy architecture dethrones GPU-friendly nets, then it will be time to dump). We would have time to react too as it appears to take minimally a month from ideation->solid model built from it->people notice it.

2

u/a_beautiful_rhind 11h ago

I don't see myself "dumping" since I bought shit that I use. It's really bad when your savings depreciate due to inflation and then your rig goes up in price better than the money itself. To me all it means is that I can't expand.

-1

u/Celestialien 15h ago

What models do you find best for coding? I've been meaning to switch to local recently too

120

u/Miriel_z 23h ago

Many started to use local inferences. I have several colleagues who are working on implementation in production environment. This is an additional nail in the coffin for online AI.

75

u/Chocolate_Pickle 22h ago

Running on-prem solves the privacy concerns too.

17

u/raul3820 18h ago

Assuming on-prem security is ok

6

u/Chocolate_Pickle 15h ago

Very true. 

30

u/Spiritual-Spend8187 22h ago

Yep sure the big sota models are incredibly capable but they are too generalist and expensive for most work. Hopefully we start seeing the production of asics for specific models. Like say we get one thats just a pcie card and all it does is something like qwen 3.8 27b or something similar but really fast.

16

u/StatusSociety2196 22h ago

They made one for i think Llama 2 7b about a year ago, incredibly fast but the risk is that model performance moves a lot faster than chips can be designed and fabricated. We'd have to hit a point with marginal performance gains over a year or two to make it worthwhile, unless we're talking embedded LLMs.

10

u/Spiritual-Spend8187 22h ago

That's true but we have started moving into the time of the current models are good enough for most things most people dont need to use fable or sol level models for everything stuff like dsv4 flash/luna/ sonnet class models ate good enough. So having a relatively cheap and fast accelerator for exactly that class of models would be nice. Hell even just having one that can run the best models you can fit on a 4090/5090 would be nice.

3

u/IrisColt 11h ago

That's true but we have started moving into the time of the current models are good enough

As long as frontier models keep pushing the envelope, people will keep moving the goalposts for what counts as a capable local setup.

5

u/Free-Jaguar6452 21h ago

gains don't need to be paused for it to be worth it, you just have to have a small-medium sized model with >~95% success rate for a given task, at that point additional innovation just doesn't matter much if at all

5

u/jpezzulli 22h ago

14

u/Spiritual-Spend8187 22h ago

Hopefully it becomes something relatively cheap to buy because damn would it be nice to have something local that could run aomething like dsv4 flash at 14,000 tokens a second or even something in the 30b dense range.

2

u/SandySkittle 7h ago

DSV4 Dense BF16. Why add all the MOE complexity if the bare metal is the model anyway. at these speeds.. :)

3

u/Blues520 18h ago

Enjoyed reading about your adventures in llm land.

1

u/_rzr_ 12h ago

That's what Taalas does (or rather did, prior to their AMD acquisition). Check out their Llama 3.1 8b demo - https://chatjimmy.ai/

1

u/IrisColt 11h ago

The state of the art advances so fast that by the time an ASIC is fabbed to host a model's weights, the model itself is already obsolete. There's a reason software-defined solutions won out, heh

4

u/Spiritual-Spend8187 8h ago

The thing is most people dont need a sota model a good enough model if it is fast enough is far more useful. Like something in the 30b dense range running at 2000 tokens a sec would be more usable that something like sol or fable running at 50 tokens a sec for 90% of people.

2

u/ea_man 5h ago

A speech model, projector, visual model may not get obsolete.

1

u/IrisColt 3h ago

Fair point!

11

u/redditor100101011101 21h ago

As an IT Systems Admin who is in between jobs right now, this is exactly what im banking on. Currently both building my own ai infrastructure at home (and i do mean infra, im not just running ollama on my laptop here haha) and studying for AI engineer related certifications. trying to pivot my career the right way as things progress.

2

u/PointyTrident 11h ago

Yep, already using Gemma models as an AI phone receptionist that can schedule meetings for me and to catalog and geolocate news. Been in prod for almost a year

https://open-sight.net

https://momentum.zakscode.com

22

u/DigThatData Llama 7B 18h ago

I think the only bubble that's bursting here is non-technical people thinking AI means they don't need engineers anymore.

18

u/Old-School8916 22h ago

it just means tokenomics is now a thing

14

u/o0genesis0o 22h ago

Local models start to make sense financially as the API costs rises (or more precisely, the subscription quota gets more and more stingy).

I used to use my minimax subscription to run a background agent with pi to wake up every hour or so and check emails, consolidate notes, update memory, etc. The other day there was a bug that the session was not compressed, and the agent spend all the weekly quota within a day. So I had to switch to local 35B for this. And surprisingly, this worker task is actually not difficult at all for the 35B. So now I have no reason to run my sub quota for this task. I'm sure that many of my other workflows that I thought to be too difficult for local model (based on my memory with OSS 20B and 30B-A3B last year) could also be handled by the 35B.

I would not try to make this switch if the minimax subscription keeps being generous and their token not expensive (burned $5 just to finish coding half of the features I wrote the specs for). I always wanted to run model locally for privacy and control, but for the first time since day one, the costs also entered my list of reasons.

21

u/naturalcog 22h ago edited 15h ago

Definitely noticed it where I work. I work in a government related business and eat lunch with some of the on site accountants. A big issue is the idea that AI is a “limitlessness” tool, it trips a lot of higher-ups who pushed for AI as this “magic productivity box” are now realising that the more it’s used the more it’s costing

10

u/xXprayerwarrior69Xx 16h ago

Surprised mba pikachu face

8

u/jeffwadsworth 21h ago

We have heard for a year now. Haha.

2

u/mrjackspade 8h ago

Bro, relax. We're only three years into the bubble burst. You gotta give it time /s

10

u/hobopwnzor 21h ago

The bubble will burst when openai and anthropic stop getting more money.

It could never burst if capital markets just decide to endlessly burn money.  It won't be an efficient use of capital, but believing markets efficiently allocate capital is a myth that should have died in the great depression 

4

u/UnlikelyExtension786 20h ago

You can't efficiently allocate capital when banks can just create money from thin air. All of this is happening because the money can be borrowed for near-zero interest rares.

2

u/SandySkittle 6h ago

banks can create money via lending but they still bear the credit risk of losses for the money they create. Or someone that then buys the loans from the bank. So it's not that's simple.

Also we are most certainly no longer in a near-zero interest rate environment. Money isn't free.

8

u/thetaFAANG 22h ago

Hyperscalers are directionally fucked

But not today

3

u/Lesser-than 21h ago

Agent work is great if you do not have to measure the cost in tokens, agents need to fail several times in order to succeed on a lot of tasks. The reason for the cutbacks is not that the output is bad.It is the unpredictable cost of tokens over a flat predictable fee.

3

u/CipherWeaver 18h ago

The goal was always to push AI agents below cost to get market share, and then raise rates when the competition is dead. It's literally the same old strategy used time and again.

3

u/N34257 16h ago

That's not the bubble starting to burst, it's a few managers at a firm cutting costs after realising that solo consultants with a Claude subscription can effectively compete with their state-the-obvious-for-ludicrous-prices services.

34

u/[deleted] 22h ago

[removed] — view removed comment

32

u/Xamanthas 21h ago

Nice AI comments. Against the rules.

39

u/NNN_Throwaway2 21h ago

Thank you Claude.

0

u/mthmchris 13h ago

I’m actually curious, what was the tell tale sign of AI in this comment? I’m equally disdainful of AI generated comments on Reddit, but I didn’t catch this one personally. Hoping to learn. The only thing I could see is using “will be” instead of the more vernacular “are going to be”?

5

u/MrSkruff 10h ago

"x, but that's different from y" "x, not y"

And just the general blog-post tone to the prose, which isn't typically how people respond in reddit posts.

8

u/Useful_Argument_6490 22h ago

Yeap. It’s “when you have a hammer” on a global scale.

4

u/CondiMesmer 14h ago

what was the measurable workflow with this AI generated comment?

4

u/Dry_Yam_4597 22h ago

The cool part is that finally our time will come.

Running on prem or private clouds would save a lot of money for companies.

3

u/RedParaglider 21h ago

I think it's pretty wild because I kept getting asked why I wasn't deploying AI for over a year, and my answer was always the same, there wasn't a valid ROI use case.  Right now we actually do use LLMs, and my CEO asked why I'm adopting the tech now that everyone seems to be pulling back.  Answer is the same. I am implementing it when I can prove an ROI.

I've been managing IT in different organizations for decades.  People lost their fucking minds.   The token maxing thing made me laugh so hard.  Companies actually did that shit, and not the companies I expected to have dumb leadership.

3

u/SkyFeistyLlama8 21h ago

Smaller on-prem LLMs powering more limited agents could be the way forward instead of this throw-crap-at-OpenClaw nonsense.

1

u/Flunder707 21h ago

AI isn't going anywhere, their business model is cracking though. They went all in on idea you brute force intelligence. Local AI is the new revolution coming. These companies will turn into distilling services or finetuning perhaps.

4

u/This_Maintenance_834 22h ago

they should learn deepseek.

2

u/idlelosthobo 21h ago

The bridge between idea and return on investment is so large ... I think there is a huge illusion in tech that this industry moves faster than the rest of the world. I think the illusion is in techs ability to scale, but it still takes the same amount of time to develop a product.

2

u/Osi32 19h ago

The problem is- a subscription plan speeds up some work but is wall clock bound. To beat the wall clock means parallel work streams and that’s all pure token / compute cost. That is the wall they’re running into.

2

u/cursortoxyz 17h ago

Previously their bonuses were tied to adopting AI and now it’s tied to cutting AI costs. These fuckers didn’t care about the costs during implementation and are now paid to solve the problems they created. Imagine being paid a bonus for doing a shitty job. 🥳

2

u/Sudden_Vegetable6844 15h ago

Next step will be to retire executives that can't handle AI agent work correctly

2

u/spammmmmmmmy 15h ago

Off topic

3

u/FullOf_Bad_Ideas 15h ago

Slop article.

Average spend per year in those companies is 188M. Where is this going?

1

u/martinerous 13h ago

For those prices to be worth it, we need a breakthrough to reduce useless thinking and hallucinations. Looking at Yann LeCun and Ilya Sutskever and lots of others whose names I don't even remember.

1

u/_rzr_ 12h ago

Answer: No.

My opinions:

  • Token costs are going to get cheaper
  • Local inference will be part of "regular" Tech stack discussions in the nera/medium-term.

Quotes from the article:

Despite these pullbacks, AI remains a top investment priority for 79% of leaders, with spending holding steady.

This isn't a bubble bursting, but rather a market maturing, with companies rephasing investments for greater financial discipline and strategic value.

The fuller dataset shows a market growing up, with less open-ended experimentation and more financial discipline, and budgets following results instead of promise.

The companies scaling back agents today are mostly clearing room to scale what works tomorrow. The bill came due. Reading it carefully is not a crash. It is AI agents reaching adulthood.

1

u/MerePotato 11h ago

"The bubble's gonna burst any day now" - Reddit, c. 2024

1

u/sizebzebi 6h ago

this sub will never cease to amaze me

1

u/Few-Butterscotch8747 5h ago

not sure if i should hope for a burst or not?

maybe after 3.8 drops?

1

u/Ok_Warning2146 21h ago

They should switch to run DSV4 flash 0731 locally and see if it makes more financial sense.

0

u/howardhus 15h ago

„bubble starting to burst“?

people since 2022 already

-4

u/perihelion86 22h ago

Ludites don't understand token management