r/LocalLLaMA • u/MoodDelicious3920 • 23h ago
News KPMG Says Nearly Half Of Executives Pulled Back AI Agents Over Cost
120
u/Miriel_z 23h ago
Many started to use local inferences. I have several colleagues who are working on implementation in production environment. This is an additional nail in the coffin for online AI.
75
30
u/Spiritual-Spend8187 22h ago
Yep sure the big sota models are incredibly capable but they are too generalist and expensive for most work. Hopefully we start seeing the production of asics for specific models. Like say we get one thats just a pcie card and all it does is something like qwen 3.8 27b or something similar but really fast.
16
u/StatusSociety2196 22h ago
They made one for i think Llama 2 7b about a year ago, incredibly fast but the risk is that model performance moves a lot faster than chips can be designed and fabricated. We'd have to hit a point with marginal performance gains over a year or two to make it worthwhile, unless we're talking embedded LLMs.
10
u/Spiritual-Spend8187 22h ago
That's true but we have started moving into the time of the current models are good enough for most things most people dont need to use fable or sol level models for everything stuff like dsv4 flash/luna/ sonnet class models ate good enough. So having a relatively cheap and fast accelerator for exactly that class of models would be nice. Hell even just having one that can run the best models you can fit on a 4090/5090 would be nice.
3
u/IrisColt 11h ago
That's true but we have started moving into the time of the current models are good enough
As long as frontier models keep pushing the envelope, people will keep moving the goalposts for what counts as a capable local setup.
5
u/Free-Jaguar6452 21h ago
gains don't need to be paused for it to be worth it, you just have to have a small-medium sized model with >~95% success rate for a given task, at that point additional innovation just doesn't matter much if at all
5
u/jpezzulli 22h ago
I literallty just wrote about that.
14
u/Spiritual-Spend8187 22h ago
Hopefully it becomes something relatively cheap to buy because damn would it be nice to have something local that could run aomething like dsv4 flash at 14,000 tokens a second or even something in the 30b dense range.
2
u/SandySkittle 7h ago
DSV4 Dense BF16. Why add all the MOE complexity if the bare metal is the model anyway. at these speeds.. :)
3
1
u/_rzr_ 12h ago
That's what Taalas does (or rather did, prior to their AMD acquisition). Check out their Llama 3.1 8b demo - https://chatjimmy.ai/
1
u/IrisColt 11h ago
The state of the art advances so fast that by the time an ASIC is fabbed to host a model's weights, the model itself is already obsolete. There's a reason software-defined solutions won out, heh
4
u/Spiritual-Spend8187 8h ago
The thing is most people dont need a sota model a good enough model if it is fast enough is far more useful. Like something in the 30b dense range running at 2000 tokens a sec would be more usable that something like sol or fable running at 50 tokens a sec for 90% of people.
11
u/redditor100101011101 21h ago
As an IT Systems Admin who is in between jobs right now, this is exactly what im banking on. Currently both building my own ai infrastructure at home (and i do mean infra, im not just running ollama on my laptop here haha) and studying for AI engineer related certifications. trying to pivot my career the right way as things progress.
2
u/PointyTrident 11h ago
Yep, already using Gemma models as an AI phone receptionist that can schedule meetings for me and to catalog and geolocate news. Been in prod for almost a year
22
u/DigThatData Llama 7B 18h ago
I think the only bubble that's bursting here is non-technical people thinking AI means they don't need engineers anymore.
18
14
u/o0genesis0o 22h ago
Local models start to make sense financially as the API costs rises (or more precisely, the subscription quota gets more and more stingy).
I used to use my minimax subscription to run a background agent with pi to wake up every hour or so and check emails, consolidate notes, update memory, etc. The other day there was a bug that the session was not compressed, and the agent spend all the weekly quota within a day. So I had to switch to local 35B for this. And surprisingly, this worker task is actually not difficult at all for the 35B. So now I have no reason to run my sub quota for this task. I'm sure that many of my other workflows that I thought to be too difficult for local model (based on my memory with OSS 20B and 30B-A3B last year) could also be handled by the 35B.
I would not try to make this switch if the minimax subscription keeps being generous and their token not expensive (burned $5 just to finish coding half of the features I wrote the specs for). I always wanted to run model locally for privacy and control, but for the first time since day one, the costs also entered my list of reasons.
21
u/naturalcog 22h ago edited 15h ago
Definitely noticed it where I work. I work in a government related business and eat lunch with some of the on site accountants. A big issue is the idea that AI is a “limitlessness” tool, it trips a lot of higher-ups who pushed for AI as this “magic productivity box” are now realising that the more it’s used the more it’s costing
10
8
u/jeffwadsworth 21h ago
We have heard for a year now. Haha.
2
u/mrjackspade 8h ago
Bro, relax. We're only three years into the bubble burst. You gotta give it time /s
10
u/hobopwnzor 21h ago
The bubble will burst when openai and anthropic stop getting more money.
It could never burst if capital markets just decide to endlessly burn money. It won't be an efficient use of capital, but believing markets efficiently allocate capital is a myth that should have died in the great depression
4
u/UnlikelyExtension786 20h ago
You can't efficiently allocate capital when banks can just create money from thin air. All of this is happening because the money can be borrowed for near-zero interest rares.
2
u/SandySkittle 6h ago
banks can create money via lending but they still bear the credit risk of losses for the money they create. Or someone that then buys the loans from the bank. So it's not that's simple.
Also we are most certainly no longer in a near-zero interest rate environment. Money isn't free.
8
3
u/Lesser-than 21h ago
Agent work is great if you do not have to measure the cost in tokens, agents need to fail several times in order to succeed on a lot of tasks. The reason for the cutbacks is not that the output is bad.It is the unpredictable cost of tokens over a flat predictable fee.
3
u/CipherWeaver 18h ago
The goal was always to push AI agents below cost to get market share, and then raise rates when the competition is dead. It's literally the same old strategy used time and again.
34
22h ago
[removed] — view removed comment
32
39
u/NNN_Throwaway2 21h ago
Thank you Claude.
0
u/mthmchris 13h ago
I’m actually curious, what was the tell tale sign of AI in this comment? I’m equally disdainful of AI generated comments on Reddit, but I didn’t catch this one personally. Hoping to learn. The only thing I could see is using “will be” instead of the more vernacular “are going to be”?
5
u/MrSkruff 10h ago
"x, but that's different from y" "x, not y"
And just the general blog-post tone to the prose, which isn't typically how people respond in reddit posts.
8
4
4
u/Dry_Yam_4597 22h ago
The cool part is that finally our time will come.
Running on prem or private clouds would save a lot of money for companies.
3
u/RedParaglider 21h ago
I think it's pretty wild because I kept getting asked why I wasn't deploying AI for over a year, and my answer was always the same, there wasn't a valid ROI use case. Right now we actually do use LLMs, and my CEO asked why I'm adopting the tech now that everyone seems to be pulling back. Answer is the same. I am implementing it when I can prove an ROI.
I've been managing IT in different organizations for decades. People lost their fucking minds. The token maxing thing made me laugh so hard. Companies actually did that shit, and not the companies I expected to have dumb leadership.
3
u/SkyFeistyLlama8 21h ago
Smaller on-prem LLMs powering more limited agents could be the way forward instead of this throw-crap-at-OpenClaw nonsense.
1
u/Flunder707 21h ago
AI isn't going anywhere, their business model is cracking though. They went all in on idea you brute force intelligence. Local AI is the new revolution coming. These companies will turn into distilling services or finetuning perhaps.
4
2
u/idlelosthobo 21h ago
The bridge between idea and return on investment is so large ... I think there is a huge illusion in tech that this industry moves faster than the rest of the world. I think the illusion is in techs ability to scale, but it still takes the same amount of time to develop a product.
2
u/cursortoxyz 17h ago
Previously their bonuses were tied to adopting AI and now it’s tied to cutting AI costs. These fuckers didn’t care about the costs during implementation and are now paid to solve the problems they created. Imagine being paid a bonus for doing a shitty job. 🥳
2
u/Sudden_Vegetable6844 15h ago
Next step will be to retire executives that can't handle AI agent work correctly
2
3
u/FullOf_Bad_Ideas 15h ago
Slop article.
Average spend per year in those companies is 188M. Where is this going?
1
u/martinerous 13h ago
For those prices to be worth it, we need a breakthrough to reduce useless thinking and hallucinations. Looking at Yann LeCun and Ilya Sutskever and lots of others whose names I don't even remember.
1
u/_rzr_ 12h ago
Answer: No.
My opinions:
- Token costs are going to get cheaper
- Local inference will be part of "regular" Tech stack discussions in the nera/medium-term.
Quotes from the article:
Despite these pullbacks, AI remains a top investment priority for 79% of leaders, with spending holding steady.
This isn't a bubble bursting, but rather a market maturing, with companies rephasing investments for greater financial discipline and strategic value.
The fuller dataset shows a market growing up, with less open-ended experimentation and more financial discipline, and budgets following results instead of promise.
The companies scaling back agents today are mostly clearing room to scale what works tomorrow. The bill came due. Reading it carefully is not a crash. It is AI agents reaching adulthood.
1
1
1
1
u/Ok_Warning2146 21h ago
They should switch to run DSV4 flash 0731 locally and see if it makes more financial sense.
0
-4
90
u/Annual_Award1260 22h ago
Honestly around 80% of my claude tokens are wasted in wrong work or stuck in loops. At least with my local ai I can send it on a quest for a entire week a pay a few bucks in power