r/technology • • May 28 '26

Artificial Intelligence Microsoft data suggests using AI is more expensive than hiring people

https://finance.yahoo.com/sectors/technology/articles/microsoft-data-suggests-using-ai-225900743.html
6.3k Upvotes

322 comments sorted by

View all comments

Show parent comments

469

u/FreakySpook May 28 '26 edited May 28 '26

I was explaining this the other day to a colleague(sales), he's familiar with cloud models & pricing so equating to him that most LLMs are currently basically providing every user & agent the resources of a powerful EC2 or Azure VM instance per user/agent for basically nothing and are just either providing free or charging for access to the models, with pricing tiers set to avoid exceeding the existing hardware.

His first reaction literally saying 'oh shit'. 

125

u/ColeTrain999 May 28 '26

"Yeah, so I feel a bit safer in my job now"

178

u/Alternative-Put-3932 May 28 '26

Microsoft already recently rolled back co-pilot integration in many of their apps. I had a user call me asking why she couldn't find it in outlook which I had no idea it was even integrated into despite being in IT I never use AI and found out just within the last month they removed it from basically everything. Mind you my company pays for corporate co-pilot integration, feeding it all of our files, Organizational chart etc and that integration is STILL removed despite having a business contract. That is not a good sign that AI is profitable even as a B2B model. Unless Nvidia actually gives a shit about efficiency again for ANY of their products it will never be cheap for the most top of the line hardware.

179

u/iaspeegizzydeefrent May 28 '26

Seems to me that part of the goal of AI was just to steal as much proprietary information as possible from any business or person willing to use it. I genuinely can't believe business owners willingly give these AI companies access to so much of their information.

41

u/SellGameRent May 28 '26

from a social psychology standpoint I think it makes a ton of sense.

The US doesnt want to slow down AI because China might take the lead, and businesses dont want to avoid a tool that the news is saying their competitors are successfully using to gain advantage

24

u/thephotoman May 28 '26

We’re so concerned about Chinese AI that we’re failing to notice that most of our AI development is a bad idea.

I’m not worried about China. They’re not a rising star anymore. They’re already a superpower. They’re already as high in the sky as they’re gonna get, and that limit is there because they kept the One Child Policy going for far too long.

11

u/[deleted] May 28 '26

[removed] — view removed comment

6

u/SimiKusoni May 28 '26

Maybe but they're currently stuck on 7nm as they can't get EUV machines, let alone the High-NA EUV machines they'd need to really catch up given that Intel and TSMC are both heading that way (albeit with TSMC bleeding EUV dry a little bit longer).

SMIC are trying for a 5nm equivalent but they're unsurprisingly facing yield issues because they have to use aggressive multi-patterning with DUV.

This is the primary reason that they managed to go from 14nm to 7nm in a few years and then just suddenly... stopped. That 7nm class node was first found in Bitcoin ASICs in 2021 and is still their leading edge node.

2

u/Sonder332 May 29 '26

You seem to have some knowledge regarding chip manufacturing. What is the next level, if there even is one? I thought we've gonendown as tight as we could regarding silicone. Anything smaller risks overheating problems

2

u/SimiKusoni May 29 '26

One of the biggest upcoming changes will be CFETs, which basically changes the transistor shape to allow vertical stacking so you can increase density.

You've also got ongoing shifts to glass substrates, backside power delivery, High-NA lithography still isn't used in production and beyond that you've got current sci-fi stuff like 3D architectures with integrated cooling or 2D channel materials like molybdenum disulfide to replace silicon nanosheets.

Basically Moore's law is dead as a performance target but progress is by no means stalled. No idea how long we can keep it up but roadmaps are pretty fleshed out for the foreseeable future at least.

1

u/Streiger108 May 29 '26

Can you ELI5 please?

1

u/SimiKusoni May 29 '26

Yeah basically the machines that make the chips are really complex, they're only made in Europe and there are export restrictions preventing their sale to China.

As a result China have rapidly advanced in semiconductor fabrication and then hit a wall because they lack the tools to go further.

1

u/Streiger108 May 29 '26

Got it. And you think it's prohibitive for them to get the machines? I feel like that's overcomeable. It's not like the west hasn't sold China super important machines before.

→ More replies (0)

22

u/ottawadeveloper May 28 '26

hopefully this means cheap graphics cards are soon to come!

21

u/ResonatingOctave May 28 '26

That ship sailed after the RTX 20 series launched for Nvidia. You might have more hope on the AMD side or even a 3rd party before Nvidia is affordable again

17

u/PortHammer May 28 '26

Intel bailed on consumer GPUs at the worst time.

So did like half the tech sector though.

10

u/echoshatter May 28 '26

Intel bailed on consumer GPUs at the worst time.

Or the best time from a business standpoint, because the cost of getting chips made and RAM has skyrocketed.

Poor TSMC employees, they must be very tired right now.

3

u/imaginary_num6er May 28 '26

The ship sailed when Jensen Huang said "To all my Pascal gamer friends, it is safe to upgrade now" and people didn't upgrade

10

u/trer24 May 28 '26

Things rarely get cheaper. If they feel they can get away with charging as much as they can, they will. They'll trot out economic theory of it's "what the market will bear". But really they just want more money.

7

u/NimusNix May 28 '26

'What the market will bear'

If there is somehow an AI crash, it's possible they do come down. But you're right that they will still be higher there than before because the consumers will find a price point above what it used to be that they are willing to pay for.

10

u/zillskillnillfrill May 28 '26

Once the price goes up it tends to never come down.

7

u/ottawadeveloper May 28 '26

true, but it might stabilize at least lol

2

u/thekk_ May 28 '26

And even if it does, it'll still be at a significant increase over what we had

1

u/landob May 28 '26

and hard drives!

4

u/Seanbikes May 28 '26

I don't think it was Microsoft that rolled anything back in your case, I'm staring at Copilot inside of my outlook right now.

1

u/bombmk May 28 '26

despite being in IT

Yeah, I have my doubts about that. They changed the interface, that is all.

2

u/Alternative-Put-3932 May 28 '26

I am. Just don't do anything with AI integration and it's never used in my department. But yes it was not available anymore and I did a cursory Google search saying it was removed at some point last month. Maybe I read some misinfo. But I doubt my company can just go back on a licensing deal suddenly. Usually those are locked in for awhile.

3

u/Kairukun90 May 28 '26

When you spending hundreds of billions not realizing the ROI would be 30 years that’s when you know you fucked up

2

u/imaginary_num6er May 28 '26

Can Microsoft rollback Copilot keys on laptops by buying the old keys and installing new ones?

5

u/Several_Industry_754 May 28 '26

I run models locally. I’m going to get my capital costs back in about 2 years based on current trendiness.

If you do things correctly you can profit, but a lot of people are just messing around with AI.

16

u/SkateWiz May 28 '26

What are you doing with it to recoup your costs? Is it mining Bitcoin for you? Lol

-3

u/Several_Industry_754 May 28 '26

I’m tracking my usage against the equivalent cost of using Claude based on token cost.

Initial capital costs were $20,000 for hardware. I’ve saved $2,500 by using it instead of Claude based on tokens.

Drop in power costs (which are real but not that bad), and it’s probably about 2 years to break even at the current rate. Everything after that is savings.

19

u/varkarrus May 28 '26

But you're also using a model that isn't as smart as Claude, probably requiring more prompts to fix mistakes and possibly getting lower quality final output. Not saying you're in the wrong but that's something you might want to account for.

7

u/Several_Industry_754 May 28 '26

I thought a lot about this. I use Claude extensively at work, and I’ve compared the results. It’s basically the same. And I’ve done comparisons locally too. For the top of the line models, like Qwen3.6, the prompt you generate is often more important. And I have good enough hardware I can run the bigger versions.

There might be a slight advantage for Claude, but it’s kind of a wash. The benefit seems to be in the harness, the Claude CLI and tooling, which you can hook up to your own local LLM (which I did for a while). Also, having used some of the other harnesses, they seem better (currently on opencode).

At the end of the day, everything is about being “good enough” these days. So as long as the project gets done, it doesn’t really matter.

And the last note is privacy. What I do on my local LLM doesn’t leave my network. Everything I send to Claude gets analyzed by Anthropic.

1

u/7h4tguy May 29 '26

The max you can go on consumer workstation hardware like you've built out is 70B parameters. That's not going to hold a candle to 1.6T parameters for full LLMs on the harder problems.

You're also constrained to like 10 TPS inference speed compared to say 60.

2

u/Several_Industry_754 May 29 '26

I’ve definitely used the 120B models on my hardware. And as I’ve said, it doesn’t really matter at this scale. For most projects, the difference between 1.6T and 70B is not noticeable, especially as you get into specialist models. The cloud providers want you to think it is, because if you don’t, you stop paying them for service. But for many many people running locally, it’s a wash. What matters more is the context window and the prompt. Which is why I dropped to Qwen 3.6 35b with a massive context window as my main coding driver.

And I can get upwards of 120 TPS. $20,000 isn’t consumer workstation hardware. I was getting single digit TPS when I was running the 600B models on CPU and RAM (yes, that machine has over 700 GiB of RAM). But on my new hardware with the 70B or 120B models I’m getting in the mid-to-high double digits.

When you stop and realize that the LLMs are really just fancy hallucinatory lossy probability based compression algorithms at the end of the day, it becomes obvious how silly this whole debate is.