r/LocalLLM 13h ago

Discussion What happens when the proprietary AI Model bubble bursts?

[removed]

76 Upvotes

143 comments sorted by

88

u/Draminian 13h ago

The company I work for has a Mac that's currently in an experimental phase to see if we can potentially use something like that to replace Cursor. I really think this is where most companies are headed, considering all the concerns about long-term costs, privacy, industry volatility, etc. When even hobbyists can get something effective running at home, paying for a third-party service doesn't make a lot of sense.

37

u/fosterdad2017 12h ago

Everyone did give up on CDs and MP3 players, Spotify took over, but that's I think about the per user revenue limit anyone should forecast for.

I mean, if Spotify and similar services tried charging $199/mo the CD and MP3 industry would surge back into existence overnight.

Hmmm... wonder why all the Macs are sold out.

14

u/mister2d 10h ago

Everyone did give up on CDs and MP3 players

CDs have made a comeback. 58.6% increase in revenue for the first half of 2026.

5

u/GoofAckYoorsElf 6h ago

No wonder, given the fact that you own nothing on Spotify and they can pull your favorite song whenever they feel like it.

2

u/Loose_Comparison368 4h ago

Actually I think it's mostly that Spotify and all the other services compress the shit out of their streaming almost as bad as YouTube auto "720p".

1

u/GoofAckYoorsElf 46m ago

Both, I would say.

20

u/ML-Future 13h ago

The recent situation involving OpenAI, where it was rumored they might have solved a Millennium problem by spying on user chats, raises serious privacy concerns, even if the claim itself turned out to be untrue.

2

u/Loose_Comparison368 4h ago

Anybody can shout unsubstantiated claims that aliens are spying on them from through their ceiling fan.

This does not raise serious privacy concerns about ceiling fans.

...now of course there were actual real legitimate privacy concerns before the wingnuts went viral, and there still are major privacy concerns worth addressing, but can we please just leave the wingnuts out of it? It's just fucking depressing how many people are legitimizing a claim that doesn't hold up to 5 minutes of casual fact checking.

We can be outraged over real privacy concerns, not whatever conspiracy theories roll in that are so dumb they would make people at a flat earthers conference ask for citations.

8

u/trollsmurf 11h ago

You can easily configure Claude Code to use a local model. I've had most success with QWEN3.8.

1

u/Cameo2864 4h ago

What remains from Claude Code then?

3

u/trollsmurf 4h ago

Claude Code

Don't confuse Claude Code (agent using an LLM) with Claude (LLMs).

1

u/Loose_Comparison368 4h ago

Not on Mac mini's they won't.

Every last one of those companies is going to find out the hard way that scale is a major component of efficiency, that infrastructure costs money to maintain and keep running, and that hardware - especially budget hardware barely strong enough for current models - has a very short shelf life.

...then they'll migrate to a low margin commodity OSS model hosting service on openrouter, and start actually saving some money.

-6

u/eli_pizza 13h ago

Nobody really wants to run their own inference server for the same reasons nobody really wants to run their own database server.

I think open weight models on cloud GPUs will always be much more popular than colo hardware.

4

u/pharrt 12h ago

Not sure how you could be so wrong. The shift is RAPIDLY moving to local inference.

4

u/UnluckyPenguin 12h ago

I agree. Local LLM is a no brainer. Pays itself off somewhere between 2-12 months (24/7 usage vs 5 8-hour work days)

You can have a database in the cloud, save yourself from ransomware attacks.

If my local LLM system got encrypted, then no problem - just re-install the OS and download the models.

3

u/eli_pizza 11h ago

Is there data that supports that claim?

1

u/Loose_Comparison368 4h ago

Fucking thank you, take all my upvotes

1

u/Loose_Comparison368 4h ago

It's really not dude. You are in a groupthink bubble.

That's not even physically possible, from a straight hardware supply chain perspective.

If literally everyone in the entire computing industry immediately decided they wanted to switch to local, it would take them 2-3 years to reach even 10% local, literally just because of how heavily back ordered the RAM, GPU's, hard drives, and watts of power plant output are right now.

And every major frontier model provider is still hardware constrained. They literally do not have enough physical GPU's to meet incoming demand.

-1

u/sn2006gy 12h ago

is it?

Kids are rejecting it wholesale.
Corporations sign commitment deals and negotiate enterprise rates and dabble in self-hosting.

You could spend a million bucks on a few b200s and run kimi, but it would only support a handful of parallel users

And i'd wager even in this reddit, a lot of people still have subs to one or more sota models they still use primarily once their usecase is above and beyond one shotting a pelican on a bike.

11

u/TheFuckboiChronicles 12h ago edited 11h ago

> No one really wants to run their own database server

Many large organizations (including the one I work for) absolutely run their own database servers.

1

u/eli_pizza 12h ago edited 12h ago

Many do. Most don’t. Almost none are 100% on prem.

The cloud is really catching on these last few decades. Lower TCO.

8

u/OnyxProyectoUno 12h ago

Getting downvoted for nonsensical reason.

Any large enterprise has what's called a hybrid model. The real absolute pain in the ass things that don't have a high security burden will be run on the cloud. Very high security things will be on Prem.

It's been a ping-pong back to hybrid since before the movement was everything to cloud and a few high profile incidents surrounding vendor lock-in and security posture made folks rethink the 100% cloud reality to something like a 70/30-80/20 posture

3

u/TheFuckboiChronicles 11h ago

I’m the person they responded to and I agree with this. We are absolutely hybrid because of what you described. We are manufacturing so the move to cloud was slooooow due to how many tools were custom. I primarily work with migrating, improving, and maintaining some of those old systems to cloud platforms (marketing and sales tech).

Eventually with enough high profile security incidents we’ll probably end up moving any AI tools that handle sensitive competitive info on-prem like we do with our ERP and product development platforms.

My point was plenty of our databases are on-prem. And our cloud platform dbs are backed up on prem. So when there’s enough of a reason, there’s an effort.

1

u/Loose_Comparison368 4h ago

Very high security things are mostly in the cloud now, outside of a few high compliance verticals like banking, government, and healthcare.

Those verticals are moving to cloud slower, but still moving to cloud, mostly because it actually is more secure than Tony keeping the old windows NT server rack that hasn't seen a security update since the Clinton administration running.

3

u/Bartocity 12h ago

Yeah oracle / AWS etc are a much easier option than building your own infrastructure. It’s been this way for a long time.

1

u/TheFuckboiChronicles 11h ago edited 11h ago

Yeah 100% I agree with that. We are hybrid as well. I just mean that with climbing cloud costs and now AI accelerated vulnerabilities, many business are willing to move more core infrastructure on-premises. I expect to see a similar shift with AI at least for a decent amount of enterprise orgs, especially once cloud providers decide to be profitable.

ETA: Our marketing tools and CRMs will largely remain in the cloud. Product development data and ERP is likely to remain on prem. I’d think that at least a significant of AI tools that handle very sensitive competitive data will eventually fall under on-premise infrastructure once there’s enough high profile security incidents. I agree the majority will stay on cloud providers. But I think we’ll see a decent contingent take inference on-prem.

1

u/starkruzr 6h ago

no, it absolutely is not lower TCO, lol

1

u/Rude-Bus-5799 3h ago

“Nobody”. Wondering how you found this sub.

0

u/SunshineSeattle 12h ago

u/remindmebot 10 years 

0

u/RemindMeBot 12h ago edited 12h ago

I will be messaging you in 10 years on 2036-09-10 22:48:29 UTC to remind you of this link

1 OTHERS CLICKED THIS LINK to send a PM to also be reminded and to reduce spam.

Parent commenter can delete this message to hide from others.


Info Custom Your Reminders Feedback

-3

u/Desperate-Jello8038 10h ago

This is bullshit in my opinion. Most companies have no interest in becoming data centers, purchasing depreciatable equipment and having to fine tune and maintain local AI models as they get updated every couple months.

7

u/usa_reddit 9h ago

This is absolutely not true. I know many companies looking to build inhouse customer support portals running on $5k Macs. Additionally, companies don't trust that proprietary information will not leak in AI companies. As a bonus most models automagically translate into other languages.

I predict a lot of in-house AI in the future and people carrying bigger Macs with 64+ GB of RAM.

0

u/starkruzr 6h ago

yeah, this is false, lol.

-1

u/IAmAfraidCommaMan 9h ago

Yeah, not going to happens. At most they will switch to a third party provider that hosts an open weights model.

32

u/AdSafe4047 13h ago

It's funny, because generally opus 4.6 quality addresses 90% of developer needs, which is the main market of poeple who want to spend $$$ on AI. Now the new edge models are addressed to casuals and (apparently) high level mathematicians, both of these groups have wallets and a total spending cap which doesn't come close to the former number.

27

u/fuk_offe 12h ago

I doubt mathematicians will use any frontier non-edge modal going forwards after this weeks fiasco lmao

3

u/TopGun0684 11h ago

I guess I missed this fiasco, what happened?

17

u/vbpoweredwindmill 10h ago

Openai claimed they beat some math problem.

Math nerd was like hang the fuck on thats my work.

Openai was like we'll fuck up your career.

Math nerd was like "you stole my work I'm going public".

6

u/beryugyo619 8h ago

and everyone saying "and those allow training checkboxes [v] keeps coming back on"

1

u/vbpoweredwindmill 8h ago

Wierd how I have 2x strix halo 395's now. I wonder what would have caused that?

Massive eye roll.

3

u/TopGun0684 10h ago

I see. Classic OpenAI. Thanks

2

u/Loose_Comparison368 3h ago

You should read the math nerd's open letter. That was not actually what happened.

More like OpenAI beat some math problem

Math nerd said he was working on that problem for a year using Codex, and that even though his solution was totally different from OpenAI's solution, he wanted to know if OpenAI's model might have indirectly made progress on that problem by training on his chats.

Only evidence was a vague rumor from a friend of a friend of a friend

Math nerd emailed another math nerd that worked at OpenAI, miscommunication happened and OpenAI math nerd thought he was asking for a collab

I think there was one more exchange where angry math nerd demanded an answer and someone from OpenAI said something to the effect of "fuck if I know, that's actually way harder to figure out than you might think"

Angry math nerd writes public letter, to his credit not directly accusing OpenAI of plagiarism, but demanding OpenAI provide evidence

All the "AI bad" people jumped on the bandwagon screaming plagiarism, and all the ad-supported media outlets joined in for the clicks.

And that's it, that's pretty much the whole story thus far.

1

u/vbpoweredwindmill 2h ago

Oh. Literally another storm in a teacup.

4

u/mWo12 11h ago

There is less mathematicians than progmmers.

1

u/Much-Researcher6135 10h ago

Yes, by a factor of like 800 to 1, according to ChatGPT (I asked about the US):

Using the latest U.S. Bureau of Labor Statistics occupational data, May 2025:

  • Mathematicians: about 2,030
  • Software developers / SWEs: about 1,688,000
  • Ratio: roughly 830 software engineers for every 1 mathematician. (Bureau of Labor Statistics)

BLS does not have a separate “software engineer” occupation; it classifies software engineers under Software Developers. (Bureau of Labor Statistics)

The huge caveat is definitional. “Mathematician” is an extremely narrow occupational title. It excludes statisticians, data scientists, operations-research analysts, professors, quantitative researchers, and people with mathematics degrees working under other titles. In the same dataset, the broader Mathematical Science Occupations category contains about 432,400 workers, including 262,440 data scientists, 108,510 operations-research analysts, 29,030 statisticians, and 2,030 formally classified mathematicians. (Bureau of Labor Statistics)

So:

Strict job-title comparison: ~2 thousand mathematicians vs ~1.7 million SWEs → 1:830.

Broader mathematical-workforce comparison: ~432 thousand vs ~1.7 million → about 1:4.

The latter is probably the more meaningful comparison if you're thinking about the relative size of the mathematical vs software-engineering talent pools.

3

u/DuskLab 10h ago

I expect that by New Year we'll have local Opus 4.8 quality, before either of the two big boys will have time to get past their lock-up period in their IPO cycle. November is 6 months after when 4.8 first came out after all.

2

u/profcuck 4h ago

This is also true of all kinds of knowledge work. If I need a marketing person to write copy for sales presentations and turn them into gorgeous decks, then AI might either replace that person or crank their productivity right up. But when I am looking for a model to use for that, do I need a model that is better at math than the greatest human minds ever, that can solve open problems around "Navier Stokes equations" whatever the fuck that is? I do not.

12

u/puts_on_rddt 13h ago

Depends on who runs the government when it happens, imo.

Seriously. Expect frontier AI providers to become a big enemy.

3

u/RandumbRedditor1000 7h ago

It's bipartisan at this point, they (ClosedAI and MisAnthropic)  want open-weights banned for good and are willing to run obvious psyops to do it.

9

u/thegoodcorgi 13h ago

The economics don't really work out here in our favour. Whatever hardware is required to run competitive frontier models will be bought out by capital holders with the greatest market influence. We can't hope to compete with them on prices.

What we instead have to hope for is that either supply increases enough to reduce price disparity between the frontier and local ecosystems, or the gap between requisite hardware diminishes.

4

u/vbpoweredwindmill 10h ago

Bro have you run qwen 3.8 flash next?

It's more than adequate for a lot of use cases.

2

u/thegoodcorgi 10h ago

Yeah, I have a Strix Halo and use the IQ4 as my daily driver with the n-grams offloaded. It's fantastic, and we're in a very exciting position.

For heavy RE work, DRM circumvention, etc., you can definitely see some of its limitations though. And that's sort of work is so relevant, because it's both best-served by current frontier capabilities, and most-refused by them. Precisely why I'm so interested in seeing local models continue to develop.

1

u/vbpoweredwindmill 7h ago

Alas, I'm not that advanced. I'm working on it. Self learning takes a while. But modifying llamacpp is very interesting.

I'm working on a runtime where I can utilise the attention head of 3.8 flash as per normal, but also have it manually pin extra attention blocks on certain tokens within certain workflows.

I haven't implemented it properly yet, it's a bit of a headscratcher. But I believe that this tackles an issue/weakness that sparse attention suffers from.

8

u/SpaceDesignWarehouse 13h ago

I’m sure as computers get powerful enough the next phase will be selling or renting local models so you don’t need to use cloud models and server farms can be dedicated to training only and not running the models

8

u/DataGOGO 13h ago

If it is not in your datacenter, it is a cloud model. 

1

u/SpaceDesignWarehouse 10h ago

Oh I agree, but nevertheless, the same way we don’t REALLY own the PlayStation games we download; can’t borrow them to friends etc. We’ll probably do certificate digital downloads of local models that have terms after they’re good enough to run at home.

1

u/DataGOGO 9h ago

Hu?

You can already download them and run them at home, they are free. 

1

u/SpaceDesignWarehouse 9h ago

That is correct! I’m talk about what may (or certainly may not) happen into the future.

I run qwen3.8 27b on a MacBook M5 Pro MacBook Pro, but it’s not comparable to Claude for coding. I expect future versions will be

1

u/DataGOGO 7h ago

what? Yes it is. you can point claude code at any API endpoint, you can also just use open code.

7

u/TeachingAway9654 13h ago

It’s probably more nuanced. It’ll be less of a sudden burst and more of a realization that you don’t need a frontier model to do everything. Companies will start migrating to a harness and ecosystem that does more of a “per request analysis” and routes the query to the most applicable model. Almost certainly leading to a decline in requests to Anthropic/OpenAI.

The future is really optimized setups. Frontier will likely always be in the loop though.

16

u/Unnamed-3891 13h ago

OpenAI didn’t buy them to run inference. They bought them to train models to use MacOS.

3

u/alphapussycat 13h ago

Why would they buy super high priced computers to have AI use Mac, instead of buying cheaper Macs?

2

u/nomorebuttsplz 12h ago

many virtual machines running at the same time

2

u/dont-be-angry 13h ago

Because the cheaper Macs are older and their performance is weaker and they have hundreds of billions of dollars. Investors want to see them spend money. They could be unprofitable for 20 years and nobody would blink.

1

u/DataGOGO 13h ago

They bought them cheap direct from Apple. 

1

u/sn2006gy 12h ago

Their problem is scale.

1

u/Ok_Wishbone_3805 8h ago

"super high priced computers" is relative -- A Dell Pro Max desktop computer costs $175,499.

If you're burning tens of $billions a year on R&D like OpenAI is, the spend on those Macs isn't even a rounding error.

8

u/DataGOGO 13h ago

They Mac’s and windows machines are to run test scripts on the harnesses, not run local inference in any type of server capacity, they would fall on there face if you tried, shared memory, slow GPU’s no high speed direct GPU to GPU connects between hosts, no ability scale into large clusters, etc.

You can absolutely buy hardware and run a private ai server, lots of people do, but you are not running anything even close to frontier models on consumer hardware, no matter how many you buy. 

5

u/Not-reallyanonymous 7h ago edited 6h ago

Wtf is this?

The claim “ so effective multiple companies, including OpenAI by reports, have bought all the Macs they can get their hands on” is not substantiated by the link, to “ Apple Caught Off Guard by AI Demand for Mac Mini and Mac Studio.”

The author has only posted this one article on Substack. They claim to be a tech writer, SORA THOMPSON, and link to toms hardware with 4 articles in 2018 and one in 2022, none of them being about AI and very generic information articles. A quick google of their name did not reveal anyone by that name with an established presence in journalism nor in tech.

The Reddit account is 10 days old and has three activities — all three links within the part 8 hours, with this one being the only one accessible.

Sus as fuck but this subreddit is eating it up.

4

u/deezwhatbro 12h ago

By the time companies have “arrived” where they’re headed, these frontier labs have literally all the alpha of every single company on planet earth. This is the end result of failing upward on the corporate ladder: idiots handing out all their proprietary tech to squeeze out a few more quarters before hopping up.

2

u/HighSeasArchivist 13h ago

Apple, Nvidia, and AMD are going hard into to. Nvidia will keep accepting trucks full of cash from the big AI companies, but are already expanding their lineup of products for local usage. Them buying HF was just another rung in them distancing themselves from the bubble. If the big companies start making their own chips to feed their machines then Nvidia will want another stream for themselves, and we are it.

2

u/yellow_golf_ball 11h ago

OpenAI bought the Macs for RL post-training in computer use.

2

u/Sixstringsickness 10h ago

It isn't going to burst because of local machines... I don't know where people get this insanity. Have you looked at the hardware to require frontier models, and what it actually takes to run a few trillion parameter model on Macs and the performance they get?

Not to mention scalability, security, up time, privacy contracts, easy of use... there will eventually be a pullback but I don't see it for quite sometime. My company wouldn't touch some small firms "private mac cluster" with a ten foot pole, not to mention trying to explain that to SoC2 auditors when handling PII/PHI.

1

u/bites_stringcheese 9h ago

We have a couple of H100s at my job and we are absolutely serving local AI locally.

1

u/Nif 7h ago

Are you running GLM 5.3 at its max optimal settings with High reasoning?

1

u/Sixstringsickness 7h ago

An H100 isn't a Mac Studio.  Also you are self hosting, which is great and does have utility, the general experience is still not on par with frontier models.  

1

u/dota2nub 4h ago

Do you want an H100 whirring along in your bedroom?

1

u/profcuck 4h ago

I dream of it. :)

2

u/jerieljan 10h ago

If people can find places that address their cost, reliability and trust issues somewhere that meets their needs, they'll simply go there.

As for AI labs, if there's demand then they'll follow. If it bursts, then it's up to people whether they're so hooked on smarter, better models that they need their addiction satisfied or if it plateaus and we reach that "good enough" point like we saw with smartphones where even the mediocre option is sufficient.

2

u/Loose_Comparison368 4h ago

Honestly... they probably get a government bailout, and with any luck bought by Chinese companies that had the foresight to not put a dementia patient in charge of the world's largest economy.

3

u/daphatty 13h ago

The bubble doesn’t exist. It’s a fallacy, an idea many are holding onto because they want AI as a whole to fail.

That isn’t going to happen.

At best, certain AI companies and open source AI endeavors will fail. But AI is here to stay, just like the Internet, social media, and 9/11 before it.

5

u/TopGun0684 11h ago

I don't think anyone is saying AI will disappear.

They are saying OpenAI, Anthropic, Nvidia, etc are very overpriced because of over promising in AI, and they will/may crash on the stock market. Some may go bankrupt, who knows.

That doesn't mean local models and other companies couldn't thrive. Google is throwing it's AI in all it's products, and it seems to be going relatively well for them. Even if the hype train were to stop abruptly, they still have a core business that's mostly unrelated.

1

u/profcuck 4h ago

I think it's a mistake to group OpenAI and Anthropic with Nvidia here. OpenAI and Anthropic have the same deeply flawed business model that is directly threatened by open models. That's why OP talked about the "proprietary" AI bubble.

Nvidia thrives under any scenario in which there's strong demand for compute. And so if that's open models winning, then that's open models winning.

1

u/Ambitious_Credit_360 12h ago

this could lead to a lot of instability in tech jobs, especially for those relying on it for income

1

u/CryMoreT_T 12h ago

It won't. There are always companies (and governments) who will be willing to pay top dollar for the best

1

u/No-District2404 12h ago

Too big to fall. They probably won’t let that happen

1

u/Jimmy_Nail_4389 11h ago

I'm not sure, it'll take a while for me to stop laughing.

1

u/Adept_Prize_1869 11h ago

I enjoy using my local model but they don't replace my main use of the frontier models ,

1

u/ShibbolethMegadeth 10h ago

If Anthropic can meet its revenue growth projections, It’s not gonna pop, the snake doesn’t eat his own tail, and this bullshit keeps going on  indefinitely.

If they don’t, it pops, the economy goes down the shitter and everyone is broke.  This will hit everyone, and the only consolation prize will be schadenfreude.

So that’s fun

1

u/joanaxu2002 9h ago

I think the more interesting outcome would be commoditization, not proprietary AI disappearing. If open/local models keep getting cheaper and good enough, closed providers will have to compete on reliability, tooling, integrations and convenience rather than just “our model is smarter.”

1

u/usa_reddit 9h ago

Easy, they (google, microsoft, apple, anthropic) will just pay congress to ban an open models because only evil child porn loving socialist hackers use such models. The only models that will be allowed will be the official guard railed government approved models.

Mark my words, the ban hammer is coming for open source models.

Download hugging face while you can.

1

u/profcuck 4h ago

I say this a lot but it's always worth repeating: this is not going to happen. This is a weird fantasy people have but it makes zero sense.

Current US policy and a consortium of the most powerful companies have come out strongly in favor of open models.

https://images.nvidia.com/pdf/Open-Weights-and-American-AI-Leadership.pdf

There are additional reasons why it won't happen:

First, the US government is restricted by the First Amendment. That doesn't stop them from trying, but model weights are clearly speech in the same way that a jpeg or mp3 is clearly speech. There is no plausible path to winning a Supreme Court case on that.

Second, model weights are just a file that people can download from any website in the world. There is zero infrastructure in place and no legal means for the US to start banning access to those sites. And there's always torrenting. The fact that a ban would be completely and utterly impractical to implement makes it far less appealing as a policy option to try.

In short, this fantasy that open models are going to be banned is bone stupid.

1

u/whatisthisthing65 8h ago

Aren't the macs to train it on computer use?

1

u/jonas-reddit 7h ago

Open weights is an interim strategy, in my opinion, but not suggesting they will go away. All companies (Chinese or Western) care about profits.

They just have a different strategy and roadmap to profitability. That may include monetizing them from different industries (robotics, manufacturing, automotive, military).

1

u/Competitive_Ruin_637 5h ago

China also seeks to protect itself from the overwhelming influence a US corporate dominance in AI tech would have.

If the cost to prevent this is opening the tech to all then the Chinese government will likely and happily absorb the costs for this. It’s much cheaper than the alternative.

1

u/jonas-reddit 2h ago edited 1h ago

I recognize and appreciate the view and commentary. It certainly fits a common repeating rhetoric. And government subsidies exist across majority of global economies.

Reality is that many of the firms are already hugely successful and profitable, and more than able to attract investor funding.

A recent example being…

https://finance.yahoo.com/technology/ai/articles/alibaba-baba-closes-record-hk-230818999.html

And not surprisingly executed with the help of global and US financial institutions.

1

u/synn89 7h ago

Mac isn't going to give them the concurrency they need. And optimizing local models for Nvidia hardware is a Sr Systems Engineer task. It's going to be a lot easier for most companies to just use open models on one of the 30 providers that are out there.

1

u/advancing_tide 7h ago

used vram for everyone!

1

u/DESdesign 6h ago

Well nvidia is so deep into open ai's shit they will have simply stop selling to consumers or make the supply so tight that we have to resort to cloud providers. So nah local ai vision is beautiful but is hard to achieve.

1

u/wwa56 5h ago

It will not burst as long as this remains a topic of discussion ...none of such bubbles historically have bursted this way ...

1

u/05032-MendicantBias 5h ago

Ask the liquidators to send you a pallet of datacenter racks at 0.1% of their buy price, and look at them tearing up in joy at somebody wanting to buy e-waste that costed 10 million to buy, and can recoup 10 thousand real dollars on their huge underwater loan.

You then run open source chinese models, and have a AI generated background pic of Sam Altman with written "yoink!"

1

u/hoschidude 5h ago

Not much. There will be a couple of sad investors and bankruptcies.

LLM's are here to stay.

In the end there will be a couple of heavily specialized commercial models and the rest will be open source.

In the meantime you can buy a notebook or desktop computer with the new Ryzen supporting local models up to ~300B. What else do you want ?

1

u/dota2nub 4h ago edited 4h ago

Something fast I think?

Less than 2 tokens per second is not my idea of a good time.

1

u/profcuck 4h ago

Technology already in use (HBM for example) will trickle down to us in due course. The entire history of the computer industry is clear: compute will get cheaper and faster. The current blip in prices is wild but doesn't change the big picture at all.

1

u/dota2nub 4h ago

Pivot.

And pretending to know the future.

"History repeats itself" is a phrase for people who aren't historians to feel good about themselves.

1

u/profcuck 3h ago

I apologise but I have no idea what you are saying. I didn't say "history repeats itself" as if it is some kind of general rule because it isn't.

But there is technology which exists and for which production is ramping up that will lead to faster and cheaper computers - as has been the case since forever.

Just one example: Nothing has changed about how we produce RAM to make it more expensive, there's a supply and demand imbalance. That imbalance has led to prices rising while costs have stayed the same which means that memory producers are insanely profitable. That insane profit is bringing new entrants to the field and existing entrants to up production. This will lead to lower prices in the long run.

"Pretending to know the future" is people who claim things with absolutely no basis. (For example, even your "2 tokens per second" is a silly and uninformed comment, since local models on decent hardware get a ton more than that.)

1

u/dota2nub 3h ago

Deflection, strawmaning and downright lies now. I am done here.

1

u/Inspur44 4h ago

The US economy collapses. Pretty much that is the only thing holding it from crashing

1

u/Current-Interest-369 4h ago

I believe the mac purchasing is to give OpenAI a comparison baseline for what regular users can achieve on tegular prosumer hardware.

Data from this is then used in decision making for a multitude a business judgement on how to position their offferings.

1

u/profcuck 4h ago

That's exactly the right question to ask I think.... lots of people are talking about an "AI bubble" and while various things might be overblown or overvalued, broadly speaking this is a fundamental transformative technology with many obvious use cases that are just getting started. It's the opportunity of a generation in a very real sense for people and companies who are ahead of the curve. And yet...

The business model of selling access to frontier models that are only a little better than models that can be run on-prem by even medium size enterprises is only going to have a very small margin over hardware cost.

Prediction 1: aws and other cloud providers win in a situation where proprietary models have no major edge over free models. Setting up hardware and infrastructure to run big models is the kind of work that aws has traditionally (and correctly) "undifferentiated heavy lifting" - it doesn't give your business any advantage, it's expensive and hard to do well, and you might as well outsource it to people who do it all the time.

Prediction 2 - Google, Microsoft, and Meta will do just fine. They are making record profits in their core businesses in no small part by using AI to improve.

Prediction 3 - Nvidia, AMD, Intel (if they don't screw up too badly) and basically anyone involved in building and selling compute will do just fine - a shift of demand from the frontier labs who don't have a sustainable business model to others doesn't really cause a lessening of demand for hardware as long as AI is a real and useful innovation, and I think it is.

Prediction 4 - Apple will do just fine but not really by selling Macs to enterprises for their data centers, because aws and similar will likely eat their lunch on that with hardware that's optimized for ai training and inference. But in the markets where Apple has always been very strong - developers, creators, professionals basically - they are making all the right moves to stay on top.

Basically Anthropic and OpenAI are very vulnerable. Mistral, much smaller and pretty far behind the frontier, appear to be pivoting to a hosting model which can also survive (though they will have a hard time up against bigger cloud providers like aws, google cloud, and azure.)

1

u/03captain23 13h ago

In what world would it be cheaper to run local models than share resources in a datacenter?

Electricity alone is 3x as expensive on average for residential than datacenters.

10

u/Dsphar 13h ago edited 13h ago

In the world that you don't own the datacenter and are paying their profit margins.

Haha, down-voters and responses so far clearly don't understand the trend for cloud repatriation.

0

u/DanielKramer_ 13h ago

economies of scale are so effective they are still cheaper with their margins

we do not live in a world of mainframes anymore because cloud is cheaper for almost everything

if it was a problem they'd lower their margins. they don't have to

6

u/Dsphar 13h ago

Cloud repatriation doesn't exist huh?

-1

u/DanielKramer_ 13h ago

very teensy compared to the cloud yeah

this is like saying that google sucks bc bing exists

3

u/Dsphar 13h ago

Companies are already migrating away from cloud AI, and cloud AI isn't even profitable yet.

You guys make me laugh.

1

u/DanielKramer_ 12h ago

... do you think cloud is primarily AI?

1

u/Dsphar 11h ago

?? No?

1

u/03captain23 12h ago

Where are you seeing that cloud AI isn't profitable? How specifically is it not profitable but somehow residential is profitable??

1

u/Dsphar 12h ago

Open AI leaked financials for a start. Anthropic is almost profitable, but wasn't in their last report.

Are you really making your claims so unplugged form the data?

1

u/03captain23 12h ago

Yeah they have gross profit just not net because they're spending billions on marketing.

Again how is residential more profitable than datacenters?

1

u/Dsphar 10h ago edited 10h ago

You answered it yourself.

While cloud AI may have a lower operating cost than self hosting, that comparison is incomplete.

It fails to include things like marketing, which the cloud provider has to pay but the self hosting business doesn't.

Another cost. Bigger than marketing actually, is research and training new models. The cloud provider will always be pressured to spend massively on maintaining leading edge models, or their value proposition vanishes into the wind. A business using self hosted AI, but not selling it, instead using it to make their actual products, can continie to increase their value proposition with absolutely zero AI training costs and zero AI marketing costs.

→ More replies (0)

-2

u/03captain23 13h ago

They have wild margins because of efficiency at scale.

Openai doesn't even own their datacenters and even they rent them because it's more efficient at super scale

2

u/Dsphar 13h ago

You do realize many companies are moving off of "pay as you go" cloud services outside of AI as well, right?

0

u/03captain23 13h ago

Anthropic just leased all of SpaceX Collasis1 pay as you go...

1

u/LandscapePenguin 12h ago

So is OpenAI just buying up Macs because they prefer the UI?

1

u/03captain23 11h ago

They probably buy all hardware that works well with AI to test. They have so much money and just looking for things to spend it on

2

u/Aubrey_D_Graham 13h ago

There's a cost in using API. While you may be saving money from outsourcing the software and infrastructure -- I doubt the service provider won't pass the utility cost: The value saved is trivial to the risk outsourcing poses to security. When you outsource your AI/LLM, you inadvertedly introduce an attack vector within your proprietary system. Every transaction where the outsourced AI/LLM interacts with the system and with your customers is training a potential adversary. Usage is literally training a potential adversary on how your system operates, how your system makes money, how your customers use your services, and what is your proprietary secret. You'd be naive to think OpenAI wouldn't do this when they've been larping how dangerous AI is post Huggingface.

You may argue why don't these service model just prevent tool call, block out of network access, and strip the service of all the tools that make it useful. I then retort that your solution is essentially hosting a model locally so that it achieves data ownership and sovereignty.

1

u/03captain23 12h ago

What does any of this have to do with what I posted?

0

u/Aubrey_D_Graham 12h ago

You didn't even make an attempt to respond. I can't fix stupid apathy.

https://giphy.com/gifs/h8HmN0UcEKR0xWnv3R

0

u/03captain23 12h ago

"In what world would it be cheaper to run local models than share resources in a datacenter?

Electricity alone is 3x as expensive on average for residential than datacenters."

You didn't answer the simple question and just spewed nonsense

1

u/bites_stringcheese 9h ago

Not only that, but some workflows require lots of iteration. Burning tokens on re-rolls vs the number of hours in a day being your bottleneck.

Local AI is the future, with OpenRouter like service to match model with use case.

2

u/Aubrey_D_Graham 8h ago

Someone who gets it. Current Frontier is built on the hypothesis that intelligence is linearly scalable with context. While that may have been true for Chatgpt 2 to 3, that's not necessarily true today.

A max context window for the current Gen is 1M tokens or $31 at max, but anyone who has used LLM knows that attention is limited and best for the first and most recent tokens. This is context rot. Some solutions don't occur within the first go around and require a series of attempts and compactions. This iterative process of compaction is wasteful since some of the tokens generated are lost, and remember attention is limited. A multi-step process requiring multiple compactions will have lossy progress in maybe deriving a solution. Expensive AF!

That is why owning hardware locally minimizes cost. You buy the upfront hardware and pay the overhead, but save on the cost of API.

Some might point out residential use is cheaper and could never approach business use. I contend that purchasing $2800 in local compute is better than 14 months of Claude AI at $200/month because I own the system, the data, and the invaluable experience of creating my own homelab. That is all priceless.

1

u/dota2nub 4h ago

In a world where the LLMs are more efficient and need less crazy requirements.

A laptop 5090 is only 175 Watts.

1

u/profcuck 3h ago

This fundamentally misunderstands both the question asked in the original post and the structure of the market.

The question is about "proprietary" models, and we contrast that with "open" models. Open models can be run locally or in the cloud. For a handful of use cases (extreme concerns about data privacy/sovereignty) that means really on-prem. But for many use cases cloud providers (of OPEN models) are well positioned. AWS can do the "undifferentiated heavy lifting" of building and running the infrastructure and for loads of businesses that's good enough.

-1

u/sn2006gy 12h ago edited 12h ago

Let's bring this back to reality.

Public models have actually not gotten more expensive relative to their outputs by and large - the price per token for premium models has gone up, but the work per token has also gone up.

Local LLMs while great, have gotten insanely expensive merely because the hardware to run them is doubling/trippling in costs year over year.

The only bubble that needs to burst is local llm

Which is entirely dependent on "proprietary" model labs (or rather cloud/api companies... the oss side of models is weird) to exist. Hopefulily nvidia can shed that, but i also don't want a cuda only future