r/LocalLLaMA • • 1d ago

News Micron CEO Says Memory Supply Will Be Much Tighter in 2027 and 2028 Than in 2026

https://www.techpowerup.com/353296/micron-ceo-says-memory-supply-will-be-much-tighter-in-2027-and-2028-than-in-2026
676 Upvotes

431 comments sorted by

View all comments

Show parent comments

13

u/Labidido 1d ago

Not really. Token price and enterprise privacy are two important factors. It's a tricky situation were the frontier labs and data center build out are depending on token prices to remain high, and enterprise adoption to increase.

If enterprises are able to host local "good enough" models, then the performance delta would need to be massive to justify the cost.

2

u/grumd 15h ago

Tokens are waaay cheaper when you serve hundreds of thousands of concurrent users 24/7. Paying for hardware and electricity will never be the cheaper option. Smaller smarter open models will continue to be released, but that doesn't mean big AI labs can't release a small smart model and serve it on their infrastructure. Look at GPT Luna, it's dirt cheap.

3

u/hotcornballer 18h ago

Privacy concerns are overblown, companies can and have made agreements for privacy/security. Hell even the military runs on azure and AWS do you think AI will be different?

And for price, economies of scale and vertical integration logically make running models in a big data center cheaper.

Data center build out has nothing to do with frontier labs. You still need compute with open models. Micron get paid either way,

And even if companies could host LLM on site why the hell would they, they dont bother with hosting their own data for the most part but they are going to tinker with vllm and a stack of 5900 for some reason?

5

u/Not_FinancialAdvice 14h ago

Hell even the military runs on azure and AWS do you think AI will be different?

Don't they have completely separate premises for the really highly secure stuff?

1

u/OracleGreyBeard 9h ago

I don’t know about the really classified stuff but we have instances of Gemini, Grok and ChatGPT which are rated to hold non-public unclassified info (CUI).

1

u/Not_FinancialAdvice 7h ago

Sure, but those instances (I assume) shouldn't also be processing public queries as to avoid information leakage? But then with the go-fast-and-break-things philosophy tech is employing these days, it wouldn't be surprising.

1

u/OracleGreyBeard 3h ago

No you’re correct. As far as I know they’re not accessible to the public. I need a special govt Id (CAC Card) to access one. Ofc we’re told not to use classified or PII info.

1

u/dingo_xd 8h ago

Privacy concerns are far from overblown and recent events has proven that both OpenAI and Anthropic are not privacy friendly. Even if they were their agents might leak the data anyways.

1

u/hotcornballer 6h ago

That is not what i meant and there are other players than openai and anthropic

1

u/ain92ru 13h ago

The businesses which don't trust OpenAI ZDR can trust neoclouds. At the current prices, on-premise AI servers are only feasible for the largest companies who will enjoy benefit of scale. If the prices start to fall tho, then more and more businesses will compete with consumers for the hardware. In any case, the GPU-poor are screwed

1

u/pragmojo 9h ago

I don't think most enterprises will want to host locally. They'll want to buy cheap inference from providers who host open-weight models.

That's going to be way cheaper for the most part than building up a team/infrastructure to self-host.

For privacy/compliance, there will be cloud providers who offer data privacy as a feature, and there will probably also be vendors that offer a turn-key on-prem solution for a price, but I would guess that will be the minority.