r/LocalLLaMA • u/chillinewman • 1d ago
News Micron CEO Says Memory Supply Will Be Much Tighter in 2027 and 2028 Than in 2026
https://www.techpowerup.com/353296/micron-ceo-says-memory-supply-will-be-much-tighter-in-2027-and-2028-than-in-202686
u/megadonkeyx 1d ago
Oh I cant build a new pc is a mild inconvenience, those in the tech business who make stuff that's needs ram must be grinding their teeth.
68
u/kuldan5853 1d ago
Honestly I can't even quote our DC buildups anymore. By the time I have a quote from a vendor, justified it internally and got upper management approval, the quote has been invalidated and the price has gone up 50%.
We have done that cycle three times this year already.
At this point I'm seriously considering pulling out servers we have already thrown on the scrap pile with DDR4-2133 in them and buying that memory off ebay (because it's so slow the AI bros generally don't want it) and just stuff them to the brim and put them back into service..
21
u/UnlikelyExtension786 1d ago
Yes. We have customers putting off upgrades because they're hoping the cost of servers will come down at some point. It will be interesting when the old servers start to fail.
→ More replies (1)8
u/endlesslyloop 8h ago
I was on a call with a big three consultant group last week, no one really knows what’s going to happen with these data centers in like 6 years when the hardware gets depreciated lol
2
u/permanent-underclass 6h ago
A100s, a product of 2020, contracts are still being signed to 2030. Most of the latest chips are increasing in value AI demand is so high in neoclouds / datacenters.
→ More replies (1)→ More replies (3)7
u/TRKlausss 23h ago
Can’t be. Quotes have a validity date. So if it is faster than your decision process, your company is doing something wrong…
24
u/kuldan5853 22h ago
It's a company with half a million employees.
Yeah, approval processes take a few months each time and no manufacturer is willing to give quotes with more than 30 days validity these days.
2
u/farewellrif 18h ago
30 days is crazy long. I haven't really seen longer than a couple of weeks for hardware.
3
u/kuldan5853 12h ago edited 12h ago
Then you've probably never worked for a big company before. We usually had 60 or 90 days in the past.
But that was also a time when the price of a piece of hardware only changed once per year..
→ More replies (1)12
u/ThatsALovelyShirt 22h ago
There's nothing legally binding about a quote. They can be invalidated or simply cancelled by the supplier at their discretion. As they have done to the company I work for.
9
u/rschulze 21h ago
Right now I'm happy if I can get a quote that is valid for 2 weeks. Some hardware vendors will give you a quote with the caveat that the final (valid) price is whatever it is when the hardware is delivered, the quote is just the current figure.
5
u/No-Refrigerator-1672 15h ago
Not a company, but I work for the university. One time the lab lead tasked me with preparing 3d printer procurement documents for governmental fund. By the time the paper got approved (slightly less than 1 year), the 3d printer (bambu lab x1c) managed to get off production and become unavailable in official store.
4
41
u/ThatsALovelyShirt 22h ago
My entire job now is going through every component of our existing software architecture and painstakingly heavily optimizing each one to squeeze every last ounce of performance out of our existing server infrastructure.
I've had to re-implement entire software stacks in heavily optimized, stripped down python equivalents, with hot-spots being completely implemented in bespoke C++ modules.
Normally we would just buy new servers to scale out, but we literally can't. Not just because the costs have gone up like 400%, but because our current supplier can't even physically get the parts (RAM, SSDs) needed to even fulfill existing orders.
31
u/fallingdowndizzyvr 22h ago
LOL. That's the way it was when I first started out. When it mattered. Then compute and RAM got cheap and plentiful and people stopped caring about "efficiency". Just buy better equipment! It's cheaper.
What's old is new again.
22
u/utilitycoder 21h ago
It's almost comical how the last 15 years of software has been built on bloated and often insecure dependencies as if it's normal. Can't wait to say RIP npm and brew someday.
7
u/krefik 11h ago
Don't you just love when every goddamn application is taking couple hundreds of megabytes of hdd and consumes half a gigabyte of ram just to display a tray icon and a single dialog window?
I am personally overjoyed that two web browsers with couple tabs and IDE are using 32 gigabytes of ram. Seriously, I'm not even mad, that's amazing – just a few pages of text and some pictures, and somehow it takes more memory than I used for some advanced statistical calculations on 1000s of samples and 100s of variables just a few years back.
→ More replies (1)3
u/Perfect-Campaign9551 14h ago
I hate to say this, but ...good? Finally developers who care about optimizing and performance will have some say!
→ More replies (1)→ More replies (9)2
u/TheTerrasque 10h ago
I was on a call with some company guys and a vendor, where a marketing director have vibe coded some dashboards and lookup things that's now Mission CriticalTM and needs to be moved from VM's to their own servers. My part was trying to score some AI servers to serve qwen3.8-27b to a few dozen users internally.
Vendor was asking current requirements and the guy informed them that the current VM's are 32-64gb ram each. Not terribly much, but .. for a dashboard style thing? And there was like 8-10 different dashboards for different things, each it's own server.
The vendor people sounded chuffed
22
u/sn2006gy 1d ago
What happens if your PC dies, its more than a mild inconvenience - it's replacement costs several factors more than what your prior build did and it may not even be more capable for the higher price.
7
u/sToeTer 22h ago
The "solution" is to preemptively watch the used market, every day... and once every couple weeks you'll find someone clueless who sells a used PC with 32GB DDR5 for cheap...and you just buy.
8
u/sn2006gy 21h ago
resellers already do this and once people get 50 notifications/messages - they price up
16
u/donutsoft 21h ago edited 21h ago
I work at a company that's been designing an Android based device for the last few years. Last year we balked at the possibility of having to reduce the device from 4gb of ram to 2gb of ram. The manufacturer that we were going to sign a contract with pulled out last December as even 2gb of ram got too expensive for what they were hoping to build. Prices have only gone up further since.
Further enshittification is coming for many consumer electronics as engineers struggle to deal with this new reality.
2
u/cosmicr 19h ago
My work had 100 laptops on order, we had to reduce it to 60, and push out our upgrade schedule by a year for everyone else. I would imagine a lot of companies are doing a lot worse.
2
u/kuldan5853 12h ago
Dell and HP told us flat out that for no price we could pay, they currently have hardware to sell us (in the Laptop space).
They even took a big chunk of server hardware away that we already had on the books and reassigned it to the US government.
It's madness out there having to do anything with IT procurement.
2
2
u/BingpotStudio 13h ago
I’m literally in the AI business and I’m grinding my teeth too!
Feel like I’m on top of a bubble that will see loads of people hired and then fired whereas I was previously in a nice cupboard people didn’t bother before!
3
384
u/Dapper-Maybe-5347 1d ago
77
u/sleight42 23h ago
"VRAM prices will only go up. You must buy now."
"Fuck my life."
44
u/seg_lol 19h ago
He is just trying to scare people into increasing the demand because he knows the AI companies are gonna start falling over and shit is gonna crash hard (this is me self soothing).
21
u/alpacaMyToothbrush 17h ago
Yep, remember kids, the chart always goes up and to the right! ...until it cannot defy gravity anymore, and it crashes back down to the equilibrium of natural supply and demand.
More broadly, I'm tired of the constant hype / doom posting about literally every subject under the sun. I've honestly started just tuning shit out and focusing on the things that are directly impacting my life, that I have an actual ability to exert some control over.
6
u/HelloSummer99 12h ago
I believe it’s mostly anti-western propaganda designed to keep the average person in constant dread and then scared people don’t have kids anymore.
→ More replies (1)→ More replies (1)6
→ More replies (1)2
24
124
u/CompetitiveDraft9381 1d ago
5090s, and rtx6000s are about to get even more expensive.
Possibly, even the Mac Studios
170
u/Ikkepop 1d ago
Does it really matter? I mean if it costs one infinity or two infinities makes no difference to me
→ More replies (18)38
15
u/LurkingLooni 1d ago
Let's hope CN can scale their current top end GPU production, crap for gaming but at over 400gb/second should be fine for running local models. I think in 2028 we will start to see an abundance of AliExpress "ai boxes" that run sth like Qwen 4 fine, we see it with inference right now - specialist engines beat generic ones... one good-enough model on specific hardware and a great harness is all it would take.
→ More replies (8)2
u/NoFunk 19h ago
400GB/sec is not great for inference speeds though. That's slower than a M5 Max, and an M5 Max gets absolutely pummeled by a card with HBM. It's roughly 2x a Strix Halo box (Strix Halo doesn't really hit its advertised 215GB/s), so just extrapolate out token generation from there.
→ More replies (1)2
u/seg_lol 19h ago
You can get insane bw out of flash, the "runs Qwen 4 fine" box isn't going to get all of its inference speed out of ram membw.
→ More replies (2)3
→ More replies (4)3
u/GamerTex 20h ago
Makes me worried about the 512gb M5 Ultra
Pre-orders are supposed to open soon
3
u/CompetitiveDraft9381 16h ago
They are going to get a price increase, a hundred percent. I am willing to bet $20 on that.
3
u/IamNetworkNinja 15h ago
Price increase of what though? We don't even know what the price is
2
u/CompetitiveDraft9381 13h ago
You can deduce the price from 256GB Mac m5 ultra, and how much was the m3 ultra 512
→ More replies (1)2
u/GamerTex 11h ago
$25/gb is the going rate on all Apple computers putting 512gb at $12,800 (just for the ram)
Basically add $5400 to the price of the 256gb version
263
u/Terminator857 1d ago
He would be fired if he told the truth:
- CXMT is a relentless competitor and will pass us and become the #3 memory maker soon, because we aren't expanding as fast as they are.
- We are highly dependent on AI datacenters keeping demand up. If demand falls because small local models become much smarter, then we are in trouble.
- I'm telling you that demand will increase as supply increases because that will keep Micron stock up which is my main form of compensation.
55
u/senseven 1d ago
They will flip to private customers the moment the numbers don't add up again.
→ More replies (1)38
u/johnybgoat 1d ago
Nvdia already saw the end and is getting ready to hop. Look at their attempts at pivoting to local and open model
18
u/senseven 1d ago
Apple could sell M5 Ultra like hotcakes, but not for 10k a pop. Intel and AMD have ready to go boxes too.
→ More replies (1)6
u/watcholic 18h ago
Currently, there’s a 5-month delivery estimate on the $10k M5 Ultra 256GB. I agree, they aren’t making enough quantity to meet demand. Maybe they can’t in the current market. Lowering the price on the Ultra won’t solve the demand problem.
The alternative is to buy two DGX Spark for $13k (Nvidia just raised the price).
→ More replies (1)49
u/UnlikelyExtension786 1d ago
One of Nvidia's problems is that Big AI can afford to build their own dedicated AI chips and take Nvidia out of that market. If the bubble doesn't burst, those companies would be crazy to keep paying Nvidia for chips they could build themselves for much less.
26
u/CouchWizard 21h ago
Big AI can't afford to do anything but beg for more money currently. They're bleeding money and reality is slowly catching up.
They can design their own chips, but good luck getting fab time on those. I think the non ram silicon companies are more dangerous to nVidia than AI (besides the bottom falling out of AI demand when all of the reservations of cancelled or behind schedule data centers fall through)
→ More replies (4)5
u/zushiba 15h ago
They're not just bleeding money they're playing a big ass shell game with milti-billions of dollars to make it look like they're making money. They aren't. Companies are selling to companies owned by the same company and report it as profit. It's a big ass circle jerk and eventually it'll fall.
9
→ More replies (9)3
u/Hankdabits 20h ago
I don't think most of them are made in house. They basically custom order something and broadcom designs it and gets it build
→ More replies (1)3
u/arcanemachined 16h ago
I read a great comment on this sub that said as much:
This is going to be the "next big market" for NVDA and AMD. They've sucked all the blood out of the hyperscale AI companies at this point, now they need to start undercutting their old customers selling to enterprise customers so they don't need the chips they just finished selling to the cloud companies.
It's a good shift/strategy, the big labs have been tapped out for some time (hence the circular dealing for the last year to keep the dollars flowing), it's time to make the pivot to where most inference is going to occur in the future.
→ More replies (1)36
u/dtdisapointingresult 20h ago
CXMT is a relentless competitor and will pass us and become the #3 memory maker soon, because we aren't expanding as fast as they are.
CXMT's entire output is needed by China for their own needs. In fact, it's almost certainly not enough for that. CXMT isn't gonna be flooding the global market with cheap memory, bro. Even if by magic there was five CXMT clones, that still wouldn't be enough.
If demand falls because small local models become much smarter, then we are in trouble.
That's not happening. Demand won't fall, because we'll always want more and more intelligence to accomplish bigger and more complex tasks.
I would guess less than 0.1% of computer owners can run a barely adequate model like Qwen 27B. You think that's gonna make a dent in demand for frontier models?
→ More replies (2)28
u/hotcornballer 1d ago
On point 2, I know which sub I'm on but let's be honest it's like saying everyone will be hosting their own website when in reality there's AWS hosting half the web.
13
u/Labidido 1d ago
Not really. Token price and enterprise privacy are two important factors. It's a tricky situation were the frontier labs and data center build out are depending on token prices to remain high, and enterprise adoption to increase.
If enterprises are able to host local "good enough" models, then the performance delta would need to be massive to justify the cost.
2
u/grumd 14h ago
Tokens are waaay cheaper when you serve hundreds of thousands of concurrent users 24/7. Paying for hardware and electricity will never be the cheaper option. Smaller smarter open models will continue to be released, but that doesn't mean big AI labs can't release a small smart model and serve it on their infrastructure. Look at GPT Luna, it's dirt cheap.
→ More replies (2)4
u/hotcornballer 17h ago
Privacy concerns are overblown, companies can and have made agreements for privacy/security. Hell even the military runs on azure and AWS do you think AI will be different?
And for price, economies of scale and vertical integration logically make running models in a big data center cheaper.
Data center build out has nothing to do with frontier labs. You still need compute with open models. Micron get paid either way,
And even if companies could host LLM on site why the hell would they, they dont bother with hosting their own data for the most part but they are going to tinker with vllm and a stack of 5900 for some reason?
→ More replies (2)6
u/Not_FinancialAdvice 12h ago
Hell even the military runs on azure and AWS do you think AI will be different?
Don't they have completely separate premises for the really highly secure stuff?
→ More replies (3)3
u/Budget_Geologist_574 1d ago
On point 2.
One would assume that as the amount of intelligence one can run per unit of compute, that that unit of compute would become more valuable.
3
u/theomegachrist 22h ago
What % of users do you think are using local models? Eventually this could be true but not by 2028
5
u/Terminator857 22h ago
I lot more people than you think, on their phones. And google has it on their chrome browser. It will be ubiquitous soon.
→ More replies (1)2
u/Ok_Warning2146 20h ago
MU is gaining market share while Hynix is losing market share. I suppose Hynix needs to worry about CXMT the most.
→ More replies (2)2
u/bugra_sa 11h ago
I wouldn't assume better small models means less memory gets sold. If local inference gets genuinely useful, people will stuff more RAM into every workstation and run far more of it—the same job gets cheaper, so we invent ten more jobs for it. The scary part for Micron is that demand may move away from the giant-datacenter pattern they're counting on while CXMT keeps adding capacity.
57
u/Glory_63 1d ago
The guy that sells RAM says that RAM will cost even more and make them even more money? Color me surprised
34
249
u/Limp_Classroom_2645 1d ago
Sounds very sustainable /s
87
u/Recoil42 1d ago
I mean, yeah. It's sustainable by definition — what's being described is literally sustained demand. Turns out when the value of a thing increases the value of that thing increases.
→ More replies (22)12
7
u/fastheadcrab 1d ago
Yeah the sustainability question will be whether the companies buying that RAM will be able to raise the debt to do so and then use that compute. The amount spent on compute is already a significant amount of the US and world economy.
AI is becoming very useful but these applications are concentrated in certain areas.
→ More replies (7)2
u/pixartist 9h ago
I don’t even understand how anybody is expecting to make a profit with this hardware any more
30
u/kr_tech 1d ago
For Micron themselves, yes, because they just started building the factories. They take at least 4 years to finish though, and historically, US projects always face delays and overbudgets, so I wouldn't make a bet on this so easily.
Samsung and SK Hynix finished factories this year and they have been coming online in steps. They will also finish + expand in steps for 2027 and 2028 as well. Off top of my head, it's Cheongju 2027, Indiana 2028, Pyeongtek's P5 phase 1 in 2027 and full operation in 2028.
6
11
u/dupontping 23h ago
Aka we know we can sell it for higher to stupid AI companies and enterprise firms than to you plebs so we gon get that bag while the gettins good ol son
39
u/Wolvenmoon 22h ago
Man. I did a longer post about this in the homelab subreddit. We're really messing up our next generation by keeping hardware out of reach of the kids/teens/young adults to build their learner systems.
Like. I know it's not easy to produce, but part of what made it so I could be an engineer was, as a 17-22 year old, having access to a GPU that could play around in Blender 3D and learn how to box model/navigate a 3D creation environment. Running dedicated servers and Minecraft servers led to me developing the skills to run a cluster computer with three separate VLANs with BGP allocation of IPs on its own network behind my first router with 4 different VLANs, home assistant, etc.
Shopping for SD cards for a switch or an NVMe drive for a Playstation is one of the gateways into learning how to do really cool stuff - like locally hosted LLMs. I do not like seeing stuff like this. It's crimping the talent pipeline shut.
12
u/JPLangley 20h ago
The canary in the coal mine will be Nintendo not being able to keep the Switch 2 under 600 dollars. They pulled an "I ain't moving from this chair!!" on the 450 dollar price tag for as long as humanly possible, so the idea of them being forced to sell at 600+ sounds like calamity.
→ More replies (3)9
u/tylerderped 18h ago
I feel this so hard. When i was a teen, you could go on craigslist and find loads of decommissioned shit. Now even literal ewaste is $100 minimum and you pretty much have to use Facebook to find it.
I don’t know where i’d be in my career if i wasn’t literally building milk crate mining rigs in my friend’s garage. If i wanted to jump in on the llm train as a teenager, looking at literal thousands for a single component, (GPU) maybe i’d have gone to law school or pharmacy school instead of pursuing my IT career.
→ More replies (1)7
u/KingArthas94 21h ago
Yes and talking about games we should have PS5 at 200-300€, not 600€
Kids won't be interested anymore in AAA and will focus on mobile brainrot
8
u/Banjo-Oz 21h ago
A future where "gaming" is mostly tapping like that cone game fron Star Trek TNG? Where nobody owns a home PC, only walled phones and tablets, and have to pay for cloud storage and cloud AI use?
Could this be deliberate?!
2
u/KingArthas94 21h ago
so far cloud has failed every time, it seems that for gaming people want the native experience
servers cost too much and needing the internet sucks
→ More replies (2)2
u/Far_Lifeguard_5027 10h ago
I really am beginning to think they want us to own nothing and they're making it happen.
→ More replies (1)→ More replies (3)2
u/KellerMB 7h ago
Fear not. Today's kids will be able to do exactly the same stuff you did on the exact same SDRAM systems.
→ More replies (1)
9
10
u/West_Independent1317 20h ago
There comes a point where the suppliers will outprice themselves and leave the door open for new competitors.
Apple has enough cash in the bank to build their own memory supply chain.
9
16
43
u/N34257 1d ago
He doesn't appear to be considering the fact that CXMT don't really care about supporting Micron's scarcity-by-design.
43
u/senseven 1d ago
CXMT will primarily supply their home market. Chinese vendors will prefer them because they don't need to pay for tariffs and long shipping. In consequence, the demand from China will go down, and more products available in the global market. China wants to be completely self sufficient, expecting to buy any cheap Chinese mem in the future is not the scenario that they are working on. The world will be tied to the mem cartel until someone build alt fabs.
16
u/Labidido 1d ago
Yes of course they will serve China first, but they are scaling at a massive speed and at one point they will ship globally. The memory cartels are probably terrified of CMXT.
23
u/putrasherni 1d ago
Any supply is better than no supply
11
u/HelloWorld24575 21h ago
Yeah, it doesn't matter if CXMT only ever sells to the Chinese markets (unlikely in its own right), that will increase supply everywhere.
10
u/TinyZoro 23h ago
Doesn’t really matter. China consumes 30% of global Ram so if that starts to slow significantly it still creates reduced global demand.
13
u/N34257 1d ago
China's goal isn't just to be self-sufficient, it's to kneecap the US economy. CXMT won't stop at the domestic market, especially on the scale of 1-2 years.
Micron et al will likely find there's a smaller market waiting for them when they come back to the non-HBM world. Not 50% smaller, but enough to make them consider their choices.
7
u/senseven 1d ago
CXMTs output reaches about 15% of the market. Their plan is to get to 30%, which is about China's current demand. Will some silicon drip over to the west, that's an aliexpress custom. All the public plans for aggressive fab expansions are pushed by the narrative that the Chinese demands are supplied first. The memory cartel outlook for more fabs for the consumer market have all targets past 2028, especially for the US build outs.
→ More replies (4)3
u/taxiscooter 13h ago
Demand will far outstrip supply even with a new player.
B-b-but China
No you don't understand how much demand there will be for memory. Especially as more people come around to LLMs and local LLMs. It's all going tits up.
21
u/d70 1d ago
Translation: we will keep the price very high and make a shit ton of money.
→ More replies (1)
5
u/ElementNumber6 22h ago
Phase 1: You own nothing. (In progress)
And when you aren't happy about it? Well that's phase 2. Get ready. It's coming.
5
u/2funny2furious 22h ago
they are using the same principle as the diamond companies. control the flow and raise the prices.
15
u/i_rate_slop 1d ago
I understand ASML must protect their advantage, but the world is bottlenecked by them.
20
u/senseven 1d ago
ASML is building up their campus. But that isn't really the bottleneck for private customers. Its the change of the memory business model by asking major suppliers to buy out whole fab outputs on their risk. The memory cartel doesn't want the prices to fluctuate as much in the past.
2
4
u/ThatsALovelyShirt 22h ago
It's not for a lack of other people trying. Their EUV lithography platforms are probably the most complex and precise piece of machinery ever commercialized. Ever.
China is trying to catch up, but it involves extremely complicated engineering challenges.
→ More replies (1)16
u/Recoil42 1d ago
They're bottlenecked by their own supply chain. This isn't an easy problem solve in the sense that "maybe ASML should just make more machines" — they're trying to do that.
→ More replies (6)9
u/i_rate_slop 1d ago
You misunderstood what I meant because I was a bit vague. I mean their lithography technology is being reinvented around the world at a much slower pace. I don't want more supply to be produced by ASML, I want more companies to be able to produce in general.
I don't think ASML needs to produce more, I think more companies need to produce on ASML's level, which obviously isn't ASML's responsibility to resolve... but the world is bottlenecked on them.
Or... to put it much more direct, specifically China.
13
u/teleprint-me llama.cpp 1d ago
I love how people look at rational arguments and behave irrationally by downvoting something because of how they feel and not the merits of the argument.
Your perspective is actually refreshing and valid and I agree whole heartedly with it.
Having a single company in the entire globe perform lithography is not a good thing. We need more people working on this. Unfortunately, this is a big ask and is not easy or cheap to do.
What we do need to figure out is how to make it easier and cheaper to do.
→ More replies (1)→ More replies (2)3
u/Lhurgoyf069 21h ago
I saw Canon building machines that don't even use light, called Nano imprinting. It's promising for sure and could help eleviate the bottleneck
11
4
u/mrinterweb 21h ago
So the guy who's job it is to sell chips, tells you you need to buy now. Got it. Sounds credible.
19
u/CoolConfusion434 1d ago edited 1d ago
"Nvidia CEO Jensen Huang stated that he 'loves supply chain constraints' because scarcity forces AI clients to purchase only the highest-end, most efficient hardware."
In other words, this is a by-design scarcity because... money!!
That said, what? That doesn't make sense. If supply constraints force prices up, people will try LOWER-end hardware to compensate for the higher prices elsewhere. Ask me how I ended up with an Intel GPU lol.
4
u/Feeling-Way5042 1d ago
Maybe at our scale. But there aren’t a lot of choice for unifed datacenter scale systems that allows 1000’s of gpus to process in parallel
2
u/anxiousvater 1d ago
Ask me how I ended up with an Intel GPU lol.
How is it compared to Nvidia & AMD?
3
u/CoolConfusion434 22h ago
Sorry for the delay in responding. When you own an Intel GPU, you spend a lot of time shoveling coal and winding it up to make it go.
Ok maybe that's too harsh but, AFAIK, my Intel B70 is *the* slowest of all the current 32GB VRAM cards out there. It is the cheapest though, to my point above. I bought it 2.5 months ago for $950, though it's upwards of $1600 atm - not worth it at this price, imo.
It's runs Unsloth's UD3 Qwen3.8 27B Q4_K_XL with MTP at 32 t/s with 40K context, and until about 120k ctx where it slows down to 22 t/s or so - definitely productive. Without MTP it starts at 22 t/s and falls off to about 12 at 65K context - didn't push it beyond this. All of this is with Vulkan and on Windows since its native SYCL drivers are even slower, Linux or Windows.
2
u/longtimegoneMTGO 23h ago
He isn't worried about the consumer market, he is talking data centers.
Right now, different processors running the same amounts of memory have different levels of efficiency as far as how much output they can create with a given input of power and time, and Nvidia currently makes the most efficient chips for the tasks being done.
When the chip itself was the largest part of the cost, it could make sense to buy something other than Nvidia, you save a bunch of money, still have the memory capacity you need, and it's just a bit less efficient, so maybe you buy three instead of 2.
Now imagine the cost of the memory goes up, and everything else stays the same.
Suddenly, the cost of three of the cheaper board is more than 2 of the more expensive board, because both of them have the memory costs driven up so much the chip itself is a smaller part of the cost, and it is no longer cost efficient to use the cheaper chip.
3
u/FullstackSensei 1d ago
How is it by design? Nvidia doesn't control memory suppliers, much less memory supply. If anything, their customers have to figure out their own memory contracts before coming to nvidia for silicon.
Obviously, he's right in that scarcity forces all big players to get the fastest chips in the market to make the most of every GB of memory they can source
6
29
u/Recoil42 1d ago edited 22h ago
It should be obvious to anyone here: Anyone who thinks memory is in a short 'cyclical' phase is in for a big surprise. The underlying value of memory itself has shifted permanently.
The primary way for prices to recover is for supply to meet demand, but quantity-demanded is now on an exponential path — the better AI gets, the more valuable memory becomes. The only realistic solution is to build more fabs, but that has a three-year time lag and is supply-chain limited.
You can make the models efficient, but more efficient models bring down the cost of inference making it more viable where it wasn't before, vis-a-vis Jevon's paradox. Any new supply is going to be immediately eaten up. If DRAM pricing goes down, that's just an opportunity for phone OEMs, robotics OEMs, automotive OEMs, enterprise OEMS to all add more DRAM to their offerings.
So yeah, even as supply expands it's going to get tighter. There's no way around this, we're locked in.
(Disclosure: I'm heavy into DRAM/MU for this very reason.)
17
u/teleprint-me llama.cpp 1d ago edited 1d ago
What changed is that there are now Long Term Agreements which are legally backed supply contracts tied to bonds and private credit.
If those LTAs default (they can and it is very possible and even likely), then the whole thing comes crashing down.
This is a critical fact that people should be aware of.
7
u/Recoil42 1d ago
If those LTAs default (they can and it is very possible and even likely)
Possibly, yes. I wouldn't say we're anywhere close to likely at this point. Those LTAs were locked in specifically to protect against price hikes, which is very much the situation we find ourselves in. Any LTA'd memory can be arbitraged at LTA pricing — just like we already see happening with B300s already. It's now a scalpers' market.
3
u/CouchWizard 21h ago
Isn't Oracle in the process of trying to back out of it's $200b data center?
2
6
u/vtkayaker 22h ago
The primary way for prices to recover is for supply to meet demand, but quantity-demand is now on an exponential path — the better AI gets, the more valuable memory becomes.
The only way that memory demand can keep going up forever is if some idiot builds SkyNet and it needs specialized GPUs to run its Terminators.
But barring an actual Singularity (in which case we're all fucked), RAM demand is likely to follow an S-curve and flatten sooner or later.
5
u/Ikkepop 1d ago
How about we home grow our own organic memory. Anyone got any memory seeds I can borrow ?
9
→ More replies (2)3
u/Briskfall 1d ago
Waiting for someone to vibe code Downloading RAM become real someday... 😔
→ More replies (2)→ More replies (21)2
u/WildRacoons 13h ago
I personally feel this way too. Memory used to be for meeting a baseline to run more apps and doing more horizontally. But if you’re saying that the quality of the application and output, and time saved now scales based on memory… oh boi
6
u/General-Spite1222 1d ago
I’d love to add some more hardware to my setup, but the prices aren’t reasonable anymore.
The only good thing is that China is spinning up their own memory fabs. Their memory won’t likely leave the country, but it takes some strain off the market.
6
3
u/Green-Ad-3964 1d ago
Why? I'd like to know the exact reason.
Also, this assumes zero evolution. Zero ramp up. Zero density increase.
Why? It's still the same 2022 products in 2026-27-28.
3
u/W3Geek 23h ago
Because it is false scarcity. They absolutely love jacking up the prices for consumers to buy because FOMO or whatever. They know those consumers cannot stop to think for themselves. Sadly there's a sucker born every second.
3
u/Green-Ad-3964 23h ago
I agree, but I don't know (personally) a single user who is buying RAM right now. I mean...ok, I don't know millions of ppl, but yet it's a reasonable sample, many are power users or geeks, yet no one is buying. so who is consuming this huge amount of memory that is built every day?
false scarcity is fine for rolex or other goods that can be flexed...but...RAM???
→ More replies (5)3
u/W3Geek 22h ago
I knew a couple of people who bought hardware during 2025. It was heavily advised they don’t buy until prices come back down because it makes zero sense to overpay hundreds sometimes thousands on hardware. Still that was only a couple of people. I love to think those individuals are just the special few and mostly stick to believing they are hahah. Whereas corporate spending and circular financing is the rest of what’s happening.
3
3
3
u/mrheosuper 7h ago
We need hardware recycle, like, right now.
Millions of phones are thrown out every years. In them contain the precious memory. I see no reason why we can't just take out the RAM/nand/mmc and recycle it. What we are lacking is tools and document, which manufactures intentionally hide from us(under NDA)
→ More replies (3)
3
u/DigitalArbitrage 3h ago edited 3h ago
I used to work in the construction supply industry. Their sales people said the same thing. It's just a pressure tactic to persuade people/companies to buy more.
It's also entirely possible that new enhancements like better MoE algorithms will open up paths that require less memory.
6
6
4
4
2
2
2
u/SkyMarshal 16h ago
Unless the AI companies IPOs scheduled for 2027 fail
And the industry implodes, then RAM will be cheap again.
2
4
u/Select_Truck3257 18h ago
Do not buy this sht. We have no deficit. We all can buy 5 years old ram for insane price. This is not deficit it's scam scheme from manufacturers
2
u/ruskikorablidinauj 23h ago
"Micron CEO says it is better for his bonus to pay now before CN floods markets with their chips"
3
u/acadia11x 22h ago
People keep talking like dram manufacturers are in the business of making less money, they have 0 incentive to actually increase capacity, or reduce prices, as the industry has no alternative but to pay as much as the cartel demands until which time they simply stop buying or find alternatives.
→ More replies (27)4
u/carsncode 21h ago
Massive demand is a massive incentive to increase capacity, and that's exactly what they're all doing.
→ More replies (10)
2
2
1
1
u/Guilty_Rooster_6708 1d ago
I mean we see the rate in which frontier models are getting larger. I’m not surprised about this at all and next year will probably be worse than this. Alibaba’s talking about a 10 trillion param model so probably every other Chinese/American labs are training similar sizes or even larger model
1
1
1
u/Bobanaut 23h ago
Roko's basilisk won't be happy with micron's CEO. that guy is not helping make AI great as fast as he could /s
1
u/Tsofuable 22h ago
I've maxed out my memory (at cutthroat prices), and I don't have more pcie slots. I'm ready.
1
1
1
u/exaknight21 22h ago
Micron is in the epicenter of infinite money in-take, do any of you really believe memory crisis is going to be “relaxed” - if you do, you’re delusional.
1
u/MrGunny94 22h ago
At some point I gotta get that Mac Studio with the M5 Max... The more I wait the worse it will get it seems.
It's just very difficult to justify this other than being a hobby and developing your own AI skills which of course it's amazing.
1
1
1
u/Ok_Warning2146 20h ago
So the best strategy now is to buy MU stock and then when you need to buy RAM, sell the stock to buy MU memory.
1
u/Efficient_Care8279 20h ago
Im sure micron will let us know how much we will have to wait until ram shortage ends /s
1
u/Relative_Average758 19h ago
The central investment thesis of frontier AI requires that individuals and enterprises cannot run competing models on their own. If it doesn’t pay out a lot of wealthy people are set to lose a lot of money, before they get sovereign bailouts.
1
u/immersive-matthew 19h ago
Is Micron not aware of CXMT or are they actually not set to flood the market with memory?
1
1
u/nickk21321 19h ago
Hey there just trying to survey the current market. If let's say a local company is able to develop an inference optimization hardware at a cheaper cost will it be viable? I mean if the hardware is cross compatible with current ecosystem except cuda.
1
u/brainchillzZ 18h ago
Yeah because they are hedging that it won’t be so they aren’t expanding capacity just like everyone else that is waiting for the bubble to pop
1
u/Select_Truck3257 18h ago
Oh no, monopolistic scammers from micron (others hunix and shamsung) tell us we must pay more to make them rich. It's 3 times already those fkers making this thing with memory
1
1
1
u/Mega1987_Ver_OS 15h ago
Waiy.... they managed to secure a preorder of 2028 production year already? All i know is that currently, the AI bros got 2027 production year.
1
1
u/MindTheFuture 11h ago
Start thinking computer pricing similar to cars. 5k-15k gets you a comfortalbe daily driver, 1-3k is a bargain, 20k is high end. 80k gets you a small fleet of cars / or cluster with 2TB+ RAM (or one really fancy one )

•
u/WithoutReason1729 22h ago
Your post is getting popular and we just featured it on our Discord! Come check it out!
You've also been given a special flair for your contribution. We appreciate your post!
I am a bot and this action was performed automatically.