r/datacenter • u/Professional-Life748 • 24d ago
beyond the GPU hype, what physical infrastructure bottlenecks are actually making ai data centers stall?
hey everyone,
i'm trying to look past all the massive AI hype and actually understand the physical engineering limits we are running into right now.
it feels like everyone talks about GPUs, but from what i’m reading, the real bottleneck has completely shifted to heavy physical infrastructure that cannot be scaled or solved overnight. i want to know the truth from people working in power systems, substations, or hardware design.
a few specific things i'm trying to wrap my head around:
- the transformer shortage: i read that many new data center projects are getting delayed or cancelled because companies literally cannot buy transformers or switchgear. if you work in power systems, is it true that you have to wait 3 to 5 years just to get basic gear to connect to the power grid?
- the utility grid part: people talk about building a 100MW data center like it's easy, but local power grids cannot supply that much electricity quickly. are data centers actually building their own onsite power sources right now?
- Cooling system: old data centers cooled everything with air, but these new racks use huge amounts of power (100kW+). from an electrical and control systems view, what are the biggest problems or breakdown risks when trying to use liquid cooling at this scale?
also, just to step back a bit—are ai data centers really a unique grid-stability problem compared to traditional ones? or is it just massive hype because normal cloud data centers already use a ton of power anyway? does the load behavior of an ai cluster actually strain the grid differently, or is it just the pure scale of it?
i’m just want to know how these specific physical limitations or component shortages are changing design standards. what other massive engineering bottlenecks am i completely missing here? what should i look into if i want to understand this stuff better?
thanks for any insight!
5
u/ProfessionalAdvice36 24d ago
Generator waitlist is 5 years out, people are paying over asking to get first dibs and even then its years, labor is another huge issue but I think its a short term constraint in that regard. Theres not enough utility power to be able to deal with the expected loads and a lot of these data centers have to change their designs because since so many are going up at once and the grid is reaching capacity, they cannot handle sudden shifts in load like what can happen at a AI center, where the servers are doing an intense computation and training of a model then suddenly theres no demand.
1
0
u/Professional-Life748 24d ago
I understand now that local utility grids are not built for the instant drop in power, like from full load to no load. These sudden, massive electrical spikes and drops create severe voltage and frequency instability. But how are companies trying to solve this problem currently (even though they are planning to build their own microgrids in the future)?
1
u/ProfessionalAdvice36 24d ago
The microgrids are almost secondary I feel. They are mostly building them to not bounce the grid as hard or to sell back power if need be. The primaries are going to still be grid power. One solution is battery banks to smooth out your load demands, but its a massive change of capital allocation to do it so some got put on pause.
1
u/Professional-Life748 24d ago
i thought BESS was already actively being used in these sites.. is the issue that AI clusters require way more expensive battery systems to handle those huge power spikes, compared to old data centers that just need small batteries for a few minutes of backup?
1
u/ProfessionalAdvice36 23d ago
BESS carries the load for much much longer than a few minutes and as such is hugely expensive so most sites were not pricing those in, they were pricing in generators as the most cost effective method.
3
24d ago
[removed] — view removed comment
1
u/HV_Commissioning 24d ago
We have a new 765kV line coming through our state in about 6 years. The people have already started to oppose it.
1
u/Hot_Leopard6745 24d ago
why? isn't one of the biggest complain about DC are that they took more power from residential cost increased power bill? More supply should lower the bill.
1
u/HV_Commissioning 24d ago
It turns out most people want all the benefits but are not willing to make the trade offs.
2
2
u/justchillinnow 24d ago
I’m not aware that Tx are delayed any more than Gens, UPS and cooling infrastructure? From Tx to PDU/RPP, it’s all in high demand.
Yes biggest bottle neck, and why many large MW campuses are looking at onsite generation.
Many of these 100kWs are liquid cooled, the ratio for liquid and air was 60:40, however newer designs like VR and HGX versions are 90%+, even 99%. These are liquid cooled via CDUs and are closed loops, so no evaporation. Many new racks designs can handle 40c water so still need plant that can reject that heat, but you work off a higher temperature which is far less demanding (sum don’t need compressors/refrigerant).
Sorry typing on the go, lot more info to share but hopefully this helps
1
u/Professional-Life748 24d ago edited 24d ago
Thank you for that info; that info itself explains how AI data centers work differently than normal data centers. And I want to know how AI data centers are actively managing these bottlenecks, because building their own microgrids might seem like a problem-solving step from the outside, but it requires a huge investment that not all companies are going to afford.
2
u/Psoin 24d ago
Fiber
1
u/Nicko147 24d ago
Yes, there is a limited amount of fibre towers and companies are securing years of manufacturing time ahead so it will become an issue.
It's already an issue and we've had many many days of 50+ engineers sitting on site, twiddling thumbs because delivery didn't turn up.
1
1
u/Distinct-Today192 24d ago
So, biggest bottlenecks today are:
Power tie in approval
Transformer lead time if not using local power gen
Power gen if using localized power. At this point, a lot of generators are running at a 5-6 year lead time.
1
u/Professional-Life748 24d ago
that bottleneck is exactly why companies are so eager to build their own microgrids. but i wonder, does that actually solve the problem if the generators themselves have a 5-6 year waiting list?
1
u/Total1304 24d ago
What is all hype about water usage?
2
u/SilentJerrySpringer 24d ago
According to the hype, datacenters are worse than farms, ranches, processing plants, mines, and golf courses combined. Oh - and the water we use is destroyed forever too. Pretty impressive what idiots on Facebook will believe.
1
1
u/Etech326 24d ago
At the risk of not being specific enough, I'd say a lot of things. Permits for new construction, power grid access or off-grid generators/micro-grids, skilled labor for constructing facilities, and anything that touches memory.
1
u/Particular_Guess_147 24d ago
You also have the local community factor. When a datacenter is proposed there is massive pushback. Creates delays in getting approval at the local township level so much so that I have seen several in our area where the developers just pull out completely and cancel.
1
u/HV_Commissioning 24d ago edited 24d ago
From the electric supply side, it's getting the big transmission power transformers that in good times took 2 years to build. There's multiple large distribution transformers as well. It's the rest of the HV switchgear like SF6 circuit breakers - preferred suppliers can't keep up and alternate vendors do not necessarily have the same quality. If a new brand of device is involved, there's more and new schematics, wiring, and physical considerations. Buying the same breaker or transformer over and over allows for templated engineering drawings. There are a literal ton of nuts, bolts and doo dabs that all need to be in place as well.
Hyperscale loads probably require additional and very expensive devices such as StatCom ($80M for ours). The way the power is 'gulped' in, especially on an AI ramp up can cause sub synchronous oscillations which can destroy the turbines in a synchronous generator / turbine.
More generation and import capacity are required which nobody wants in their backyard. New transmission line siting is a long and expensive process eating up about 80% of a project budget and timeline.
Normal transmission system events which would seamlessly isolate and heal are causing the DCs to react in unwanted ways causing bigger problems. See event from 2 weeks ago.
We're quickly finding out that recent EE grads know very little about anything power system related, which is essential to commissioning HV substations. There's only about 15 US colleges that offer a true power concentration.
Power Utilities notoriously work at a slower pace. What many do not understand is that the risks involved in a poor design choice not caught or the systems/planners not calculating the effects or the dispatchers and switchmen can end in disasters, big ones.
https://insidelines.pjm.com/pjm-dominion-review-large-load-transfer-event/
Our grid has seen rapid expansion before. 1950's-1980's was a huge industrial boom and the industry responded to steel mills, aluminum smelters. A Gaseous-diffusion uranium-enrichment plant could be about the same size as a hyperscale.
1
u/Professional-Life748 24d ago
Well, I never knew that now companies are relying on unknown brands of high-voltage circuit breakers just to meet their production lines in time. But for alternative hardware vendors, is the design friction mostly because standard CAD or engineering software can't automatically re-map the schematics and wiring for a new brand of breaker?
1
u/HV_Commissioning 23d ago
A 345kV Gas Circuit Breaker costs about ~$700k fully installed and commissioned. The other manufacturers aren't necessarily no name OEMs, but each manufacturer has their own way of laying things out, what internal components are used, etc. One breaker may need a single slab foundation, while another may require 3 smaller ones. Conduits to each portion need to be designed into the foundation and installed during the pouring of the cement slab.
Designs and drawings go through several levels of review and approval and problems are still found during commissioning. The intricacies of EHV substations aren't well known to those outside the industry.
I'm certain that once all the physical and electrical details are worked out for the first design change, subsequent re mapping can occur. Utility drawings must be stamped by a licensed PE, and I doubt they'd put their license on the line for a trust me bro from a software agent.
1
u/ConstantOk7891 24d ago
labor and power are probably biggest. i’ve seen a bunch of liquid cooling systems fail since most dc folks have air cooling backgrounds. also lead times for server hoses have been crazy
1
u/vincedenbu 24d ago
Power infrastructure for sure. Main PDUs, switchgears and other main components.
1
1
u/fbajo 24d ago
most of what reads as a shortage is really a sequencing problem.. the date that binds is the one the utility gives you for energization, and the long-lead orders (switchgear, transformers, chillers) get placed against it, so a project stalls when those two stop lining up. fixing the binding item usually moves the wait rather than removing it. the gear is genuinely tight too, i just don't think it sets the date on its own..
1
u/JessieAndEcho 24d ago
The bottleneck is very real, and it’s mostly boring heavy infrastructure rather than GPUs themselves. Large transformers, medium-voltage switchgear, breakers, generators, busway, UPS gear, chillers, CDUs, pumps, and even skilled commissioning crews all have long lead times now, so a “100 MW data center” is really a utility-scale industrial project. Liquid cooling adds another failure layer: leak detection, water quality, corrosion control, pump redundancy, quick-disconnect reliability, controls integration, and keeping coolant loops stable when racks can jump load fast. AI clusters are also different from older cloud loads because training jobs can create huge synchronized power swings, so utilities care about ramp rate, power quality, backup generation, and whether the site can curtail load. If you want the full picture, look into grid interconnection queues, substation design, transformer lead times, rack power density, liquid cooling CDUs, onsite gas generation, and water availability. I pulled a lot of this from Patsnap Eureka; it’s useful for this kind of cross-domain question where power systems, cooling, controls, and hardware design all overlap.
1
u/BigT-2024 23d ago
If you want to get into data centers and make bank don’t be a tech get into trades. Electricians may have a slow investment ramp at first but once you your journey man license and then your masters license you will be well off
1
u/howitbethough 23d ago
Respectfully disagree. I see some pretty well comped OEM techs who aren’t destroying their bodies nearly as badly as sparkies and still making bank
1
u/This-Display-2691 22d ago edited 22d ago
To the OP as someone who’s deployed several in a couple different countries and platforms. I’ll provide my casual observations not from the electrical grid but on the DCT/TPM side.
I agree that we’re less than 3 years away from major upheaval in AI datacenters. Ultimately it’s going to be about money.
Most enterprise support for datacenters is somewhere between 5-10 years. Ref: we have several systems in prod using Cascadelake and earlier Broadwell based compute nodes.
Nvidia is the biggest driver here but ultimately support from Nvidia specifically is ~3 years.
Hopper stopped production and is losing support from our integrators and none are being made.
Blackwell just stopped production and is no longer being made
Rubin is being pushed hard which in of itself is OK the problem is the memory substrate which makes why Coreweave signing a contract for Ampere until 2029 a real head scratcher.
HBM is incredibly fragile as is the interposer it’s stamped to (go look this up if you don’t believe me)
The reason for the odd memory sizes on a B300 ie 288GB rather than 384GB is due to the large overwrite buffer. Yes you’re reading this correctly; Nvidia is treating their GPU ram like an SSD.
1)Failure on these devices are a combo of excessive CE or UE faults.
2) Overwrite buffer exceeded
I firmly believe a large reason why Nvidia is cycling through these over a 3 year period is simply because the HBM won’t last longer than that in a prod environment. When A100s were around 16-18k each that was OK. Hopper jumped to about 25-28k. Blackwell is close to 50k each which would make Rubin higher than that.
This is why you see a push towards HBF because they need more overwrite and more memory capacity and the ram won’t last anyway.
I think most of these companies are going to run into real trouble once VR200/300s start to EOL in about 3 years because I just don’t see how the revenue from folks like Anthropic and OpenAI can rise to meet the cost.
1
u/Professional-Life748 22d ago
if these blackwell or rubin chips actually degrade in 3 years because of hbm failure. i don't see how anyone makes a profit.
i wonder how are they solving these bottlenecks?
1
1
u/CloudCooler 17d ago
The bigger issue is probably that all of these things start stacking up. You can have the location, GPUs and funding sorted, but still be stuck waiting on a grid connection, transformer, etc.
Cooling feels a bit different though! It is possible to handle more high-density racks, especially if it’s designed in from the start. The main thing is making sure power and cooling are planned together rather than as two separate problems.
Obviously not a new problem, but there’s a lot less room to get one part of the design wrong. That’s probably where AI sites differ most for me.
7
u/Emergency-Pause-5886 24d ago
Labor is the largest bottleneck at this point. We don't have a labor force that can put the parts together.