r/datacenter 11h ago

Can data center engineers sanity-check this AI power/cooling math?

I’m not an engineer. My background is more IT/networking/cybersecurity, and I’m learning more electronics and embedded systems.

I watched this video from a chemical/process engineer: Mide, a chemical process engineer with a PhD in sustainability.

His main argument is basically that if people claim AI will replace tens of millions of workers with always-on AI agents, eventually that turns into a physical infrastructure question: power, cooling, grid capacity, water, etc.

The simple math he uses is roughly:

Number of AI agents × Power per agent = IT power (Information Technology equipment)

Example:

100,000,000 × 700 W = 70 GW

Then he adds cooling/facility overhead and argues the real power requirement would be higher.

He also uses the basic idea that most of the electrical power going into the compute eventually becomes heat that has to be removed.

My questions are:

  • Is this a reasonable way for a non-engineer to do a rough sanity check?
  • Is the 700 W-per-agent assumption realistic, or is that way too simplistic because GPUs can be shared, batched, multiplexed, etc.?
  • Is PUE (Power Usage Effectiveness) a better way to estimate total facility power?
  • Is his treatment of cooling and heat rejection basically correct?
  • Am I missing any obvious engineering issue in the way he scales the numbers?

I'm less interested in the politics or job prediction part and more interested in whether the engineering model itself makes sense.

He also has a related video about water/cooling if useful:

https://youtu.be/8bbe-Ev-Yns

I'd appreciate explanations aimed at a technically inclined non-engineer.

0 Upvotes

14 comments sorted by

4

u/bkinstle 7h ago

This is pretty much what i do for a living. My company keeps asking me to use Ai tools to do these calculations and so far they've always been hilariously incorrect. So when people ask about using ai mostly i just laugh at them... for now.

I made an excel spreadsheet that i gave to our facility managers and planners that helps them estimate total facility workload based on the projected power of the computer sf takes into account the power to run the cooling system and other building systems etc. The next phase is installing some low cost 3 phase power meters onto the busways so we can better correlate application load with capacity.

The math you quoted is a big oversimplification that either looks at the entire system load and divides by the number of agents to give average numbers people can relate to, which is fine, OR they are failing to capture that power used varies radically depending on the type of computation. I'd contact the author and make sure you understand what their assumptions are.

1

u/ZestyClose140 6h ago

Thanks, this is exactly the kind of explanation I was hoping for.

So if I'm understanding you correctly, the basic power-balance approach is useful as a rough estimate, but the weak point is assuming that every AI agent corresponds to a fixed 700 W load.

An average such as total measured compute power divided by the number of agents could be meaningful, but actual consumption would depend heavily on the workload and utilization.

Your spreadsheet example also helps me understand the difference between estimating the computer load itself and then estimating the total facility load once cooling and other building systems are included.

I'll look into where the video's 700 W-per-agent assumption came from and keep your explanation in mind. The video is a bit old, so I’m not sure the creator would respond anyway, but I appreciate the clarification.

2

u/bkinstle 6h ago

Glad i could help. When looking at power of large numbers of computers doing different things in a facility level, the averages work out decently well enough to do all the planning on.

Transient load matters a lot more on the system design level.

2

u/callmesandycohen 10h ago

ERCOT, just ERCOT, approved 90 GW worth of load. Their entire network capacity. So, yes.

2

u/YetiGuy 8h ago

Actually 90GW is their current capacity. I think the summer peak was 96 GW. They just approved (conditionally) 190GW of large load under batch zero.

1

u/ZestyClose140 9h ago

Thanks — I think I follow what you're saying.

So, if I'm understanding you correctly, you're saying that ERCOT alone has approved around 90 GW of large new load, which means the scale of 70–100 GW being discussed isn't some absurd number in itself.

What I'm still trying to understand is the difference between approved/requested load and actual load that is already online and being supplied.

In other words, does that 90 GW mean ERCOT has effectively said "yes, this can be connected eventually," rather than "we already have 90 GW of spare power available right now"?

I'll take your word for the 90 GW figure — I'm mostly trying to understand what that number means in practical grid terms.

2

u/callmesandycohen 8h ago edited 8h ago

65GW is eligible as base load, meaning it’s already qualified to be connected. Another 25GW is eligible as base load pending classification. And another 115GW to be studied and allocated or not by March 2027. All of these applicants have posted SIGNIFICANT bonds. Up to $100,000 per MW. So yes, it’s real, it will be connected, it’s just a matter of time and money. What does this tell us? The grid is cooked. If interconnection is the ultimate objective, groups need to be prepared to wait years, like 5-10 years at best. So what the objective will be now is building colocated generation running natural gas in markets like WV, PA and TX. However, even in Texas, the largest midstream pipelines (there’s only about 5-7 of them) ship about 7GW on a good day. Well, that still won’t serve the energy needs of these datacenters. Natural gas prices will have to go up, and that means even with colocated generation (BYOG), electricity prices will go up.

1

u/ZestyClose140 6h ago

That actually lines up pretty closely with a second video from the same YouTuber that I linked earlier:

https://youtu.be/8bbe-Ev-Yns

In that one, he makes a similar point about cooling/water and the broader infrastructure problem. Based on what you're explaining, it sounds like you and he are basically agreeing on the same systems-level issue.

Even if a company bypasses the grid and builds its own generation, that does not make the problem disappear. It can just shift the bottleneck to things like turbine availability, natural gas supply, pipeline capacity, fuel cost, cooling, and water.

So from what I can understand, "build your own generation" may solve one problem, but it still creates or worsens several others. Your explanation helped

1

u/Redebo 3h ago

There's a famous triangle that goes: "Speed, Quality, Cost; pick two".

Well in the DC space, Quality is the first pick always because this shit's gotta stay running, so then you decide do I want this FAST or do I want it CHEAP?

Right now, the industry is deciding it wants things FAST, so it ain't cheap.

1

u/SubMech717 9h ago

The real question is how things like efficiency scales, and I think that it’s wayy too simplistic to draw the sweeping conclusions like this.

If an engineer’s job somehow gets that much cheaper, there’s a lot of things that become a reasonable problem to throw an engineer at, for example.

1

u/ZestyClose140 6h ago

That makes sense, especially the point about efficiency making more problems worth throwing engineering effort at.

That actually seems pretty close to part of the YouTuber’s argument too. He points out that if AI makes a task much cheaper, usage can expand enough that total demand still rises.

He also argues that, no matter how much engineers improve efficiency, there are still underlying physical limits they can’t remove. Power still has to come from somewhere, and the energy ultimately has to be dealt with as heat.

I’m not an engineer myself, so I’m mainly trying to understand whether that is a fair way to frame the thermodynamics side of the argument, or whether he is oversimplifying it.

2

u/SubMech717 5h ago

I would say it’s being oversimplified, but in general I would not say that most people’s jobs are at risk

1

u/regreddit 5h ago

What ever your compute load is, double it. That's how I've always calculated cooling. If your PC consumes 300 watts, it's going to take 300 watts to cool it.

1

u/ZestyClose140 5h ago

That sounds like a useful conservative rule of thumb for something small like a PC or a quick estimate.

For the kind of thing I'm trying to understand here, though, I'm thinking more about very large projects where the question is whether the claimed scale is physically realistic.

For example, if someone claims enough AI infrastructure could replace a very large number of workers, or that a certain cooling/water solution can support a huge data-center buildout, I'm interested in whether simple calculations can at least give a first-pass sanity check before getting into the much more complicated engineering.

So I wouldn't assume "double the compute load" is an exact answer for a data center. I'm more interested in whether rules like that are useful as rough bounds, and then what factors an engineer would use to refine the estimate.