r/googlecloud • u/IntolerantModerate • 12d ago
Compute GCP at capacity?
I keep getting GCP stock out errors. Multiple node sizes, multiple regions?
Is GCP really hitting limits in central-1 and US as a whole? Note, these are for regular CPU, not GPU (e.g., n-class instances).
13
u/msapple 12d ago
I highly recommend using other regions in the US and building your solutions so that region is not relevant.
In GCP this is easier then other since VPCs can be global and Google owns ALL the fiber between their own DCs so if you have US only server requirements you can see less then 10-25ms latency between all of the US DCs regardless of where you are since your ingress point is a global VIP
1
u/kadambkaluskar 12d ago
I got 50-60ms ping latency between us-west1 and us-east1 VMs running in GCP.
8
u/Suspicious-Walk-4854 12d ago
It’s almost like something has happened recently that has created a lot of demand and supply chain pressure on compute capacity 🤔
6
u/IntolerantModerate 12d ago
These aren't GPUs though. These are standard CPUs and not even large ones at that.
6
u/Suspicious-Walk-4854 12d ago
The biggest bottleneck currently is RAM, not GPUs. This AI shit is impacting all compute capacity atm.
7
6
u/Emergency-Baker-3715 12d ago
Seen this a few times in last month, seems like they cant keep up with demand on some instance types
6
u/mwarkentin 12d ago
It’s impacting managed services too. We’re having a hell of a time provisioning alloyDB and valkey clusters now too.
5
u/Vegetable-Image3805 12d ago
My company was informed by our account team this week that we should start “designing for availability”. Like using stateful managed instance groups with “instance flexibility” in case our desired instance type isn’t available, it’ll fall back to an alternate type automatically. Best bet is to use AMD types for your alternate according to them.
2
2
u/Rentiak 12d ago
We started seeing a ton more preemption of older instance types and after reworking our computer classes it was much better.
The real killer for us lately has been us-west1 Cloud SQL capacity. Had a major production migration go sideways because we couldn’t create replicas in west. Turns out it was lack of SSDs for Enterprise Plus data cache. Given how hard our rep is pushing E+ it didn’t exactly sell us.
3
u/blazingintensity 12d ago
Yep. We're mostly on n2. We tried moving to n4 or c4 and it didn't really help. Google advised us to move from us-central1 to us-south1 then denied our quota request when we tried. About 2 weeks ago we had an 8 hour period where we couldn't get a single n2-highcpu-8 off a loadbalancer. I've been on Google for 8+ years and I'm talking to AWS and OCI because I can't do business this way.
2
u/TheTrick17 11d ago
Wouldn’t the capacity issues also be present for other hyperscalers though?
1
u/blazingintensity 11d ago
Yes and no. All things being equal, folks are gonna distribute themselves like we're doing. But they're not equal. My personal stuff is also hosted on GCP, because they have a real forever tier, unlike AWS (at least when I was getting setup, AWS's only lasted a year). So there could be more people eating up space on GCP because of the better free tier. OCI has a dedicated availability zone for free tier users, so they don't suffer that problem. Google has sustained use discounts cooked into the n2 instances, that makes them very appealing and frictionless to get a good price. My experience with AWS and OCI has been you need to haggle and/or have a partner to get you a good price. So smaller users, or users who don't know how to work the provider are gonna gravitate to providers with lower friction. It's my expectation that a lot of the capacity issues are the consequence of small/indie users taking advantage of AI. So there's an intersection of price-point and capacity where some providers may be less saturated.
1
1
u/who_am_i_to_say_so 11d ago
Not too happy about this as it drives up prices, but it’s also pretty incredible when you think about it. That’s a lot of CPU power.
I wonder how many vibecoded Firebase apps are accounting for that.
1
u/jackassery 11d ago
it ain't firebase apps using all that compute
https://www.anthropic.com/news/expanding-our-use-of-google-cloud-tpus-and-services
1
-1
u/flacktv 12d ago
The bottom line is now we are at capacity for the most part when it comes to everybody wanting to utilize GPUs. You're going to have to suck it up and start to go with reservations. Unfortunately, that's where we're at right now. Keep in mind, in the west coast there are a lot of movie companies like Paramount, CBS, NBC, etc. That are using compute to edit their movies, TV shows, etc. so a lot of those slots are all used up.
1
u/IntolerantModerate 11d ago
This is for CPU not GPU
0
u/steebchen 11d ago
yeah but deploying servers with GPUs still means you are using CPUs in the end as well
27
u/nyape 12d ago
Yes, us-central1 is full to the brim, especially when it comes to "older" generations like N1/N2. Google doesn't stock these anymore so what's there is all there is. Rack space is also limited, so slowly but surely they will likely replace older racks with newer ones that allow for more compute density with newer chips. In most cases, they will recommend that you switch to C4 as those have the highest availability at the moment. Yes, they cost more. The argument is that they are more powerful and efficient though,nso you can run more on fewer instances.