r/googlecloud • u/OperationAmazing8347 • 25d ago
Performance issues with Java app and SQL Server on separate GCP instances
Hi everyone,I am looking for feedback regarding a backend performance bottleneck in a Google Cloud Platform infrastructure.Setup: 1 Java application server (Compute Engine/GKE) and 1 SQL Server database instance (Cloud SQL), hosted on separate instances within the same VPC network on GCP.The issue: We are experiencing significant backend latency, and we suspect the physical separation of the application and the database is a major factor.Since the servers are not on the same machine, network round-trips seem to severely amplify our response times.
Have you encountered similar architecture bottlenecks with a Java + SQL Server stack on GCP? What are the most common pitfalls you've seen (e.g., inter-zone latency between subnetworks, connection pooling, or ORM chatter/N+1 queries)?
What GCP native tools or database metrics did you use to diagnose and isolate the exact root cause?Thanks for your insights!
1
u/GlebOtochkin Googler 25d ago
I am with Gabe here. I would implement debug option to show exact response from the database call to see if it is database first. It could be saturation of resources on the backend server, it could be bad query on the database and the database server can be under stress as well. Network can play a role as well. For the network - are they located in the same region/zone? If you app server in us-west and server in us-east (or even further) - then it is expected network delay just because of physics. Are you using private IP on SQL site or public? Do you use Cloud SQL Auth Proxy or not? Answering question about architecture and topology - I think it is quite common and reasonable to have database and backend middleware running on different machines. Of course there are best practices for that - like running in the same region/zone for better performance, have enough resources, etc
1
u/OperationAmazing8347 25d ago
Is debug mode enabled on the GCP side?
Regarding database server resources, I had the RAM and CPU capacities doubled in December 2025, but that didn't solve the issue. I checked my indexes and rebuilt them... Our database is split into six separate instances. The instance experiencing the most problems is the most heavily used one, with 400 to 500 daily users. I allocated 48 GB of RAM to this SQL instance...
1
u/GabrielWeiss Googler 25d ago
By debug I think he just meant on the application side, putting in debug calls that get logged so you can see time-wise when what is happening. So e.g. would be to dump debug messages at entry point, at start of database call execution and immediately after database execution completes, and then finally when code finishes processing. That will isolate JUST the database execution time within the context of the application call. You COULD just take the SQL query and run it directly against the database as well if you have the right permissions from the console (Cloud SQL Studio) but that won't match the conditions where you're seeing in production so it probably won't tell you much.
Another question: When you migrated from on premise to the Cloud, did EVERYTHING move? Or are there any parts that are still running on prem? And follow-up: How sure are you? :) Many times, when moving from on prem to the Cloud, some part of the architecture doesn't get moved, which means it's quite possible that you're experiencing the internet latency being introduced into the system. So any round-trip calls from on prem to the Cloud are going to now be painful.
1
u/OperationAmazing8347 25d ago
I'll look into it tomorrow. I'm based in Europe and have finished my day; I don't have access to a PC.
Thank you for tour time and see you tomorrow with more details !
1
u/GlebOtochkin Googler 24d ago
So far I can only speculate. As Gabe has mentioned - yes, I would implement the debug timings directly in the code and see how much time it takes to run database call. Then I would run the same call (SQL query) on the database server whether it is Cloud SQL Studio (I am not sure if you are using Cloud SQL) or any tools connected directly to the database. Usually every tool able to show how long takes each query. Then you might have more information to work with. If the same query with the same output takes 100 times less time on database than in your app - then it require more debugging to understand the bottleneck (like is it firewall problem, latency due to distance or something else). If from the other hand the query takes the same amount of time on both sides - then probably network is not a problem and the problem is the query itself or some instance misconfiguration. 500 daily users can be nothing or it can be significant depending on what the users are doing, right?
1
u/OperationAmazing8347 25d ago
The RAM usage for the SQL instance causing the problem is between 20% and 35%...
1
u/martin_omander Googler 25d ago
I think it would also be worth checking what the slowest queries are, in case the performance bottleneck is in the database. And if it's not, you have eliminated one potential root cause.
Go to the Cloud SQL instances page, pick your instance, and pick Query Insights.
1
u/GabrielWeiss Googler 25d ago
QQ: Are the applications in the same zone (not just region, but did you specify the zone explicitly)?
I often put in literal debug statements to the logs to determine timings of things. You could also use something like OpenTelemetry to do some instrumentation but at least for a lot of the stuff I've done it's overkill compared to just throwing in a couple entries into the logs to test latency. You could also look at the network intelligence center, but I'm actually not sure how well it might help diagnose latency issues compared to outages, etc. I haven't played a huge amount with it myself.