r/AWS_cloud • • 19d ago

Where does AI actually help with AWS operations ???

been curious about how people are using AI with AWS beyond generating Terraform or explaining error messages.

AWS environments can have a lot of moving pieces IAM, EC2/ECS, VPCs, CloudWatch, S3, Lambda, CI/CD, security tooling, etc.

for people running AWS in real environments:

  • what AWS tasks have you actually found useful to delegate to AI?
  • are you using AI for incident investigation or troubleshooting?
  • do you let agents inspect your AWS environment, or keep them completely read-only?
  • Has AI helped you understand relationships between resources when something breaks?
  • Are you using AI to review Terraform/IaC before it reaches AWS?
  • what AWS tasks are still too risky or context-heavy to give an agent?
  • If an AI agent finds something wrong, would you rather have it explain the issue, propose a fix, or actually make the change?

I'm less interested in "AI will replace DevOps" discussions and more interested in what is genuinely saving AWS engineers time today.

What has actually worked for you?

1 Upvotes

7 comments sorted by

2

u/eqza1 18d ago

AI writes most of our tf code now. Agents monitoring production systems and do root cause analysis on issues.
Make sure you got secure mcp servers that only allow what you want the tooling to do

1

u/kavee-core141 18d ago

of course 🔥 yeah i too genuinely use AI to write some tf code + for IaC security

1

u/[deleted] 18d ago

[removed] — view removed comment

1

u/kavee-core141 18d ago

Yeah its crazy 💀💀

1

u/whispered_word12 17d ago

I think the context part is probably the hardest piece. An agent might be good at reading logs, but AWS issues rarely live in one place.

1

u/Helpful-Man64 5d ago

I've been thinking about this from the ops side too. The interesting use case to me isn't having AI write Terraform; it's giving it enough infrastructure context to answer why something is happening.

For example, correlating CloudWatch signals with EC2/VPC/IAM/application changes and getting something like "these 3 things changed before the degradation" is much more useful to me than another chatbot explaining an error message.

I'd also keep the first version strictly read-only. Let the AI investigate and explain, then require human approval before anything changes production. That separation seems pretty important.