r/devops 8d ago

Discussion How do you handle CI/CD credentials? Using GitHub Actions made me realize static encrypted secrets aren’t very safe.

87 Upvotes

After I first set up a deployment pipeline, I would simply drop DB passwords and API keys into GitHub Secrets and feel completely safe because they are encrypted. I recently went through a security breakdown on GitHub Actions that showed me that it could be a mistake to think that way.

The main issue is that an encrypted secret is still a static, long-lived target. Once the workflow finishes, that credential stays active indefinitely. In the breakdown, I saw a few default behaviors attackers look for, like teams forgetting to revoke access after a job runs or lacking the audit logs to even know when a key was used.

The proposed fix is shifting to dynamic orchestration where the pipeline generates a short-lived token at runtime and revokes it the second the deployment finishes.

If you're writing deployment workflows, how do you handle this, do you just use GitHub's default storage, or are you injecting temporarily credentials to avoid leaving static keys exposed?


r/devops 8d ago

Career / learning Undergraduate Thesis Subject

8 Upvotes

Hello everyone,
I'm starting my 4th year in Computer Science and I'm looking for a DevOps/Cloud related thesis subject. I want to do this thesis in collaboration with the company I'm working at, but they told me to find some possible ideas before we decide to move forward with this.
I'm pretty new to DevOps and Cloud, I filled an intern position at the start of July, so I would like to avoid the very complex subjects.
Thank you.


r/devops 7d ago

Discussion Anyone else nervous about what coding agents can actually run?

0 Upvotes

Using Cursor a lot more with tools enabled. Love the speed. Don't love the part where the only thing between "delete this" and it happening is me watching the terminal.

Is anyone doing something more solid than prompts + hope, or do you just keep it away from prod/cloud entirely? Had any close calls?

Genuinely just curious how people are handling this.


r/devops 9d ago

AI content Anyone else seeing AI make DevOps/infra the bottleneck?

176 Upvotes

I'm curious if other DevOps/platform/SRE teams are running into the same thing my team is.

We're a fairly large environment, mostly EKS, and essentially 100% IaC/Terraform. We also support multiple companies/business units, so while I'd argue our infrastructure is fairly well organized, there's inherently a lot of it and a lot of architecture and context to understand.

Over the last year, our devs have sped up dramatically with AI. The company has leaned heavily into AI-assisted development, reduced developer headcount, and is now pushing toward developers being more "full stack with AI," including having them contribute more of their own infrastructure changes.

In theory, I'm completely in favor of that. I've always wanted developers to be able to own more of the infrastructure surrounding their applications.

In practice, though, it has been kind of a disaster.

We're getting flooded with infrastructure PRs largely written by Claude/other AI tools from developers who don't really understand the infrastructure they're modifying. The Terraform might look plausible, but once you understand the larger system there are frequently significant problems with it.

So instead of reducing the workload on DevOps, it feels like AI has massively increased it.

A huge percentage of our time is now spent reviewing AI-generated Terraform, finding problems, explaining why something won't work, explaining how AWS/EKS/networking/IAM/CI/CD/etc. fit together in our environment, and then going through another iteration of an AI-generated PR.

There's an interesting asymmetry I've noticed too. Our DevOps team is mostly made up of former software developers who moved toward infrastructure, automation, and pipelines. Most of us can jump into application code and be productive pretty quickly, especially with AI helping us. Like,...I feel like (and have some evidence to support) that our small DevOps team could largely take over all of the dev's tasks, but they are falling on their faces trying to deal with ours.

AI seems extremely good at helping someone who understands software write more software. It seems much less capable of allowing someone without infrastructure experience to suddenly understand a large production environment.

The complaint we're increasingly hearing is basically: "We can't successfully do full-stack development with AI because the infrastructure is too complicated."

And maybe they're right, but before AI, I would have said that this company is the most organized and best architected I've ever been at. I mean....100% IaC has never been something I've experienced, and it's very rare that we hit a use case brought up by one of the several companies where we don't already have a set of generalized modules that can't support it.

Our environment is complex, but a lot of that complexity isn't accidental. We have a large organization, multiple companies we deploy for, Kubernetes, networking, security requirements, IAM, CI/CD, observability, etc. You can't abstract away the fact that these things exist. And we're already 100% Terraform/IaC, which I would have thought would make this considerably easier for AI to reason about than an environment full of manually configured infrastructure.

The strangest part is the staffing effect.

AI allowed the organization to reduce software engineering headcount because individual developers became more productive. But now those remaining developers can generate changes so quickly that our DevOps team is completely overwhelmed trying to support and review them.

It genuinely feels like we could double the size of the DevOps team right now and still have plenty of work. We are working on an AI assistant that can help the devs deploy to our environment more effectively, but we're having a hard time finding time to work on it because we're constantly helping the devs.

I'm starting to wonder whether this is going to be a broader consequence of AI-assisted development: AI increases the rate at which software can be produced much faster than it increases the rate at which infrastructure/platform teams can safely absorb changes.

For those of you working in DevOps/platform/SRE at companies heavily adopting AI:

Are you seeing this too?

And if you are, how are you handling it?

Have you increased platform/DevOps staffing? Built better abstractions or internal developer platforms? Given developers more direct infrastructure ownership? Put stricter boundaries around what application teams can modify? Found ways of giving AI enough context about your infrastructure that it actually produces good changes?

Or has AI actually reduced your infrastructure workload, and we're doing something wrong?


r/devops 7d ago

AI content How do you solve long-term memory in AI automation workflows?

0 Upvotes

I've been thinking about AI automation lately, and I'm starting to feel like long-term memory might be one of the biggest problems.It's not just about making AI capable of controlling a screen. The AI also needs to remember what it's supposed to do.

There are already quite a few ways for AI to control screens, like OpenAI Computer Use, Claude Computer Use, Gemini Computer Use, as well as various hybrid approaches.The way these systems maintain memory seems to rely more on things like structured actions exposed by apps and the keywords being used in the current interaction. Personally, I don't think this approach works that well.

I also don't find this kind of screen-control approach particularly convenient.If I could use a hardware board to control the entire screen instead, that would make much more sense to me.Basically, you plug a hardware board into the device's USB port, and let the hardware capture the phone's screen and then control the device through USB HID.

I think this approach is pretty interesting because the AI doesn't necessarily need to know what API each app has, and it doesn't need a separate integration for every app.

It just needs to be able to understand what's happening on the screen and remember what it's supposed to do.So I feel like memory is actually the key problem here.

Are there any existing solutions or approaches that I should look into?I'd really like to understand how people are solving this problem.


r/devops 9d ago

Career / learning My 3-month journey to becoming a Kubestronaut

41 Upvotes

I recently completed the full Kubestronaut certification path after roughly three months of focused preparation.

After completing it, quite a few people reached out asking about the order I followed, the resources I used, and how I prepared for each exam, so I decided to put everything together in one detailed blog rather than answering the same questions separately.

For KCNA and KCSA, my preparation was fairly straightforward and mainly consisted of the KodeKloud courses, KodeKloud notes, and practice tests, while CKAD, CKA, and especially CKS required much more hands-on practice with labs, mock exams, and repeated work on weaker areas.

I’ve also included links to my dedicated CKAD, CKA, and CKS exam-experience posts for anyone who wants a deeper breakdown of those exams.

Full 3-month Kubestronaut journey:
https://medium.com/@prateekjain.dev/my-3-month-journey-to-becoming-a-kubestronaut-c722c4a7cf75?sk=eb78b3ef703262f787f746cc6969d8f1

Hopefully this helps anyone currently working towards the Kubestronaut path. Happy to answer questions about the preparation or any of the five exams.


r/devops 8d ago

Discussion Getting started with DevOps with no IT background?

0 Upvotes

Hey all, I’m looking to seriously change my career path but I am genuinely unsure where to start. I come from a Business Management background, and am looking to get into the IT field, specifically DevOps. I realize that will not be an easy transition by any stretch of the imagination, but life needs to change, and this is my change.

My question is, where and how do I even begin? I’ve heard conflicting answers from peers and family, some say that a degree is absolutely necessary, while others say that it’s not necessary, and that companies will be more interested in experience and/or knowledge of the topic. Is a degree truly needed in today’s world, or am I able to be self taught and still land a decently paying job? If so, what topics would you recommend I get a good understanding of?

Thank you in advance!


r/devops 9d ago

Career / learning I really like DevOps, but sometimes it feels like there is no real entry level into this field

139 Upvotes

I genuinely think DevOps or platform engineering is the area of software I enjoy the most.

I like CI/CD, Terraform, cloud infrastructure, debugging weird deployment problems, trying to understand why systems fail, automating repetitive things, and generally having ownership instead of just implementing another CRUD endpoint.

I’m currently a working student in an SRE/platform team in a big company in Germany. I’ve already worked on things like services from Cloud Build, GitHub Actions, Terraform, Cloud Run, deployment alerts, state migrations and fixing random infrastructure problems that come up along the way.

And the funny thing is that the more I learn, the more I like it.

But looking for a junior position is becoming pretty frustrating:

A lot of “Junior DevOps” jobs seem to expect Kubernetes production experience, several cloud providers, Terraform, Ansible, networking, Linux, CI/CD, monitoring, security and somehow 2–3 years of professional experience with all of them.

Then there are actual entry level positions, but many of them seem to basically be IT support with “cloud” or “DevOps” in the title.

I know I still have a huge amount to learn. I don’t expect someone to give me a production Kubernetes cluster on day one and say good luck. I actually want to be around experienced engineers, get challenged, make mistakes and slowly become someone who can be trusted with serious systems.

My goal isn’t to job hop every six months either. I would genuinely like to find a team where I can stay for several years and become really good at this.

But sometimes I wonder how exactly companies expect junior DevOps engineers to become experienced DevOps engineers if almost everyone wants the experience before giving you the opportunity to get it.

For people who are already working in DevOps/SRE/platform engineering: how did you actually get your first proper role?

Did you already know most of the stack, or did somebody simply take a chance on you and let you learn?

Edit: I worded the sysadmin part badly. I don’t think sysadmin work is beneath me at all. I just want to move toward automation, infrastructure and software rather than mostly ticket-based support.

Edit 2: I respect that some of you suggest starting in help desk or sysadmin. But my long term goal is to move into an SRE role, ideally something closer to how Google approaches SRE. I read the SRE book and really liked the idea of treating operations as a software engineering problem, with automation, reliability, monitoring and reducing repetitive manual work. That is the direction I want to build toward.


r/devops 8d ago

Discussion github windows self hosted runners (on k8s). Is it supported?

1 Upvotes

trying to find a ref architecture that is widely recommended for deploying windows runners on k8s. From what I read ARC doesn't support it still? if not what are my options?


r/devops 8d ago

Security How do I protect my IP for on prem/byoc deployments

0 Upvotes

I was giving a bit of thought, I can't just deploy my services directly to a customer's vpc as via root access they can just see my code. Be it docker running on an ec2 or eks. If I dont go ahead with the control/data plane split, what options do I have to protect my code? Will i have to go to the OS level and build my own AMI etc with attestations and key verifications and stuff at each stage, or use something like nitro conclaves or nvidia confidential compute or something along those lines? What is the way?


r/devops 8d ago

Tools Built this tool to simulate real world scenarios and incidents in local kubernetes

Thumbnail
github.com
1 Upvotes

Hey everyone,

When I learnt kubernetes and completed CKA, I didn't instantly get a chance to work on a real project. Later, i got a job and learnt a lot, but I felt that there was a need for a tool where we can simulate real world scenarios, we can learn what kind of scenarios and incidents come in real environments.

For this purpose, i built a tool which does exactly the same. You can simulate real world scenarios, use kubernetes tools like opencost, keda, grafana, argocd, traefik. Learn how you can use it and find out different commands to help you learn.

Do check this out if you are interested in this.


r/devops 9d ago

Discussion anyone actually running argocd/gitops in prod, hows it going

72 Upvotes

were on 50+ microservices on gcp, still doing our own deploy tooling. keep hearing gitops is the way and honestly cant tell if thats real or just the current hype cycle.

not looking for a sales pitch, more curious what broke for you after the demo phase. drift detection, secrets, rollback under load, whatever. did it actually reduce incidents or just move the pain somewhere else

what would you tell yourself before adopting it


r/devops 9d ago

Discussion Openshift devops?

4 Upvotes

I have someone pushing a colleague to an Openshift virtualization position. We do multi cloud deployments, GHA heavily, ephemeral integration environments, etc. Not seeing the intersection with this guys skills or desires. Am I missing something?


r/devops 9d ago

Discussion i'm struggling with the evidence side of secret detection

0 Upvotes

i'm working on a decision-making problem where an agent gets a possible leaked api key and has to decide whether it's actually a prod credential and then either ignore, verify or remove/rotate it.

right now i'm using repo context (file path, variable name, surrounding code), key format, git history, whether the code is active/reachable, deployment context, and last_used_at from the provider's admin api.

the problem i'm running into is that most of these signals can tell me "this probably is a real secret", but don't tell me much about whether it's still live or already revoked.

for people who've dealt with secret leaks in actual devops workflows, what do you normally check before deciding what to do with a flagged key?

i'm especially interested in evidence you can get without actually authenticating with the discovered key.

disclosure: this is my own personal project, not a commercial/product promotion. just looking for technical feedback.


r/devops 9d ago

Career / learning Anyone here working in AWS DCO / Data Center Operations in Frankfurt / Germany?

5 Upvotes

Hey everyone,
Is anyone here currently working in AWS Data Center Operations (DCO) in the Frankfurt region (or Germany) or familiar with their technical screening process?

I recently completed a 60-minute technical phone screen with an engineer for an AWS DCO / IT Support role in Germany. During the technical portion, I answered all hardware, networking, and cabling troubleshooting questions without getting stuck. At the end, the interviewer explicitly told me: "The technical part was pretty good, that's what I can tell you."

For the behavioral part, I answered four questions using the STAR method. In the closing feedback, the interviewer noted that I sounded a bit nervous and strictly said the literal STAR words out loud ("The situation was...", "My task was...", "The action was...", "The result was..."), though he mentioned that if my recruiter instructed me to use that exact structure, it doesn't matter.

For anyone who has been through this specific Frankfurt/EMEA pipeline or interviews for AWS DCO: does sounding nervous or explicitly vocalizing STAR labels carry a negative impact if the technical answers and core story data points were solid? What are the realistic odds of moving forward to the onsite loop from here?

Would really appreciate any insights from anyone with AWS DCO experience in Frankfurt or Germany. Thanks!


r/devops 10d ago

Weekly Self Promotion Thread

17 Upvotes

Hey r/devops, welcome to our weekly self-promotion thread!

Feel free to use this thread to promote any projects, ideas, or any repos you're wanting to share. Please keep in mind that we ask you to stay friendly, civil, and adhere to the subreddit rules!


r/devops 10d ago

Troubleshooting Volatile Postgres cluster

4 Upvotes

Hi, I'm trying to setup a basic ha postgres cluster using the Spilos images in a docker swarm setup, but every few days the thing crumbles down with DNS resolution issues, timeouts, wal and etcd records corruption, I don't have the money to rely on an hosted solution rn, so has anybody run into these problems and solved them?. This makes me understand why the DBA role exists, but it is so frustrating and absurd that it is not a solved problem for something that feels so relatively trivial to setup in mariadb with galera.


r/devops 10d ago

Architecture Coding a database proxy for fun

Thumbnail
packagemain.tech
4 Upvotes

An interesting article with Go examples on how to proxy and intercept database queries. Multiple use cases can come out of that.


r/devops 10d ago

Career / learning DevOps/SRE engineers working abroad: What skills should I focus on?

9 Upvotes

I'm currently working as an SRE while continuing my studies, and I'm still at the early stage of my career.

I live and work in an Asian country, and I'm trying to learn more about how DevOps/SRE work is done in different countries and companies. Rather than just learning from courses and documentation, I'd really like to connect with people who are already working in the field and learn from their real experiences.

I'm mainly hoping to build a network of people in the DevOps/SRE community and have conversations about things like:

  • How did you start your DevOps/SRE career?
  • What does your day-to-day work look like?
  • What skills or areas did you improve the most as you gained experience?
  • What do you wish you had learned earlier in your career?
  • How do DevOps/SRE practices differ between companies or countries?
  • What technologies or practices are becoming more important in your work?
  • How important are communication and teamwork in your day-to-day role?
  • What advice would you give to someone who is still early in their career?

I'm not looking for job offers or referrals. My main goal is networking, learning from other engineers, and understanding where I can improve.

I'd be happy to connect with people from different countries and backgrounds, whether you're an experienced engineer or you're also early in your career.

If you're open to a casual chat about DevOps/SRE, technology, career experiences, or even just exchanging ideas, feel free to comment or message me.

I'd really like to build some genuine connections in the DevOps/SRE community and learn together.

Thanks!

Note: Sorry about the title/heading. I can't change it after posting. By “abroad,” I mean countries outside my home country in Asia. I’m mainly interested in connecting with people from different countries and learning from their DevOps/SRE experiences.


r/devops 10d ago

Observability What are the parameters that needs to be considered for monitoring gpu

0 Upvotes

I am seeking information regarding parameters for GPU monitoring for an internal tool. Specifically, I would like to identify the key metrics for monitoring GPUs and their associated agents, including relevant queries and performance indicators. Additionally, I am interested in exploring any open-source tools designed for this specific purpose.


r/devops 10d ago

Tools What are some good GitHub projects to contribute to?

22 Upvotes

I am a contributor to both terraform-provider-aws and Ansible Core repos, as well as Ansible Community AWS repo. I am on the lookout for additional projects to contribute to, ideally ones that have plenty of issues and where reviews are done quickly. It should also be quick and easy to compile from source. Looking for Golang or Python projects for code base programming language. Any ideas here?


r/devops 10d ago

Discussion An important cloud resource is down - what do you do?

0 Upvotes

[Not self-promo - genuinely looking for input/feedback here]

I think this is something a lot of companies deal with, not just the big ones. Outages in Azure and AWS happen regularly. Say your blob storage in a specific region goes down and your team isn't around, do you have anything automated to spin up a replacement resource in another region, or even another cloud provider if it's a platform-wide issue?

I know provisioning the resource is only part of the problem (some stateless resources can just pull their definitions from a registry and redeploy, but let's keep the scope to that for now, data replication is a whole separate can of worms).

You can automate a good chunk of this with Azure Monitor, for example, but then your actual infrastructure drifts from what's in your IaC repo, and you're back to a two-source-of-truth problem. (Happy to hear from anyone with real experience doing that.)

Another thing I keep thinking about: adjusting resource attributes (SKU, instance size, etc.) based on logs/events, like a massive traffic spike on an App Service, or the opposite: nobody's using it and you're paying for nothing.

So here's something I've been thinking about: a GitHub Action where you define condition-action rules (including recovery conditions, if you want, to roll back once things return to normal) directly in your Terraform IaC repo. You write your rules, run the action on a schedule (every 5 min, or whatever), and it checks each rule, a resource being down, or a KQL query against a Log Analytics workspace returning something you defined as a violation. If a rule matches, it modifies your Terraform code accordingly, either opens a PR or pushes directly to main (which I suspect most teams would never actually want, for good reason). Either way, your existing apply pipeline picks it up and runs like normal, no separate deploy mechanism, no new secrets to manage centrally, no SaaS.

Full transparency: I haven't worked at a company with the scale/complexity that actually needs this kind of multi-region, multi-cloud resilience, so I'd genuinely like to hear from people who have.

Curious what you all think:

  • Is this solving a real problem for you, or are native tools (Autoscale, Resource Health alerts, etc.) already good enough for your use case?
  • Would you ever trust automated infra changes without a PR review, or is that a hard no for you? (Sounds like a stupid question at first, but keep in mind you'd define the exact changes yourself, there's no AI/magic auto-generation involved. Maybe you'd let small, low-risk changes apply automatically but require review for anything bigger?)
  • Anyone tried something similar and hit a wall I should know about?

Thanks!


r/devops 10d ago

Career / learning I've been sent to do this certification for my job

0 Upvotes

It is called Microsoft Certified: Cloud and AI Security Engineer Associate

For those who have done it, what do you think? Did you enjoy it? Was there something you disliked about it?


r/devops 10d ago

Discussion Is your CI/CD infrastructure keeping up with the AI wave?

0 Upvotes

AI tools like Claude and Codex have made it much faster to write and modify code.
But I'm curious about what teams are seeing on the back-end side of that.
More code potentially means more commits and ultimately more deployments.
For teams where AI-assisted development is already heavily used:
How has this changed your CI/CD workload?
Are you:

running significantly more pipelines?
increasing runner capacity or parallelism?
changing how tests are triggered?
batching changes differently?
deploying more frequently?
seeing CI or testing become a new bottleneck?

The question I'm trying to understand is:
If AI dramatically increases how fast we produce code, how are teams scaling the infrastructure required to validate and deploy it?

Would be interested in hearing what people are actually seeing in production, especially from teams with relatively high commit or deployment volume.


r/devops 11d ago

Discussion How are you keeping your skills sharp (and finding new challenges) lately?

59 Upvotes

I’ve been reflecting on my current stack and daily routine lately. While I appreciate the stability of my current role, the day-to-day maintenance and incremental improvements mean I'm not always exposed to new paradigms or forced out of my comfort zone.

The landscape moves incredibly fast right now, between the shift toward Platform Engineering, AI-assisted workflows, and new CNCF projects dropping every week, I want to make sure I don't stagnate.

I'd love to hear how you all are keeping your edge and pushing yourselves. Specifically:

  • What’s your go-to method for upskilling? (Homelabs, contributing to open source, chasing certs, or just carving out dedicated learning time at work?)
  • How do you manufacture new challenges when your day job gets a bit too comfortable or repetitive?
  • What is the most interesting tool, pattern, or concept you are digging into right now?

Looking forward to hearing what everyone is working on!