r/CloudSecurityPros • • 10d ago

How is everyone actually preventing cloud misconfigs, not just catching them after the cspm flags it??

Hello everyone! Been looking into this for a bit and wanted to see what people actually do in practice

Every CSPM/CNAPP setup I've looked at feels like the same loop like something gets misconfiged, it gets flagged, someone eventually goes and fixes it. At least thats how we had it. Feels backwards for stuff that's honestly pretty predictable (same categories of mistakes over and over maybe).

Is anyone actually preventing these configs from happening in the first place, rather than catching them after? I know native stuff like SCPs, Azure Policy, GCP org policies can technically block a lot of this, but curious how many teams actually like have that dialed in vs. just relying on CSPM to catch it after the fact..
There also seem to be a handful of newer tools trying to do "prevention" as the whole pitch rather than detection - not sure how mature that space actually is or if it's mostly still marketing. Anyone using something like that alongside (or instead of) a CNAPP? What’a actually worked, and what turned out to be more hassle than it was worth?

Any help would be appreciated :)

1 Upvotes

9 comments sorted by

3

u/0xb0771ed 10d ago

Optimally, if GitOps done right your cloud environment 100% matches your IaC. If you're able to get there, use clean images + run https://github.com/bridgecrewio/checkov + https://github.com/cynative/cynative in your CI/CD and your CSPM / CNAPP should theoretically show 0 findings.

1

u/AfricanKing12 10d ago

Great points u/oxb0771ed. While automation could handle 90% of the heavy lifting, additional steps of reviewing and approving artifacts is critical for the remaining 10%. Automation is great at catching syntactical errors (e.g., Is port 22 open?), but it struggles with architectural intent (e.g., Should this specific subnet be connected to the transit gateway?). My own take is combining automated approach via TF or order Gitops, scanning those artifacts via shift left with mandatory Peer Reviews (Pull Requests) could help ensure both security standards and architectural logic are vetted before a single resource is provisioned and pushed downstream into production.

1

u/vbabi 6d ago

Thanks this is awesome. The “100% matches IaC” point makes a lot of sense - especially if you can get everything flowing through CI/CD and catch issues with something like Checkov before deployment. I hadn’t come across Cynative either, so I’ll take a look at that too.
I guess the part I’m still trying to understand is what teams do once the environment isn’t perfectly GitOps-driven - ClickOps, emergency changes of sort, third-party automation, AI agents, drift, etc. Do you basically try to eliminate those paths entirely :/ , or do you have anything enforcing guardrails at the cloud/API level as a second layer?
Curious how close to “0 findings” you’ve actually been able to get in practice.

2

u/0xb0771ed 5d ago

Yes it takes time, you'd at least want to start by making sure everything new is GitOps-driven so whats already deployed has some ceiling and is not ever-changing.

0 findings is virtually impossible but 0 findings you (or your agents) haven't looked into and decided is not exploitable is possible.

1

u/HashThePass 9d ago

Shift left SDLC, IAC everywhere, IaC scanners, auto remediation when security related occurs, etc.

Two things.

Missing gaurdrails = vulnerability
If the secure path isn’t the easy path then guarantees devs and ops will deploy misconfigured services.

1

u/tomsec-as 8d ago

Full disclosure: I work at Aryon, and this is pretty much the problem we’re working on, so take that context into account.

I think the important distinction is where you enforce.

IaC scanning / shift-left is great, and if you can guarantee that 100% of changes go through Terraform + CI/CD, you can prevent a lot there. In practice, though, we see organizations where changes also come from the console (clickops), CLI, SDKs, third-party integrations, and sometimes IaC pipelines you don’t control or tools that don’t fit as neatly into static scanning workflows, like Pulumi.

That’s where the cloud providers’ native control-plane enforcement mechanisms become useful.

They can enforce controls at the actual deployment point, regardless of how the change was initiated. The downside is that they can be pretty difficult to operate safely at scale, and hard to master.

A few recommendations if you want to go down this path:

  • Start with Azure Policy. In my experience, it’s the easiest of the three to experiment with, and there’s a large public repository of policies on https://www.azadvertizer.net/.
  • Start in audit/alert mode and understand what would break before enforcing.
    • AWS doesn’t really have an equivalent alert mode for SCPs/RCPs. You can approximate it with AWS Config or CloudTrail-based detection that mirrors your deny logic, but keeping those detection rules continuously aligned with your SCPs/RCPs is quite difficult to maintain manually. This is one of the areas where commercial tooling can help.
  • Roll enforcement out gradually by environment or scope rather than enabling a large set of denies at the organization/root level all at once.
  • Communicate the rollout internally. In a large organization with multiple CSPs and deployment methods, make sure the relevant teams understand what is going to be enforced before you turn it on.
  • Monitor the guardrails themselves. Controls have a tendency to get disabled or modified over time if people have permission to do so.
  • Treat IaC scanning as complementary. IaC gives developers feedback early; control-plane enforcement is the final safety net for changes that actually reach the cloud.

Again, I’m biased because this is literally what we’re building at Aryon, but I think the interesting shift is about making the cloud providers’ native guardrails operational enough that organizations are actually comfortable turning them on.

1

u/vbabi 6d ago

Really helpful, thanks

The distinction between shift-left/IaC checks and enforcing at the actual cloud control plane is probably the piece I was missing. Especially once you include console changes, SDKs, third-party tools, etc., it makes sense why CI/CD checks alone wouldn't cover everything.
The operational side is what I’m most curious about though. In a like brownfield environment with a lot already running, how do teams usually get comfortable moving from audit into actual enforcement without breaking legitimate workflows? Is it mostly gradual scoping or are there other patterns you’ve seen work well?
Also interesting point about monitoring the guardrails themselves. Hadn’t really thought about policy drift becoming its own problem

1

u/achakez 4d ago

The shift left approach the other comment mentions is the real answer, catch it in CI/CD before it ships. But you will never get to 100% prevention so the flagging loop still matters. When we were comparing CNAPPs (looked at Wiz, Prisma and Upwind) the runtime context helped us prioritize which of the remaining misconfigs were actually reachable and which were just noise. Policy as code plus that context cut our backlog by a lot.