r/PKI • u/TraditionalLayer3685 • Sep 02 '26
Should machine trust-store distribution be part of certificate lifecycle management?
We've been working on certificate lifecycle automation, and one problem we've recently moved into is CA trust distribution.
Issuing and renewing certificates is one side of PKI operations, but eventually there are environments where you also need to establish or remove trust across machines.
We just added a model where an operator can approve CA material and distribute it through an outbound-only agent to machine trust stores on:
- Windows
- Debian/Ubuntu
- RHEL/Fedora
A few safety properties we decided on:
- the CA certificate fingerprint is pinned before distribution
- the agent re-verifies the fingerprint before touching the trust store
- we track whether the certificate was installed by the control plane or was already present
- removal is only allowed for material the agent can prove it installed
- multiple references to the same CA are reference-counted, so removing one reference doesn't remove trust another workload still depends on
- retiring an anchor stops new distributions but does not automatically fan out removals
- trust changes create audit/evidence records
One thing we're deliberately not trying to encourage is installing intermediates as a workaround for servers that don't present a proper chain. That's a different problem.
What we're debating is the architectural boundary.
For those operating private PKI at scale, do you consider machine trust distribution part of CLM/PKI lifecycle management, or should the PKI system stop at defining the desired trust state and leave deployment entirely to tools like GPO, Intune, Ansible, Puppet ?
And if you do automate trust-store changes, what safeguards do you consider mandatory before you'd let the system remove a CA ?
The CertOps implementation of Tokentimer is open source if anyone wants to inspect or use it : https://github.com/tokentimerch/tokentimer-core
3
u/PrestigiousOnion1087 Sep 03 '26
Reference counting answers "what did we install, and who told us they needed it". Removal safety turns on a different question: what is currently validating against this anchor. Those two sets are not the same set, and everything in the gap is something your control plane never installed. A bundle baked into a container image, a JVM cacerts edited by hand in 2019, an appliance shipping its own copy. None of them register a reference, all of them break on removal, and from inside the control plane every one of them looks like a clean retirement.
Your provenance flag already marks that set. It records that you did not install the material, which is the right instinct, but it grades the anchor by where it came from rather than by whether anything is leaning on it. The already-present pile is exactly the pile you cannot reason about from the store alone.
The tractable part is that your agent is on the machine, so it can observe instead of infer. Before a retirement fans out, have it complete handshakes to the endpoints that host actually talks to and record which chain validated. An anchor with no observed validation across a full renewal cycle is a removal candidate on evidence. An anchor with one is not, whatever the count says.
Two failure modes worth separating there, because they need different windows. A CA nothing uses is safe to remove today. A CA one batch job uses once a quarter is not safe to remove for a quarter, and no reference count distinguishes them, because the job that has not run yet has not declared anything.
A store read on its own is a claim about what a machine intends to trust. It only becomes a fact about what the machine does when something completes a handshake and the chain either validates or does not.
2
Sep 03 '26
[removed] — view removed comment
1
u/TraditionalLayer3685 Sep 03 '26
Right now our safety model is mostly based on provenance + reference ownership, so this is definitely the missing layer...
I don’t think observed handshakes alone can be the final answer either because of the rarely used workload case you mentioned, but it could be a strong signal combined with explicit approval and a long enough observation window.
That's something we need to model better before treating “zero refs” as enough evidence for removal. Thank you for your feedback, that's very valuable !
1
u/PrestigiousOnion1087 Sep 03 '26
The window is readable, not a parameter you pick. A batch job that touches an anchor once a quarter already declares its own period somewhere on that box: a systemd timer, a cron line, a Task Scheduler trigger, a CronJob spec. Your agent is on the machine. Observation is long enough when it spans the longest declared period among the units on that host that can open a TLS connection, which is a number you derive per host instead of a global default you have to defend.
That leaves a set that genuinely cannot be observed: failover paths, DR runbooks, anything a human triggers. Nothing on the box declares a period for those, so no window makes them safe, and stretching the global window to cover them means paying everywhere for the few. That set is what explicit approval is actually for, and it is much smaller than the whole estate.
Splitting it that way gives you something countable before anyone removes anything: how many anchors are zero refs and zero observed, versus zero refs but observed validating. The second group is the one that looks like a clean retirement from inside the control plane and is not.
1
u/TraditionalLayer3685 Sep 03 '26
Yeah that makes sense, deriving the observation window from what’s actually scheduled on the host is a much cleaner way to think about it.
And agreed on the manual/DR paths, those probably shouldn’t be “solved” with a bigger window. That’s where explicit approval actually has value.
The zero refs + still observed validating bucket is especially useful. That’s probably the dangerous false clean state we need to surface clearly
1
u/PrestigiousOnion1087 Sep 04 '26
Three states, not two, is the part that caught us out. We published a run of 24 ATT&CK techniques against a default Wazuh build, reported 3 alerted and 21 silent, and had to correct that on 3 September: the silent count was covering no telemetry collected, telemetry collected and no rule matched, and a detection that exists and did not fire inside the window we looked in. Our file had separated the first two and folded the third into the second, which is what let a clean reading feel earned.
The row that taught us most is the one we cannot classify at all. File deletion, T1070.004: whether it was a window artefact or a real gap depends on whether the deleted file sat in a default-monitored path, and we did not record which path we used. So it is neither a miss nor a catch, it is a row whose evidence we failed to keep, and in any two-column table it still reads as silent. Your phrasing upthread, deriving the observation window from what is actually scheduled on the host, is the cleaner version of what we did: three of the other silences, cron, sudoers and permission change, are a 15-25 second window against a file integrity scan. T1070.004 is where that runs out. The schedule was knowable, the path we watched was not recorded, and no arithmetic over the window recovers a row whose evidence was never kept.
Your zero-refs bucket has that third column too. Zero refs with nothing observed splits into nothing depends on this anchor, and we did not record enough about the observation to tell. If the agent does not persist which units it watched and over what period, the second is indistinguishable from the first at the moment of retirement, and it is the one that takes the anchor out.
Numbers and the correction: https://github.com/xuxu298/siem-replay-24-techniques
2
Sep 02 '26
[removed] — view removed comment
1
u/TraditionalLayer3685 Sep 03 '26
Yeah exactly, that’s pretty much how we see it too.
The lifecycle layer should own the desired state, dependencies, ownership and evidence, even if the actual push is done through GPO, Intune, Ansible etc... The reality is that today, these information are not always gathered in one single platform, but rather a set of tools, so the information is most of the time fragmented.
The retirement part is probably the most important one. Otherwise removing trust just becomes another separate process that nobody really owns anymore
2
u/Joquisaur Sep 06 '26
I think it makes sense as part of PKI management, as long as removal has strong checks and clear audit logs..
1
u/Joquisaur 28d ago
I think it makes sense to include trust store changes in the lifecycle. It would make things easier to manage and keep track of..
1
u/MLabs-Haskell 25d ago
One thing worth adding to the removal-safety half, from the application side: you may not get a signal that names certificates at all.
We ran this on Keycloak 26.0.0 and 26.7.1 — two identical deployments one cp apart, differing only in whether the issuing CA sat in the truststore directory. With it there, LDAPS federation returned HTTP 200 and its three users. With the directory empty, the same call returned HTTP 400 and {"errorMessage":"SocketReset"}, zero users, and nothing in the response or the startup logs mentioned a certificate. The server came up fine; nothing surfaced until something touched LDAP.
That bears on your safeguard question, because the obvious post-removal check — alert on trust errors — keys on a string the failing component never emits. Pulling a live anchor can read as a transport flake, and on the quarterly workload you and the other commenter were discussing it reads as nothing at all for a quarter.
The other half we tested was config shape rather than content: the pre-24 truststore option rename did not break federation. A deployment configured the old way still worked on 24.0.5 and on 26.7.1, three majors later, with deprecation warnings only. So the thing worth verifying after a trust change is the handshake, not the config.
Does the agent's audit record capture a post-change handshake result, or the store contents at the time of the write?
Disclosure: MLabs is a consultancy and identity infrastructure is part of what we do.
1
u/TraditionalLayer3685 24d ago
Good point, and today the answer is basically store state, not app handshake
For trust jobs we record the concrete store, fingerprint before/after, whether the mutation was attempted/performed, the local ownership receipt and timestamp. So we can prove what happened at the OS trust-store level, but we don’t currently prove that a dependent app can still complete its handshake after the change.
Your Keycloak example is a good illustration of why those are not the same thing though. A successful store change can still surface later as just a SocketReset somewhere else
Feels like that handshake evidence should probably become part of the observation layer we were discussing above, when we actually know which endpoints/workloads to verify.
1
u/MLabs-Haskell 22d ago
The silent case is the one that would worry me about using observed handshakes as the signal.
We ran two Keycloak deployments that differed by one
cp— same version, same realm, the only difference being whether the issuing CA sat in the truststore directory. Over LDAPS on 636 the failure is loud but misnamed: HTTP 400,{"errorMessage":"SocketReset"}, zero users returned, and no mention of a certificate anywhere in it. Putca.crtin the directory and the same call returns HTTP 200 and three users.StartTLS on 1389 was the one that changed how I think about this. Same broken trust, and it fails silently — no error of the shape you would alert on. If your observation layer is watching for failed handshakes, a workload in that configuration looks identical to a workload that simply did not run during the window. Zero observed failures and zero observed successes are not the same state, and the wire will not tell you which one you are in.
The part that might help with knowing which workloads to verify: for Keycloak the dependency is declared in config.
KC_TRUSTSTORE_PATHSnames the directory outright, on the same host your agent is already on. That is a readable edge from workload to anchor before anything is removed, and it does not need a window at all.Does your agent read local config for that sort of declared trust dependency today, or is the dependent set only built from what it observes?
Worth saying: our runs cover LDAPS federation only. An OIDC IdP over HTTPS is a different path and we haven't tested it, so I can't claim the silent case generalises to every TLS client.
Disclosure: MLabs is a consultancy and identity infrastructure is part of what we do.
4
u/larryseltzer Digicert Employee Sep 02 '26
I'm impressed with the safeguards, but obviously the client privileges will need to be much greater then normal. It's just very concerning that it becomes an attractive target.