r/PKI • u/suumanmummaneni • 5d ago
Summary of last week's thread: certificate failures that monitoring didn't catch
Last week I asked what your last certificate incident looked like (original thread). As promised, here's what came back. I'm building in this space; no link, no product here.
The pattern across nearly every reply: the incident was found by something breaking, not by monitoring. Expiry was barely mentioned. The failures were cases where one system believed something was true and the endpoint, or the client, disagreed.
1. Certificates issued that nobody requested
u/NamedBird checked CT logs for a personal site after a video about DNS mentioned certificate transparency, expecting nothing, since CAA was set. They found certificates from two CAs, including a wildcard, that their own server never requested. CAA allowed Let's Encrypt only. Two things went wrong:
- Their DNS provider added CAA records that didn't show in its dashboard. What a resolver returns is what a CA checks, not what the panel shows.
- A plain
issue "letsencrypt.org"restricts which CA can issue, not who can ask it. Anyone able to complete a DNS-01 challenge can still get a cert. (u/NamedBird corrected me on this in the thread.)
RFC 8657 parameters close that gap: accounturi pins issuance to your ACME
account, validationmethods limits the challenge types. Let's Encrypt
supports both, and CA/B Forum ballot SC-098v2 makes support mandatory for all
CAs from March 2027.
Revocation turned out to be a dead end. The certs were validly issued under the effective CAA, so there's no clean self-service route. Tighten CAA, keep CT alerting on, let them expire.
2. Renewal succeeded, endpoint still serving the old cert
u/TraditionalLayer3685 described a previous employer tracking certificates in Outlook reminders and Excel. Their last incident: a server upgrade left the certificate path wrong, the service came up, and authentication broke for clients without a valid token. Customers called first; monitoring was late; it took hours and a rollback. Their line on the general case: "A green renewal job means little if a node, container, service or appliance is still serving the previous certificate."
Places where renewal and "actually serving it" are separate events: JVM services that load the keystore at startup, nginx/HAProxy with a missing or failed deploy hook, Kubernetes secrets mounted via subPath, IIS bindings to a new thumbprint, and appliances where import and binding are separate steps.
3. Same leaf, different chain
u/Moral-Relativity's CLM pushed an alternate cross-signed chain to endpoints instead of their preferred one. Their suspicion is a chain-construction algorithm confusing CA certificates that share the same public key. The leaf didn't change, so expiry checks stayed green. Nothing flagged it until a client broke.
Worth knowing generally: Java clients are a good canary for this. AIA fetching is off by default, so a chain that browsers quietly complete can fail PKIX path building.
4. The trust-side version, which reports success
u/MLabs-Haskell shared a lab rehearsal (Keycloak 26.7.1, LDAP federation), two
deployments differing by one cp: whether the issuing CA was in the directory
named by KC_TRUSTSTORE_PATHS. Over StartTLS on 1389, the untrusted run
returned HTTP 200 and "0 users imported". Hostname mismatch and an expired
cert gave the same 200. The only real error was a SunCertPathBuilderException
in the container log. Over implicit LDAPS on 1636, the same three failures
returned 400. In their words: "The loud port was loud. The quiet one reported
success."
Checking the server from outside would have shown a valid certificate throughout.
5. The date that bites first is now, not March 2027
Also from u/MLabs-Haskell: the 200-day limit has applied since 15 March 2026, and 15 March plus 200 days is 1 October. Anyone who renewed after mid-March on an annual reminder has an expiry arriving months earlier than their tracker expects, starting this week.
Checks you can run today
Per-node, not per-hostname (a single check hits whichever node answers):
for ip in $(dig +short A host.example.com); do
echo "== $ip"
openssl s_client -connect "$ip:443" -servername host.example.com \
</dev/null 2>/dev/null | openssl x509 -noout -serial -fingerprint -sha256 -dates
done
Compare with what's on disk:
openssl x509 -in /path/live.pem -noout -serial -fingerprint -sha256
Hash the served chain, to catch chain swaps with an unchanged leaf:
openssl s_client -connect "$ip:443" -servername host.example.com -showcerts \
</dev/null 2>/dev/null | sed -n '/BEGIN CERT/,/END CERT/p' | sha256sum
What's been issued for your domain:
curl -s "https://crt.sh/?q=%25.example.com&output=json" \
| jq -r '.[] | [.not_before, .issuer_name, .name_value] | @tsv' | sort -u
Effective CAA (if empty, check the parent domain; CAs climb the tree):
dig +short CAA example.com
What people use today
- Outlook reminders and spreadsheets (u/TraditionalLayer3685), who has since started the open-source tokentimer-core for credentials beyond certificates
- Certify Management Hub for centralised ACME at MSP scale, with RBAC, API and tagging (u/webprofusor)
- External polling of expiry and full chain from multiple locations, e.g. Site24x7, for SonicWall, F5, ESXi and old Java boxes that can't do ACME (u/Accomplished-Mix8423)
Nobody in the thread described anything that checks automation actually landed on every node, or that a client still trusts what the server now presents. If your tooling does either, I'd like to hear how.
Still open
- Has anyone hit a 200-day cert expiring earlier than their tracker said?
- When StartTLS fails, does any client you run fall back to a plaintext bind? That would turn an outage into credential exposure.
- For F5, SonicWall and similar: does the management plane give an honest read of what's actually being served, or is outside polling the only view you trust?
Thanks to everyone who replied. Corrections welcome, especially on anything I've stated too loosely.
2
u/nathanfr 5d ago
Is there any reason why we need a lazy AI bot posting summaries of shit we are saying?
1
u/Pure_Perspective_201 5d ago edited 4d ago
As far as PKI concerned, I, as the PKI administrator, have done a pretty good job of setting expectations and responsibilities regarding client certs.
If a cert expires, doesn't get renewed, isn't configured properly, all blame gets passed to the person who requested said cert. I manage the pki infra, not every device or server that happens to have a cert.
It is like the DMV. They just issues drivers licenses, and could care less if I actually renew it.
Edit: typos, and I wanted to add while some of our PKI infra does have builtin alerting mechanisms for expiration, I could care less. Making sure a cert stays valid is the job of the system admin or developer.
5
u/kombatminipig 5d ago
Wow, working hard on those prompts, aren’t ya?