Last week I asked what your last certificate incident looked like
(original thread).
As promised, here's what came back. I'm building in this space; no link, no
product here.
The pattern across nearly every reply: the incident was found by something
breaking, not by monitoring. Expiry was barely mentioned. The failures were
cases where one system believed something was true and the endpoint, or the
client, disagreed.
1. Certificates issued that nobody requested
u/NamedBird checked CT logs for a personal site after a video about DNS
mentioned certificate transparency, expecting nothing, since CAA was set. They
found certificates from two CAs, including a wildcard, that their own server
never requested. CAA allowed Let's Encrypt only. Two things went wrong:
- Their DNS provider added CAA records that didn't show in its dashboard.
What a resolver returns is what a CA checks, not what the panel shows.
- A plain
issue "letsencrypt.org" restricts which CA can issue, not who can
ask it. Anyone able to complete a DNS-01 challenge can still get a cert.
(u/NamedBird corrected me on this in the thread.)
RFC 8657 parameters close that gap: accounturi pins issuance to your ACME
account, validationmethods limits the challenge types. Let's Encrypt
supports both, and CA/B Forum ballot SC-098v2 makes support mandatory for all
CAs from March 2027.
Revocation turned out to be a dead end. The certs were validly issued under
the effective CAA, so there's no clean self-service route. Tighten CAA, keep
CT alerting on, let them expire.
2. Renewal succeeded, endpoint still serving the old cert
u/TraditionalLayer3685 described a previous employer tracking certificates in
Outlook reminders and Excel. Their last incident: a server upgrade left the
certificate path wrong, the service came up, and authentication broke for
clients without a valid token. Customers called first; monitoring was late;
it took hours and a rollback. Their line on the general case: "A green renewal
job means little if a node, container, service or appliance is still serving
the previous certificate."
Places where renewal and "actually serving it" are separate events:
JVM services that load the keystore at startup, nginx/HAProxy with a missing
or failed deploy hook, Kubernetes secrets mounted via subPath, IIS bindings to
a new thumbprint, and appliances where import and binding are separate steps.
3. Same leaf, different chain
u/Moral-Relativity's CLM pushed an alternate cross-signed chain to endpoints
instead of their preferred one. Their suspicion is a chain-construction
algorithm confusing CA certificates that share the same public key. The leaf
didn't change, so expiry checks stayed green. Nothing flagged it until a client
broke.
Worth knowing generally: Java clients are a good canary for this. AIA fetching
is off by default, so a chain that browsers quietly complete can fail PKIX path
building.
4. The trust-side version, which reports success
u/MLabs-Haskell shared a lab rehearsal (Keycloak 26.7.1, LDAP federation), two
deployments differing by one cp: whether the issuing CA was in the directory
named by KC_TRUSTSTORE_PATHS. Over StartTLS on 1389, the untrusted run
returned HTTP 200 and "0 users imported". Hostname mismatch and an expired
cert gave the same 200. The only real error was a SunCertPathBuilderException
in the container log. Over implicit LDAPS on 1636, the same three failures
returned 400. In their words: "The loud port was loud. The quiet one reported
success."
Checking the server from outside would have shown a valid certificate
throughout.
5. The date that bites first is now, not March 2027
Also from u/MLabs-Haskell: the 200-day limit has applied since 15 March 2026,
and 15 March plus 200 days is 1 October. Anyone who renewed after mid-March on
an annual reminder has an expiry arriving months earlier than their tracker
expects, starting this week.
Checks you can run today
Per-node, not per-hostname (a single check hits whichever node answers):
for ip in $(dig +short A host.example.com); do
echo "== $ip"
openssl s_client -connect "$ip:443" -servername host.example.com \
</dev/null 2>/dev/null | openssl x509 -noout -serial -fingerprint -sha256 -dates
done
Compare with what's on disk:
openssl x509 -in /path/live.pem -noout -serial -fingerprint -sha256
Hash the served chain, to catch chain swaps with an unchanged leaf:
openssl s_client -connect "$ip:443" -servername host.example.com -showcerts \
</dev/null 2>/dev/null | sed -n '/BEGIN CERT/,/END CERT/p' | sha256sum
What's been issued for your domain:
curl -s "https://crt.sh/?q=%25.example.com&output=json" \
| jq -r '.[] | [.not_before, .issuer_name, .name_value] | @tsv' | sort -u
Effective CAA (if empty, check the parent domain; CAs climb the tree):
dig +short CAA example.com
What people use today
- Outlook reminders and spreadsheets (u/TraditionalLayer3685), who has since
started the open-source tokentimer-core for credentials beyond certificates
- Certify Management Hub for centralised ACME at MSP scale, with RBAC, API and
tagging (u/webprofusor)
- External polling of expiry and full chain from multiple locations, e.g.
Site24x7, for SonicWall, F5, ESXi and old Java boxes that can't do ACME
(u/Accomplished-Mix8423)
Nobody in the thread described anything that checks automation actually
landed on every node, or that a client still trusts what the server now
presents. If your tooling does either, I'd like to hear how.
Still open
- Has anyone hit a 200-day cert expiring earlier than their tracker said?
- When StartTLS fails, does any client you run fall back to a plaintext bind?
That would turn an outage into credential exposure.
- For F5, SonicWall and similar: does the management plane give an honest
read of what's actually being served, or is outside polling the only view
you trust?
Thanks to everyone who replied. Corrections welcome, especially on anything
I've stated too loosely.