Skip to content

Note

The certificate renewal nobody is watching

Automated renewal removed the annual panic and left a quieter problem behind it: nothing tells you when it stops. The expiry is in every handshake, so it is readable from anywhere by anyone — including the next visitor, who gets a full-page warning instead of your site.

  • TLS
  • Certificates
  • Monitoring

Written

Read the date yourself

The certificate arrives during the handshake, before any application data, so the expiry of every https host on the internet is public information. Two commands, and no access to the server:

openssl s_client -connect example.com:443 -servername example.com </dev/null 2>/dev/null \
  | openssl x509 -noout -subject -issuer -dates

-servername is the part people leave off, and it sends the name in SNI. Without it, a host serving several names answers with whatever its default is, and you have just read the expiry of a certificate no visitor of yours will ever be shown. The same read from a standard library, which is what a monitor should be doing:

import datetime, socket, ssl

host = "example.com"
ctx = ssl.create_default_context()
with socket.create_connection((host, 443), timeout=10) as raw:
    with ctx.wrap_socket(raw, server_hostname=host) as tls:
        cert = tls.getpeercert()

until = datetime.datetime.strptime(cert["notAfter"], "%b %d %H:%M:%S %Y %Z")
print(cert["notAfter"], (until - datetime.datetime.utcnow()).days, "days left")

What notAfter actually is

RFC 5280 defines the validity period as notBefore through notAfter inclusive. A client that follows the rules refuses the certificate after that instant whether or not the server has noticed, and there is no grace period and no partial failure: the page does not load, and what the visitor sees is an interstitial about an insecure connection. To a non-technical reader that reads as this company has been attacked rather than a scheduled job failed.

Two details that cost an afternoon each. The clock that decides is the client's, so a device with a wrong date fails against a perfectly good certificate. And the chain has its own dates: an intermediate that expires is the same outage as a leaf that expires, and the leaf's notAfter says nothing about it.

The window is shrinking, on a published schedule

This is what turns an old habit into a defect. The CA/Browser Forum Baseline Requirements cap the validity period of a publicly trusted certificate, and the cap steps down on fixed dates. Read out of the requirements themselves rather than from memory:

Certificate issued on or afterMaximum validity period
Before 2026-03-15398 days
2026-03-15200 days
2027-03-15100 days
2029-03-1547 days

A yearly renewal written in somebody's calendar is already a thing that cannot work, and each step makes an unwatched automatic renewal fail sooner and more often. The direction of travel is automation that is verified, not automation that is trusted.

Four ways the automation still fails

Every one of these has a renewal mechanism that is installed, configured, and was working last quarter:

  1. DNS moved. An http-01 challenge is validated against whatever the name resolves to now. Point the name at a new host, a CDN or a staging environment and the challenge is answered by something that has never heard of the token. The dns-01 variant fails one layer up, when the API credential belongs to a provider that is no longer authoritative for the zone.
  2. The challenge path stopped being reachable. A catch-all rewrite that serves the application for every path answers the challenge with a 200 and an HTML page, which is not the token. A WAF rule, basic auth added to a staging host, or a redirect into https before the first certificate exists all do the same thing.
  3. The renewal runs somewhere that no longer exists. The job was installed on a box that has since been replaced, or in an image rebuilt from a base that dropped the timer, or under a user whose credentials were rotated. Nothing fails loudly, because nothing runs at all.
  4. It renewed, and the server kept serving the old one. The file on disk is new and the process still holds the previous certificate in memory until something reloads it. This is the failure that looks fine on the host and wrong from outside, and it is why the measurement has to be taken through the network rather than with ls.

Monitor the artefact, not the job

Whether the renewal exited zero is the wrong question, because three of those four failures never run the renewal at all. The right question is the one a visitor asks: what certificate does this name actually serve, and how long has it got?

  • Check from somewhere else. A probe on the same machine shares its DNS, its clock and its fate.
  • Check every name, with SNI. The apex, the www, the API hostname, each alternative name on the certificate. A wildcard that covers them all still has to be the certificate actually presented on each one.
  • Two thresholds, not one. One that means somebody should look at the renewal this week, and one that means somebody looks now. A single alert a few days out leaves no room between a warning and an outage.
  • Alert somewhere a person will see it. The failure mode of certificate monitoring is a daily email that stopped being read, which is indistinguishable from no monitoring at all.

That is a few dozen lines of standard library and a scheduled job somewhere other than the web host. It is the cheapest monitoring on any site, and the only one whose absence a stranger can measure on your behalf.

Sources

Where this is written down

Related

Where this comes up in the work

More notes

Other things worth writing down

Next step

Tell us the version, the hardware, and what it has to do.

You will get a written scope and a fixed price against it. If the honest answer is that you do not need us, you will get that instead.