# A link check that runs before a visitor finds the 404

> A link check over your own built output, in CI: the four classes of break it finds, and why it works locally and it worked last deploy both miss them.

Source: https://plantroomlabs.com/notes/catching-a-broken-link-before-a-visitor-does/  
Published: 2026-10-02 (2 October 2026) · Plantroom Labs  
Topics: Deployment, Front end, Testing

An internal link that answers 404 is the one defect where both ends are in your own repository. It needs no crawler, no third-party service and, for the internal half, no network at all: it is a property of the build, and a build can refuse to ship.

## Why the local check passes

- **The development server answers everything.** A history fallback that serves the application shell for any unmatched path is exactly what a single-page application needs while you work, and it means nothing can 404 in front of you. The CDN serving the built output has no such fallback.
- **The laptop's filesystem is case-insensitive.** macOS and Windows will happily open `Notes/Index.html` when the link said `notes/`. Almost every production host is Linux, where that is a 404. This is the most common break that passes every local test.
- **Trailing slashes are host policy.** Whether `/notes` serves `/notes/index.html`, redirects to `/notes/` or answers 404 is a setting, and it differs between a development server, a staging host and a CDN.

## Why the last deploy passing proves nothing

Because the break is almost never in the page that changed. Rename one slug and every page that linked to it is now wrong — and none of those pages appear in the diff. On any site whose navigation, indexes, related links and sitemap are generated from data rather than typed by hand, one edit moves dozens of links at once. That is precisely the property which makes a generator worth having, and which makes reviewing a diff useless as a link check.

This site is the worked example. Every internal link on it is emitted by one script from one table of pages, and a second script reads the generated HTML — not the generator's intentions — and refuses the publish when an href names something the build did not produce. The two scripts disagreeing is the only signal worth having, and it is the reason the check has to run over the output.

## The four classes a check should find

1. **A path the build does not produce.** Collect every `href` and `src` in the output, resolve each against the tree on disk with the host's index and trailing-slash rules applied, and assert the file exists.
2. **A fragment nothing carries.** A link ending in `#section-3` where no element has that id lands silently at the top of the page. Nothing reports it: not the server, not the browser, and not the visitor, who assumes they misread the heading. Same pass, one extra set of ids.
3. **An asset referenced but never copied.** A stylesheet, a font, an image, a social card whose build step was removed six months ago. The page still renders, slightly wrong, and the 404 is in a log nobody greps.
4. **An external link that has rotted.** A different budget: the network is involved, so it is slow, flaky and rate-limited. Run it on a schedule rather than on every commit, and record each answer in a file in the repository, so a document moving is a diff somebody reviews rather than a surprise. A dead citation is worse than no citation, because it turns evidence into a claim that something used to exist.

## Two things that look like fixes

**A redirect is not a repair.** Keeping the old path alive after a rename is the right thing to do for inbound links you do not control, and it does nothing for the links you do: every internal hop now costs a round trip, crawlers re-crawl slowly, and the stale href stays in the template until somebody renames that page too. Add the redirect, then fix the links.

**A pretty error page that returns 200 is worse than a 404.** A host configured to serve a styled not-found document with a success status has made every broken link both indexable and invisible: a checker sees a 200, a crawler sees a thin page it may index, and nothing anywhere reports a failure. Assert the status code rather than the body, and keep the error document out of the sitemap, which is a list of URLs that exist and are meant to be indexed.

## What it costs

For internal links: a walk of the output directory, two regular expressions and a set lookup. No network, no dependency, and it finishes faster than the build it follows. For external links: one request per distinct URL, on a schedule, with the answer committed. The reason to run it before the publish rather than after is not thoroughness — it is that a build which refuses costs ten minutes, and a visitor who finds the 404 is a visitor deciding the site is not maintained.

## Where this is written down

- [RFC 9110: HTTP Semantics](https://www.rfc-editor.org/info/rfc9110/) — what a 404 means and what a 410 means, and why answering a not-found document with a success status is a contradiction rather than a style choice
- [Sitemaps XML format and protocol, version 0.9](https://www.sitemaps.org/protocol.html) — a sitemap is a list of URLs that exist and are intended for indexing, which is what makes an error document inside one a defect rather than an oversight
