Note
What a browser does with http assets on an https page
The certificate is installed, the redirect is in place, and the page is quietly broken: a blocked stylesheet is an unstyled page, and a blocked script is a feature that silently does nothing. The only report is a console log no visitor will ever open.
- https
- Browsers
- Front end
Two classes, and neither of them is loud
The Mixed Content specification splits a page's subresources in two. Anything that can run code or restyle the document — script, stylesheet, iframe, fetch and XHR, web fonts, workers — is blockable: requested over plain http from an https page, it is not loaded at all. Images, video and audio were the historical exception, allowed through with a warning on the reasoning that a tampered image is a smaller problem than a tampered script.
That exception has been closing. Browsers now upgrade those requests to https rather than sending them in the clear, and where the upgrade cannot succeed — because the origin does not serve https — the image simply does not appear. So the two outcomes to expect are blocked and upgraded, or blocked. Neither of them is a warning anybody sees.
What it looks like to a visitor. Not an error page. An unstyled wall of text, a map that never draws, a form that does nothing when submitted, a font that falls back to something ugly. Each of those reads as this site is broken, and none of them appears in a server log, because a request the server never received cannot be logged.
Where they hide
- Absolute URLs in templates. Written when the site was http, and correct at the time.
- Content in a database. A CMS body field with an image source on http is invisible to any search of the repository, and there may be thousands of them.
- Inside stylesheets. A
url()in CSS makes a request exactly as an attribute does, and grepping the HTML will never see it. - In
srcsetrather thansrc. The fallback got fixed; the responsive set did not. - A form action. Not mixed content in the specification's sense, and worse in practice: the browser warns on submission, and the data crosses the network in the clear.
- A redirect chain that dips. The URL in your markup is https and the third hop is not. Only following the chain shows it.
- Metadata rather than requests. A canonical, an og:image or a JSON-LD url still on http costs nothing in the browser and quietly tells every crawler the page lives somewhere else.
Find them in the built output
Search what you publish rather than what you author: the defect is as likely to be in a template partial, a vendored library or a generated file as in a page somebody wrote.
grep -rnE '(src|href|action|poster|data|srcset)="http://' dist/
grep -rnE "url\(\s*['\"]?http://" dist/
Then throw the false positives away, because the first run is full of them and that is
how this check gets abandoned on day one. http://www.w3.org/2000/svg is an
XML namespace, not a request; so are the namespaces in a DOCTYPE, an xmlns
attribute, an RSS or Atom document, and a rel="profile" link. A namespace is
an identifier that happens to be spelled like a URL, and nothing fetches it. Matching on
the attributes that cause a request, as above, rather than on the string
http:// removes most of the noise by construction.
For the half that is not in the repository — database content, a third-party embed, a redirect that dips — let the browser do the finding. Load each page that matters with the console open and read the mixed-content entries: they name the exact URL and the document that asked for it.
The transitional net, and why it is not the fix
The upgrade-insecure-requests directive in a Content Security Policy tells
the browser to rewrite http subresource requests on your pages to https before sending
them. On a site whose assets are all available over https it turns the whole class of
defect off in one response header, and it is the right first move on a large estate whose
markup cannot be fixed in an afternoon.
It is a net rather than a repair, for three reasons. The upgrade only works where the origin actually serves https, so an asset on a host that does not is still gone. The wrong URL is still in the source, so the next template copied from it is wrong too. And it does nothing for the metadata cases above, which are not requests at all. Ship the header, then fix the markup, then keep the grep in CI so it cannot come back.
Sources
Where this is written down
- Mixed Content, W3C WebAppSec editor's draft — the two classes of mixed content and what a user agent must do with each, including why images were ever treated differently from scripts
- Upgrade Insecure Requests, W3C WebAppSec editor's draft — what the upgrade-insecure-requests directive rewrites, and that it is specified as a transition mechanism rather than as a fix for the markup
Related
Where this comes up in the work
Web, apps & DevOps
Web and app builds with the DevOps around them: front end, backend, the database under it, and the build, deploy and monitoring work that keeps it up.
More notes
Other things worth writing down
The http redirect a static host does not add for you
A host that serves https does not always redirect plain http to it. What a plain-http request really does, and the one-line fix on four common hosts.
- https
- Static hosting
- Deployment
A link check that runs before a visitor finds the 404
A link check over your own built output, in CI: the four classes of break it finds, and why it works locally and it worked last deploy both miss them.
- Deployment
- Front end
- Testing
The one line that decides how a phone renders a page
With no viewport meta tag a phone lays a page out at desktop width, scales it down, and matches no media query. What the tag changes, and what not to add.
- Mobile
- Front end
- CSS
Next step
Tell us the version, the hardware, and what it has to do.
You will get a written scope and a fixed price against it. If the honest answer is that you do not need us, you will get that instead.