A catch-all route does not only rescue the URLs your router owns. It answers every URL anyone will ever request, always with a success code, and the tools meant to catch problems believe it.
One of the first checks I run after deploying any site is a request for a file that must never be
served: /.env, /.git/config, the build configuration. On a host with a catch-all, the answer is a clean 200 on every one of them, and the first time you see it
you assume the worst. Usually nothing has leaked. The response body is the site's own home page. The host
has simply decided that any path it did not recognise should be answered with index.html and a
success code.
That behaviour has a name, the single-page-application fallback, and it exists for a good reason.
A client-side router owns URLs like /account/settings that do not exist as files, so
the server must hand back the application shell for them and let the router take over. Cloudflare
Pages does this by default: if a project has no top-level 404.html, it assumes you are
deploying a single-page application and serves the root page for anything unmatched. Plenty of other
hosts and framework adapters do the same, sometimes without saying so.
The trouble is that a fallback does not just rescue the routes you meant. It answers every URL anyone will ever type, and it answers them all with yes. A surprising number of the tools I rely on to tell me the site is healthy read that yes and believe it.
The probe for /.env is supposed to tell me whether a secret file is exposed. With a
catch-all in place, it returns 200 whether the file is there or not, so the status code carries no
information in either direction. Worse, a scanner that only reads status codes will report every
sensitive path as present, and after the first dozen false alarms people stop reading the report,
which is exactly the state in which a real exposure goes unnoticed.
The fix is not to stop checking. It is to check the body. For every path that must not be served,
my post-deploy audit fetches it and asserts that the body is not the file it is guarding against,
with a positive signal: it looks for a marker that only the application shell contains, and for the
shape of the forbidden file, such as a line of the form KEY=value or a
[core] section. A 200 that is the shell passes. A 200 that is anything else fails
loudly, and so does a 200 that is empty, because an empty body is not proof of anything either.
The deeper point is that a status code is a claim made by whatever sits in front of your files, and a catch-all makes that claim unconditionally. Anything that matters has to be settled by what came back, not by the number attached to it.
The same behaviour quietly damages a site in search. A typo in a link, an old URL from before a restructure, a crawler guessing at paths: every one of them comes back 200 with the home page's title and canonical tag. Search engines learn to call these soft 404s, a page that says it exists and plainly does not. At best they are ignored. At worst, dozens of distinct URLs appear to serve identical content, and the engine has to decide which of them is the real home page.
For a portfolio site whose whole job is to be understood correctly by a search engine, that is
the wrong noise to be making. The remedy is the smallest one available: ship a
real 404.html. On Pages its mere presence switches the fallback off, so unmatched paths
receive that page with a genuine 404 status. The page itself carries noindex, links
back to the sections that do exist, and has no canonical tag pointing anywhere, because a not-found
page should not be claiming to be anything.
I run a broken-link check against every site before a content run is considered finished. With a fallback in place, that check cannot find a broken internal link, because there is no such thing: every href resolves to a 200. The failure it hides is links to pages that were renamed months earlier, all reported healthy, all delivering the home page to anyone who clicks them.
The same body-over-status rule applies. My checker now treats a 200 whose body is the root page as a failure for any URL that is not the root, and it compares the canonical tag in the response against the URL that was requested. A page whose canonical disagrees with its own address is either a deliberate duplicate, which I can list explicitly, or a fallback in disguise.
Uptime checks have the same blind spot. A monitor that requests a deep page and alerts on a non-200 will stay green through a deploy that deleted that page, because the fallback serves the shell in its place. The monitor is still telling the truth about the edge: it is up. It has just stopped telling me anything about whether the site I meant to publish is the one being served.
For the checks I care about, I pin something specific to the content: a heading, a build marker, a string that exists only on that page. The monitor asserts the string is present. It is a small change, and it converts a check of the host into a check of the deploy.
None of this is an argument against client-side routing. Some of the products I build are single-page applications, and they need the shell served for their routes. The distinction I draw is between a fallback that is scoped and one that is global.
A global fallback says every path is an application route. A scoped one says that paths under
/app/ are application routes and everything else is a file or a 404. On Pages that
means a real 404.html plus explicit rewrite rules in _redirects for the
prefixes the router owns, each rewriting to the shell with a 200. The router still works, the
marketing pages and articles remain static files, and a mistyped path outside the application
returns the honest answer.
Inside the application the router must then do the job the server has delegated to it. A route the router does not recognise should render a not-found view rather than silently showing the dashboard, and if the application has any pages meant to be indexed, those are the ones that should not be client-rendered in the first place.
The audit that runs after every deploy of these sites has four assertions that exist only because of this:
The first assertion is the cheapest and does the most work. It takes one request, and it tells me whether I can trust any of the status codes that follow. I put it at the top of the audit for that reason: before believing what a server says about the paths I care about, I ask it about a path that does not exist, and see whether it is willing to say no.
Written by Liana Grigory, Entrepreneur and Software Engineer, from work on The Care Royal, Tegula Stone and Unified Savers. Everything above describes decisions actually made on those systems, including the ones that turned out to be wrong.