Publishing is not the same as being reachable. A page that ships successfully and is linked from nowhere has all the costs of existing and none of the benefits.
While working on my own site I found a directory of six long-form technical pages. They were finished. They were deployed. They returned two hundred. They had canonical tags, structured data and meta descriptions. They were also linked from nothing, listed in no sitemap, and mentioned in no index anywhere on the site.
They had been live for weeks. As far as any search engine was concerned, they did not exist, and as far as any visitor was concerned there was no path that led to them. There was even an index page listing all six, which was itself linked from nowhere.
This is a specific and under-discussed failure, and it is worth writing down because everything about it looks fine from the inside.
Consider what the usual checks actually assert.
A build confirms the files were produced. Orphaned pages build perfectly; they are ordinary pages. A deploy confirms the upload succeeded. It succeeds. A live check confirms the URL returns a two hundred, which it does, because you fetched the URL directly. A link checker traverses outward from pages it already knows about, so it never reaches a page nothing points at, and it therefore reports no broken links, which is true and useless.
Every signal is green. The failure is not in any page; it is in the absence of an edge in a graph, and almost no tooling is looking at the graph.
In my case the cause was mundane. The site had two content systems that had grown up at different times, one under one directory and one under another. The older one was wired into the sitemap, the homepage and the machine-readable summary file. The newer one was created later, with better pages, and the wiring step was never done, because the wiring lives in files that are nowhere near the content.
That is the general shape. Publishing a page touches one place. Making it reachable touches several: a sitemap, an index or hub page, whatever navigation exists, and any secondary manifest. The first step is the one that feels like the work. The others are bookkeeping, they live in unrelated files, and they are exactly what gets skipped when a session runs long.
A second contributor, which I also hit: I checked whether a page was listed using a pattern that excluded digits, and the slug contained a number. The page was listed. My check could not see it, and I nearly "fixed" a problem that did not exist by adding a duplicate entry. A verification whose failure mode is a false negative is worse than no verification, because it produces confident action in the wrong direction.
All of these are cheap and none of them require a service.
Compare the file tree to the sitemap, both directions. Enumerate the pages that exist on disk, enumerate the URLs in the sitemap, and print both differences. Pages missing from the sitemap are orphans. Sitemap entries with no page are dead URLs being advertised to crawlers, which is the same bug wearing the opposite mask. One symmetric check finds both.
Count inbound links per page. Walk every internal link on the site, build the map, and list every page with zero inbound edges. Excluding the deliberate exceptions, which are few and which you can name, that list should be empty.
Parse the sitemap as XML, do not grep it. Parsing catches malformed entries and lets you assert on structure rather than on substrings. It also makes duplicate detection trivial, and duplicates are the specific thing a careless insertion produces.
Make the publishing step atomic in intent. If creating a page and registering it are separate actions, they will eventually come apart. Either generate the sitemap and the index from the content directory, so they cannot disagree, or make the check that compares them a gate that fails the build.
The generated version is strictly better where it is available. A sitemap derived from the file tree is not a thing that can drift, because there is nothing to keep in sync. Hand-maintained lists of URLs are a class of file that is always slightly wrong.
The reason I think this is worth more than a checklist item is that it is an instance of a broader mistake: verifying the artefact instead of the property you actually want.
I wanted these pages to be found. I verified that they existed. Those are different claims, and the first one implies nothing about the second. The same gap shows up everywhere once you look for it. A security rule that deploys is not a security rule that denies; you have to make the request that should fail and watch it fail. A backup that runs is not a backup that restores. A monitor that is configured is not a monitor that pages anybody.
In each case the cheap check confirms the thing was created and the expensive check confirms it does its job, and it is always the cheap check that gets automated. The discipline is to write down what you actually wanted in plain language, then ask whether your check would fail if that thing were false. For an orphaned page, the honest sentence is "a crawler starting from the homepage can reach this", and no test I had would have failed if that were untrue.
Added every one of them to the sitemap, then verified by parsing the XML and diffing it against the directory listing rather than by looking at the file. It should have been part of publishing them. That is the point: it always should have been, and it never is, unless something checks.
Written by Liana Grigory, Entrepreneur and Software Engineer, from work on The Care Royal, Tegula Stone and Unified Savers. Everything above describes decisions actually made on those systems, including the ones that turned out to be wrong.