← Liana Grigory
Insights

A deploy is not done when the upload succeeds

Exit code zero is not proof and a green deploy is not proof. The verification rules I ended up with for edge-deployed sites, and the specific wrong reports that produced each one.

By Liana Grigory · 1 September 2026 · 9 min read

The most expensive sentence in my own operating notes is one I had to write after getting it wrong: exit code zero is not proof, and a green deploy is not proof.

Both feel like proof. A deploy that completes without error produces the same emotional signal as a passing test, and it is a much weaker claim. All it establishes is that an upload succeeded. It says nothing about whether the thing you changed does what you changed it to do, whether the paths beside it still work, or whether the platform is currently serving the build you think it is.

What follows is the verification discipline I ended up with for edge-deployed sites, and the specific mistakes that produced each rule. None of them are theoretical. Every one is a thing I reported confidently and incorrectly first.

Refresh cached state before comparing against it

This is the one that embarrassed me most, and it has nothing to do with deploys.

I reported that three repositories were holding unpushed work. The report was wrong. Git's "ahead/behind" counts, and anything derived from @{u}, read a local remote-tracking pointer. That pointer is only as current as your last fetch. Mine was months stale. One of those repositories was in fact zero ahead and two hundred and twelve behind, and its "unpushed" commits were already upstream.

The fix is one word — git fetch — placed before the comparison rather than after the conclusion. The general form is worth extracting though, because it applies far beyond git: any number you read from a cache is a fact about the cache, not about the world. Build outputs, CDN responses, DNS, a package lockfile, a dashboard with a refresh interval. If a decision depends on the number, refresh the source first.

Print the full identifier, never the convenient one

In the same report I described a directory as ~/Desktop/ios-app. That path does not exist. The real one was two levels deeper, and I had printed a basename from a find result and mentally attached it to the wrong parent.

A basename is a partial identifier, and partial identifiers silently relocate things. So: absolute paths, full URLs, full refs, full identifier prefixes. Not because it reads better — it reads worse — but because a truncated identifier is a claim that cannot be checked by the person reading it, and it transmits your assumption as though it were an observation.

Confirm the target exists before describing its state

A configured remote is not a live remote. A DNS record is not a served endpoint. A binding declared in configuration is not a binding attached to the running deployment.

The cheap version of this check is one command: resolve the remote, fetch the HTTP status, ask the platform what is actually deployed. It costs a second and it prevents the specific failure where you carefully describe the state of something that is not there, and only discover it when an action against it fails.

Send a browser User-Agent, and bust the cache

Two separate traps that both produce the same false result.

The first: many origins, mine included, refuse or differently handle requests from default tooling user agents. A verification request that gets challenged is not a verification. You are not measuring your site; you are measuring your bot protection. So the check goes out with a real browser User-Agent, which is also more honest, because it is the response an actual visitor gets.

The second: you are behind a cache, deliberately, because that is the entire point of the architecture. A request immediately after a deploy can be answered from an edge that has not yet seen the new build. My deploys flap for something like thirty to ninety seconds, which is long enough to read the old build, conclude the deploy did not work, and start debugging a problem that does not exist.

So: a cache-busting query parameter, and a re-check after the propagation window rather than during it. And when a result looks wrong immediately after shipping, the first hypothesis is propagation, not regression.

Trigger every control you added, at least once

This is the rule that catches the most real defects, and it is the one most often skipped, because it is the only one that requires doing something rather than reading something.

A rate limit that has never returned a 429 is a rate limit you believe in. A form handler that has never rejected a bad payload has untested validation. A security header configured in a file is not a security header until a response carries it. A permission denial you have never seen deny anything is a hypothesis.

The reason this matters more than ordinary testing is that these controls fail open. A broken feature is visible because someone cannot use it. A broken control is invisible, because the symptom of a rate limit that does not work is that everything is permitted, which looks exactly like everything working. Nobody files a ticket saying the abuse protection let them through.

A specific instance, since I would rather be concrete than sound rigorous: I shipped two rate limits in one session where the arithmetic was wrong, each sitting directly beneath a comment asserting it was fine. The comment was the problem. It let me read the intent and skip the calculation. What fixed it was mechanical — compute the effective rate, print the number, compare it against the threshold it is supposed to enforce, and trip it once for real. Any rate you write down should be a rate you have printed.

Separate what you measured from what you inferred

The single highest-leverage habit in this list, and it is a writing habit rather than an engineering one.

When reporting, mark each claim as measured or inferred. "The endpoint returns 200 with a browser User-Agent" is measured. "So the deploy succeeded" is inferred, and it is a decent inference that is nonetheless wrong if the platform is serving a previous build. "The rule denies cross-tenant reads" is inferred from reading the rule; it becomes measured when you attempt a cross-tenant read and watch it fail.

The forcing function I use: if a sentence is inferred, either name the command that would settle it, or run that command instead of writing the sentence. Most of the time it is faster to run the command, which is the point.

Getting this wrong is worse than not knowing, because a confident wrong answer gets acted on. An honest "unconfirmed, here is how to confirm it" costs a minute of someone's time. A confident false claim costs however long it takes them to discover it, plus whatever they built on top of it.

When a number decides something, verify it twice by different means

Not the same check run twice — two methods that would fail differently.

Whether a change is deployed: the platform's record of the deployment, and the content of the live response. Whether a rule is in force: reading the rule, and making a request that the rule should reject. Whether a secret is absent from a repository: the current tree, and the history.

Two agreeing methods is a reasonable standard for a number someone will act on. One method is a guess with a citation. And if the two disagree, that disagreement is the most valuable thing you have learned, so it should be reported rather than resolved by picking the more convenient one.

Make the override loud

A theme that recurs across all of this: the dangerous failures are the quiet ones.

A verification script that stops matching anything and keeps exiting zero. A post-build correction that no longer applies because an upstream format changed. A check that was skipped because a condition silently evaluated false. None of these announce themselves, and all of them degrade into a default that looks acceptable.

So anything that enforces something should say what it did, with numbers, every time. Not "OK" — the before and after values. A step that prints 100 rules -> 2 tells you it worked; the same step printing 2 rules -> 2 after a dependency upgrade tells you the world moved underneath it. "OK" tells you nothing in either case.

What this is really about

None of these rules are sophisticated. They are: fetch before comparing, print the whole identifier, check the thing exists, use a real User-Agent, bust the cache, trigger the control, label inferences, verify twice, print numbers.

What they have in common is that each one closes a gap between something I observed and something I concluded. That gap is where every incorrect report I have written came from — not from insufficient care, but from care applied to the wrong step. I was careful about the change and casual about the verification, and then reported the verification with the confidence I had earned on the change.

The discipline that actually helps is smaller than it sounds: before asserting a problem, run the one command that would disprove it. It is nearly always cheap, it is nearly always available, and the times it contradicts you are worth more than all the times it does not.

Written by Liana Grigory, Entrepreneur and Software Engineer, from work on The Care Royal, Tegula Stone and Unified Savers. Everything above describes decisions actually made on those systems, including the ones that turned out to be wrong.

Home · Terms of Use · Privacy Notice