A change can be correct as a diff and broken as a rollout. Rules, data and code deploy separately, so the intermediate states are real, and one of them is always the one that locks your users out.
The change was four lines. A contractor should only be able to read a job once their
subscription is active, so the rule that had said request.auth != null now also had to
check that a mirror document existed marking them as paid. A function already wrote that mirror on
every subscription event. Deploy the rule, done.
It is not done, because those four lines are not one change. They are three changes that land at different times, and the order they land in determines whether the system is briefly wrong, briefly insecure, or neither.
A typical Firebase and Cloudflare stack has at least three independently deployable things, and almost nothing forces them to be consistent with each other.
A code review reads all of this as one state. Production experiences it as a sequence. The question a reviewer should be asking is not only is this correct, but is every intermediate state survivable.
For the subscription change, there are two ways to sequence it.
Rules first. The moment the rule is live, every request is checked against a mirror document that mostly does not exist yet, because the function that writes it has not shipped and no backfill has run. Every existing paying contractor is locked out of their own jobs. Nothing is insecure; everything is broken. Support finds out before monitoring does.
Rules last. Ship the function, run the backfill, confirm the mirror is complete, then ship the rule. During the gap the old rule is still in force, which means access is looser than you intend for a few minutes. Nothing new is exposed that was not already exposed yesterday.
The second is right, and the reason is worth stating precisely rather than as a rule of thumb. The intermediate state of a tightening change should be the state you are already in, not the state you are heading to. You are moving from a weaker rule to a stronger one; the correct interim is the weaker rule, which is a risk you have already been carrying, as opposed to the stronger rule evaluated against data that has not caught up, which is an outage you are choosing to create.
Loosening changes invert it. If a rule is being relaxed and code depends on the new access, the rule goes first, because the interim is again the state you are already in: restricted. Code that tries to do the new thing fails, which is annoying, rather than users seeing something they should not, which is not recoverable by rolling back.
Stated once: the safe order is whichever one makes the gap look like the past.
The rule is a minute. The function is a few minutes. The backfill is where the work is, and it is the step that gets underestimated because it feels like a chore rather than a change.
Things I now treat as non-negotiable for a backfill:
That last point is the one that has bitten me hardest. A backfill written as a one-off script is a second implementation of a write path, and second implementations drift. The better shape is for the backfill to call the same function the live path calls, so there is one definition of what the document looks like.
Everything above treats the browser as if it updates when you deploy. It does not. Someone has had a tab open since yesterday, someone else has a service worker serving a build from last week, and a mobile app shipped through a store has a version distribution rather than a version.
The consequence is that every rule change is also a compatibility question with clients you have already shipped. If the new rule requires a field the old client does not send, the old client starts failing writes, and the user has no idea why. The failure is silent, remote, and correlated with exactly the people least likely to report it.
The practical pattern is the same one used for database migrations: make the field optional first, ship the client that writes it, wait, then make it required. Two deploys and a wait, where one deploy looked sufficient. The wait is not caution, it is the actual mechanism.
For apps distributed through stores the wait is measured in weeks, not minutes, and some proportion of users will never update. That is a reason to be extremely reluctant about rules that require anything new from a native client, and a reason to build a version floor into the system early, while doing so is still cheap.
Not a ceremony — six lines in the pull request, which take longer to think about than to type.
That last line is the one that has changed how I sequence things. Rules roll back cleanly: they are a single artefact with a previous version. Data does not. A migration that rewrites documents in place has no previous version unless you made one, which is why the safe shape is to write a new field alongside the old one and stop reading the old one later, rather than to transform in place and hope.
The same reasoning applies well below the level of a migration, and this is where it earns its keep day to day.
Adding a field the rules validate: ship the rule that allows the field before the client that sends it, never after. Removing a field: stop sending it, wait for the clients to turn over, then stop allowing it. Renaming anything at all: do not — add the new name, write both, read the new, drop the old months later. A rename is a delete and a create wearing one commit, and the documents that already exist do not know that.
Splitting a monolithic deploy into several is usually the right instinct, with one exception worth naming: on platforms that bill per build, ten small deploys cost ten builds. The resolution is to batch the artefacts that are safe to land together and separate only the ones whose order matters. Order is a correctness property; deploy count is a cost property. Optimise them separately.
Configuration, schema and code are one system, and only one of them has a version control history that anybody looks at. Rules live in a repository but take effect the moment they are published. Data has no version at all. The client has many versions at once.
So the diff is not the change. The change is the diff plus an ordering plus a wait, and the ordering is a design decision with a right answer that is derivable rather than a matter of taste: make the gap look like the past. Write the order down in the pull request, because the person deploying at eleven at night is going to be you, and you will not re-derive it then.
Written by Liana Grigory, Entrepreneur and Software Engineer, from work on The Care Royal, Tegula Stone and Unified Savers. Everything above describes decisions actually made on those systems, including the ones that turned out to be wrong.