← Liana Grigory
Insights

Custom claims are the only authorization you have, and they are always slightly stale

On Firebase, claims make authorisation unforgeable and free to check. They are also a snapshot that can be an hour out of date, which makes granting a support ticket and revoking a security problem.

By Liana Grigory · 1 September 2026 · 10 min read

Every multi-tenant app on Firebase eventually arrives at the same question, usually late: where does the sentence "this user is an admin of tenant X" actually live?

It cannot live in the client, because the client is a program on someone else's computer. It cannot usefully live in a Firestore document that the client reads and then trusts, because that is the same thing with extra steps. The answer on Firebase is custom claims: key-value pairs the Admin SDK writes onto a user, which arrive inside the signed ID token, and which Firestore security rules can read without performing a document lookup.

That is the right answer, and it took me a while to understand that it comes with a property nobody mentions in the first tutorial you read: a claim is a snapshot, not a subscription. It is minted into a token, and that token keeps being valid, and true, and stale, for up to an hour. Every real bug I have had in this area has been a variation on that one sentence.

What claims buy you, concretely

Two things, and the second is the one that matters more than people expect.

They are unforgeable. The token is signed. A user can read their own claims, decode them, and see exactly what they say. They cannot change them, and a modified token fails signature verification. This is what makes it safe for the rules engine to trust them.

They are free to read inside a security rule. This is the part that changes architecture rather than just correctness. A rule can check request.auth.token.tenantId == resource.data.tenantId without reading a document. Compare that with the alternative shape, where every rule performs a get() against a memberships collection to find out who the caller is. That get() is a billed read, on every operation, forever, and it multiplies: a rule that does a lookup to authorise a query has to do it for the query, and your read count for a screen becomes a function of your authorisation model rather than of your data.

So putting the tenant and the role in the token is not a micro-optimisation. It is the difference between an authorisation check that costs nothing and one that scales with traffic. On a product where the goal is that the only metered line growing is the account count, an authorisation model that bills per check is disqualifying on its own.

The trusted path, and keeping it narrow

Only the Admin SDK can set a claim. That is the whole security model, so the interesting question is what is allowed to run the Admin SDK.

The rule I hold to: there is exactly one code path that grants a role, and it is not reachable by the person being granted it. Not one per feature. One.

Concretely, that means the grant function verifies the caller's own ID token, checks that the caller already holds an authority sufficient to make this grant, checks that the grant is within the caller's own tenant, validates that the requested role is a member of a fixed allowlist, and only then writes the claim. Every one of those checks has a failure mode I have actually seen or nearly shipped.

The allowlist matters more than it sounds. If the role string arrives from a request body and is written through without validation, a caller can invent one. Sometimes inventing a role is harmless because no rule mentions it. Sometimes you add a rule six months later that mentions superadmin, and there is a user out there who has held that claim since March because nothing rejected it. Validate against a literal set, and fail closed on anything else.

The tenant check matters for the same reason in a different direction. A legitimate tenant administrator making a legitimate grant, with a tenant identifier taken from the request instead of from their own token, is a cross-tenant privilege escalation with no attacker sophistication required. The tenant should come from the caller's verified claims, never from the payload.

The staleness problem, properly

Here is the behaviour that produces the bug reports.

An ID token is short-lived and refreshed automatically, roughly hourly. When you set a claim, you change the source of truth for future tokens. You do not change the token the user is currently holding. So for up to an hour after you promote someone, their app is running on a token that says they are not promoted.

Two symptoms, and they have opposite severities.

Granting is a support ticket. An administrator adds a colleague, the colleague reloads, and nothing has changed. They log out and back in, and it works. This is annoying, it makes your product feel broken, and it generates the class of ticket where the resolution is "try again later," which is corrosive.

Revoking is a security problem. An administrator removes someone's access. That person's existing token continues to satisfy your rules until it expires. You have not revoked anything yet; you have scheduled a revocation for some point in the next hour.

That asymmetry deserves to be stated plainly, because the mitigation for the first is a convenience and the mitigation for the second is not optional.

What to do about each

For granting, force a token refresh at the moment it matters. The client SDK can request a fresh token explicitly, and the reliable pattern is: the grant function returns success, the client refreshes its token, and only then re-renders anything that depends on the new authority. Doing this on a signal rather than on a timer is the difference between an app that feels immediate and one that feels unreliable.

The trap is doing it eagerly everywhere. Forcing a refresh on every navigation, or polling for claim changes, turns a rare event into constant traffic and reintroduces the ambient cost you adopted claims to avoid. Refresh on the event that changed the claims, not on a schedule.

For revocation, accept that the token cannot be recalled and design around it. Firebase provides token revocation, and the rules engine can be made to respect it, but the honest version of this is: if immediate revocation matters for an operation, that operation needs a check that is not purely token-based.

Which means being deliberate about tiers. For ordinary reads and writes, a claim-based rule with up-to-an-hour staleness is a reasonable trade, and it is what keeps the read cost at zero. For the small set of operations where an hour of retained access is genuinely unacceptable — billing changes, data export, deleting a tenant, granting further roles — the check goes through a verified path that can consult current state, and the cost of that lookup is acceptable precisely because those operations are rare.

I got this ordering wrong initially, in the way that is easy to get wrong: I treated all authorisation as one problem with one mechanism. The result was either a system that billed a lookup on every read, or a system where revocation was slower than it should have been for the operations that mattered most. The resolution was not a better mechanism. It was noticing that "how quickly must this become false?" is a different question per operation, and answering it per operation.

Size, and the thing that will bite you at scale

Claims live in a token that travels on requests, and there is a size limit — a modest one, around a kilobyte. This is fine until someone belongs to many tenants.

The shape that fails is the obvious one: a map of every tenant the user belongs to, with their role in each. It works beautifully for the first several tenants and then stops working, and it stops working for your most engaged users, which is the worst possible population to break.

The shape that survives is to keep the token small and let it point at the truth rather than contain all of it. The token carries the active tenant and the role in that tenant. Switching tenants is an explicit operation that mints a new token. Membership across many tenants lives in Firestore, gated by its own rules, read when the user opens the tenant switcher and not on every request.

This is a better model even where size is not yet a constraint, because it makes the active tenant an explicit piece of state rather than something inferred. A rule that compares against a single tenantId claim is trivially auditable. A rule that has to index into a map of memberships is where cross-tenant bugs hide.

Claims are not the whole rule

The failure I want to warn about most is the one where claims work so well that they become the only check.

A claim answers who the caller is. It does not answer whether this specific document belongs to that tenant, whether the fields being written are ones a client may write, or whether server-owned values are being tampered with. A rule that checks the claim and stops has authenticated the caller and authorised nothing.

So the rule still has to compare the claim against the document's own tenant field, still has to make server-owned fields immutable from the client, still has to validate types and sizes on write, and still has to scope list queries — because a rule that permits a read of a document the caller is entitled to does not, by itself, prevent a query that asks for documents they are not.

Claims made my authorisation cheap and unforgeable. They did not make it complete. The mechanism that says who you are is not the mechanism that says what you may touch, and conflating the two is how a system with correct-looking rules turns out to have none.

Written by Liana Grigory, Entrepreneur and Software Engineer, from work on The Care Royal, Tegula Stone and Unified Savers. Everything above describes decisions actually made on those systems, including the ones that turned out to be wrong.

Home · Terms of Use · Privacy Notice