Machine identities look like accounts and behave like dependencies. Nothing lists what breaks when one goes away, the breakage is delayed and uncorrelated, and the recovery window is finite.
Tidying up a cloud project is one of those tasks that feels unambiguously virtuous. Fewer identities, less surface, a shorter list. The identities in question had unmemorable generated names, no description, and no visible attachment to anything.
Deleting them produced no error, no warning and no immediate symptom. Over the following hours, several apparently unrelated things stopped: transactional email, a deploy, a scheduled job. Nothing announced a common cause, because from each system's point of view there was no common cause. There was just a credential that had stopped working.
The lesson is not "be careful". It is that a service account is not an account, it is an edge in a dependency graph, and the tooling presents it as the former.
A user account has an obvious owner who notices immediately. A service account's dependents are scattered across systems that do not know about each other: a function's runtime identity, a workload the platform impersonates, a key embedded in a deploy pipeline, a domain-wide delegation grant sitting in an entirely separate administrative console, a build step, a scheduler.
None of those register themselves anywhere that shows up next to the delete button. The console will happily tell you the roles an identity holds. It will not tell you what uses it, because that information does not exist in one place.
The delay compounds it. A credential that is already issued keeps working until it expires, so the first failure can arrive an hour later, and the failures arrive separately. By the time there are three symptoms, they look like three incidents.
The genuinely useful fact, which I did not know at the time and now treat as load-bearing: on Google Cloud a deleted service account is recoverable for thirty days, and restoring it preserves the identity itself, including the numeric identifier that domain-wide delegation grants are bound to. That last detail is what makes restoration meaningfully better than recreating something with the same name: a new account with an identical email address is a different principal, and every grant attached to the old one still refers to the old one.
Two practical points about doing it.
First, the identifier you need is still in the policy. Bindings that referenced the deleted identity do not vanish; they are rewritten to a form that marks the principal as deleted and carries its unique id. Reading the project's policy is therefore both the way to find out what the account had access to and the way to recover the identifier required to undelete it. The wreckage is the documentation.
Second, restoration does not restore keys. Any key file that was issued is permanently dead, and anything holding one needs a new key or, much better, a way of authenticating that does not involve a key file at all.
Thirty days is generous and it is also finite. It is enough time to notice, if anybody is looking. It is not enough time if the thing that broke is a quarterly job.
One identity, one purpose, and say the purpose in the name. Not
service-account-1. A name that says which system uses it and for what, plus a
description field that says the same thing in a sentence. The name is the only documentation that
travels with the object, and it is what a future person reads immediately before deciding whether the
thing is safe to remove.
Least privilege as blast radius, not as compliance. The usual argument for narrow scopes is what an attacker could do with a stolen credential. The argument I find more persuasive day to day is what I can do by accident. An identity that sends email and nothing else can only break email. An identity that accumulated six roles because each was needed once breaks six things, and those six things are exactly the ones nobody will connect.
Prefer identities you cannot delete by hand. Workload identity, platform-managed runtime identities, short-lived tokens — anything where the credential is issued at run time rather than sitting in a file. This removes the key-rotation problem and the key-in-a-repository problem at the same time, and it means fewer objects in a list for somebody to tidy.
Write down what depends on what. A short file in the repository listing each identity, what uses it, and what breaks if it disappears. It is not elegant, it goes stale, and it is still the only artefact that answers the question at the moment the question is being asked. A stale list that names four of the five dependents is enormously better than no list.
Deprecate before deleting. If an identity looks unused, remove its roles and wait, rather than deleting it. Removing a role is reversible in seconds and produces the same signal: if something breaks, you have found the dependent and can put it straight back. Deletion converts a reversible experiment into a thirty-day clock.
Machine identities belong to a category of infrastructure that shares three properties, and the combination is what makes them dangerous.
DNS records have this shape. So do API keys with referrer restrictions, storage bucket names, webhook endpoints, and any identifier that was chosen once and then quietly became a contract. The common treatment is the same: name it for its purpose, record what depends on it, scope it so its failure is narrow, and make removal a two-step process where the first step is reversible.
The instinct to tidy is a good one. It just needs to run against a list of dependents rather than a list of objects, and in most systems nobody has written the first list, because writing it is the sort of work that only ever pays off on a day you were not expecting.
Written by Liana Grigory, Entrepreneur and Software Engineer, from work on The Care Royal, Tegula Stone and Unified Savers. Everything above describes decisions actually made on those systems, including the ones that turned out to be wrong.