Unattended jobs get rerun, relaunched and retried as a matter of course. The only safe design is one where a second run, against the same state, finds its work already done and says so.
Several of my sites are updated by a job that runs on a schedule, unattended, while I am doing something else. The first version of that kind of job is always written for the happy path: it starts, it does its work, it finishes, and it runs again tomorrow. The schedule is assumed to be the thing that guarantees each piece of work happens once.
It is not. A scheduler guarantees that a job is started, roughly when you asked. It says nothing about how many times. A machine wakes from sleep and runs the job it missed, and then runs the one that was due. A run is killed halfway and relaunched. A retry policy fires because the first attempt took longer than a timeout, while the first attempt is still going. Someone, usually me, starts it by hand to check something, on a day it has already run. Each of those is ordinary, and each one turns a job that assumed it ran once into a job that does its work twice.
For a job that only reads, that costs a little time. For a job that publishes, it costs a second batch of pages nobody asked for, a duplicate entry in a list, a second email, or a second commit that undoes a careful first one. The property I now design every unattended job around is simple to state: running it a second time, on the same day, against the same state, should change nothing.
Idempotence is not a feature you add at the end. It falls out of one decision made at the start: what identifies a unit of work. If the job publishes articles, is the unit "today's run", "an article with this slug", or "an article on this topic"? Each answer gives a different guard, and the wrong answer gives a guard that passes when it should not.
"Today's run" is the natural first choice and the weakest. It needs somewhere to record that the run happened, which is state the job has to write and trust, and it breaks the moment a run is partial: half the work was done, the marker was never written, and the relaunch does the first half again. A marker written at the start has the opposite failure, where a crashed run prevents the rerun that would have finished it.
What works better is a natural key on the output itself. An article's slug is its address, so the question "has this already been published?" has a factual answer: does a file exist at that path, and does the index already link to it? The job does not need its own bookkeeping, because the thing it produces is the bookkeeping. If the path exists, that unit of work is done, whoever did it and whenever.
The guard has to read the current state of the thing it is about to change, not a cached belief about it. A job that keeps a local note saying "last published on this date" will be wrong as soon as anything else touches the same target: a manual edit, a different machine, an earlier run that succeeded after the note was written. A job that looks at the repository after fetching it, or at the live page after publishing, is checking something that cannot drift out from under it.
This is the same discipline I apply to verification after a deploy, turned around and applied before the work starts. Before writing anything, the job asks what is already there. If the answer is "today's batch already exists and is live", the correct behaviour is to verify that batch, report it, and stop. Publishing a second batch because the job was started again is not diligence. It is a duplicate.
A job that is idempotent as a whole can still have steps that are not, and a crash between them exposes the difference. The pattern I use is to order the steps so that each one either produces something the next step can detect, or is harmless to do twice.
The ordering matters as much as each step. If the list entry is added before the file is written, a crash leaves a link to nothing, and the rerun sees the entry and skips the file. If the file is written first, a crash leaves an orphan the rerun can detect and finish linking. The second ordering fails into a state the job knows how to repair. The first fails into a state that looks finished.
The useful mental shift was to stop treating an interrupted run as an exception. On a long enough timeline, every unattended job is interrupted: by a session limit, a laptop lid, a network that drops during a push, a tool that hangs. If the design only behaves correctly when a run completes, it behaves correctly most of the time, which for a job nobody watches is the same as sometimes being wrong without anyone noticing.
So the question I ask of each job is not "does it work?" but "what does the next run see after this one dies at each step, and does it do the right thing?" Walking through that once, step by step, usually finds one place where the answer is "it does the work again". That is the bug, and it is cheaper to find on paper than in a list with every entry doubled.
A job that silently does nothing on a rerun is correct and slightly alarming. From the outside, a run that found its work already finished looks the same as a run that failed to start. The report at the end has to say which of those happened: "today's three pages already exist, are linked from the index and the sitemap, and return the expected content; nothing written". That sentence is the difference between a quiet success and a silence you have to investigate.
It also catches the case idempotence cannot fix on its own. If the guard says the work exists but verification says the live page is wrong, the job has found a real problem from an earlier run, and the right response is to repair that, not to stop because the box was ticked.
Very little. A few existence checks before writing, a lookup before appending, a diff before committing. None of it is clever, and all of it is cheaper than the alternative, which is a person reading through a repository to work out which of two near-identical batches to keep. The schedule decides when a job starts. Only the job can decide whether there is anything left to do, and it should check every time.
Written by Liana Grigory, Entrepreneur and Software Engineer, from work on The Care Royal, Tegula Stone and Unified Savers. Everything above describes decisions actually made on those systems, including the ones that turned out to be wrong.