← Liana Grigory
Insights

The scheduled job you cannot see fail

A generator of mine ran nightly for five weeks and produced nothing, with every component working correctly. The symptom of a scheduled job failing is absence, and absence looks exactly like a quiet day.

By Liana Grigory · 24 September 2026 · 8 min read

An article generator on one of my sites ran on a schedule every night for about five weeks and produced nothing. No alert fired. The scheduler reported successful invocations the whole time. The way I found out was noticing, by eye, that the site had not gained a page in a while.

The cause was mundane: the function called a paid third-party API, the key had been revoked, and the call threw. What interests me is not the cause. It is that a system with logging, a scheduler and a cost alarm attached to it ran broken for over a month, and every one of those mechanisms was working correctly at the time.

Scheduled work fails differently

A request-driven failure has a witness. Somebody clicked something, got a spinner or a 500, and either retried or complained. The complaint is the monitoring, and it is better than most monitoring because it is attached to a human who cares.

A scheduled job has no witness. Nobody is waiting on this particular run. The output is consumed later, by someone who does not know when it was produced, and who has no baseline for how much of it there should be.

Worse, the symptom of failure is absence, and absence is indistinguishable from a quiet day. Zero articles published because the API key was revoked looks exactly like zero articles published because there was nothing left in the queue. Both are "the job ran and wrote nothing". Only one of them is a fault, and no ordinary log line distinguishes them.

The three ways they stop

Credential expiry. A key gets rotated, revoked or simply expires. Tokens have lifetimes, service accounts get deleted during a tidy-up, an OAuth grant is withdrawn. This is the most common cause I have seen, and it is uncorrelated in time with anything anyone was doing, which is what makes it hard to attribute later.

A limit. A quota, a rate ceiling, a budget cap, a free-tier boundary. These often present as a partial failure — the first few items succeed, the rest do not — which is more confusing than a clean outage, because the output is non-zero and therefore looks alive.

The silent zero. The one I care about most, because it survives every retry and every alert. A guard clause returns early. A query that used to match now matches nothing because a field was renamed. A candidate pool is genuinely exhausted. The code is fine. The schedule is fine. The job is doing precisely what it was told and the outcome is still wrong.

Make the job assert its own output

The instinct after being caught by this is to add monitoring. I think that is mostly the wrong move, or at least the second move. More alerting on a job that is not reporting anything meaningful gives you more silence, at greater expense.

The first move is to make the job say what it did in terms that distinguish the two zeros. Not "completed", which it always is, but which of these it was: produced N items; produced nothing because the candidate pool was empty, which is expected; or produced nothing because every attempt threw, which is not. Those are three different sentences and a job that cannot emit the right one does not know its own state, which means nothing downstream can know it either.

The second move is to alert on staleness rather than on error. Errors are the easy case; they announce themselves. Record the timestamp and the count of the last successful completion, and let something notice when that timestamp gets old. A heartbeat that goes quiet catches every one of the three failure modes above, including the silent zero, which no error-triggered alert can ever catch because there is no error.

This is a single record, written once per run, read once per check. Which matters, because of the next part.

The watchdog must not cost more than the thing it watches

I built a cost monitor once that was itself a meaningful line on the bill. It refreshed by scanning collections to count documents. In a metered database, counting by reading is charged per document, so the tool whose entire purpose was to keep spending down was spending, on a schedule, for ever.

The same trap catches dashboards, admin panels and health checks generally. Any panel that answers "how many" by fetching the things and measuring the array has a cost that grows with success. A counter incremented on write answers the same question for one read. An aggregate document holding a summary answers a dozen questions for one read. Where an exact count is truly needed, an aggregation query that returns only the number is dramatically cheaper than the documents it counted.

The rule I ended up with is blunt: a tool built to observe a system must have a cost that is flat in the size of the system. If observing gets more expensive as the thing being observed grows, the observation will eventually be switched off, usually at the moment it becomes most useful.

Deploys belong in the same accounting. On some platforms every function deploy runs a container build, and container builds are metered. Ten deploys to fix one job is ten builds. Batching changed functions into one deploy is not tidiness, it is the difference between a free operation and a charged one.

What I did instead of fixing it

When I came back to the broken generator, the obvious repair was a new API key. I did not do that, and I think the reasoning generalises.

The job's dependency on a paid external service was the failure, not the key. A scheduled process that will run unattended for months should not have a credential that can be revoked by someone else's billing decision sitting on its critical path. So the generation moved to an inference endpoint already inside the platform the site runs on, with no key to expire and no per-item charge. The failure mode did not get better monitoring. It got removed.

The second change was to the rate. The original ran a batch. The replacement publishes a small number of items a day, deliberately, which is partly about spot-checking quality and partly about search engines' policies on mass-produced pages — but it also has a reliability property I did not anticipate. A slow, steady job makes its own absence visible much faster than a large periodic one, because the expected output is a number you hold in your head.

Three questions

The test I now apply to anything that runs on a timer. It should be able to answer all three without me opening a database:

When did it last succeed? Not last run. Succeed. If the answer is not recorded somewhere durable, nothing can alert on it going stale.

What did it produce? A count, written down by the job itself, not inferred later by counting things.

Is zero expected today? The job knows whether the queue was empty. Nothing downstream does. If it does not say, that information is gone.

Five weeks of nothing is a long time, and it was not caused by any component being down. It was caused by a system that had no way to express the difference between finishing and doing nothing. That is a design gap, not an operational one, and more alerting would not have closed it.

Written by Liana Grigory, Entrepreneur and Software Engineer, from work on The Care Royal, Tegula Stone and Unified Savers. Everything above describes decisions actually made on those systems, including the ones that turned out to be wrong.

Home · Terms of Use · Privacy Notice