A job that writes to a repository every day is not a fast version of the person who owns it. It lacks their context, so it has to touch only what it brought and report everything it left alone.
Several of my sites are updated every day by a scheduled job rather than by me. It pulls the repository, writes new pages, validates them, commits and pushes. That arrangement saves real time, and it is tempting to think of the job as a very fast version of myself: it does what I would do, in the same working tree, with the same permissions.
That is the wrong mental model, and it produces a specific set of bugs. A person working in a repository carries context the job does not have: what they changed today, what someone else left half-finished, what the remote looked like an hour ago, which command deploys and which merely uploads. The job has none of that. It arrives in a working tree it did not prepare, runs against state it did not create, and has to behave correctly anyway.
The model I use now is that an automated job is a guest in the repository. It touches only what it brought, it checks the house before describing it, and it does not rearrange anything it cannot put back.
The first surprise is that the working tree is not clean. A previous run may have been killed at its time limit after writing files and before committing them. A person may have left edits in progress. Another job may share the same checkout. Whatever the cause, the next run starts with untracked and modified files that are not its own.
The instinct, and the default in a lot of example scripts, is git add -A followed by
a commit. In a shared tree that one command publishes everything: yesterday's half-written article,
someone's uncommitted experiment, a scratch file with credentials in it that a gitignore rule missed.
The commit message will describe today's work, and the commit will contain much more than that.
So the job stages exact paths and nothing else. It keeps a list of every file it created or
edited during the run, and that list is the only argument to git add. Before
committing it prints the staged diff summary and compares it against the list; if anything is
staged that the job did not write, it stops. Files it finds but did not create are left exactly
where they are and named in the report, so the person who owns them knows they are waiting.
That last part matters. Leaving someone else's work alone is correct; leaving it unmentioned is how it sits there for weeks. The report says what was carried over and why it was not included.
The second class of bug comes from believing local state about the remote. Commands that report how many commits a branch is ahead or behind do not ask the remote anything. They compare against a remote-tracking reference stored on disk, which is only as fresh as the last fetch. In a checkout that a job touches once a day, that reference can be very old.
The consequence is a job that reports the wrong thing with complete confidence: commits described as unpushed that were pushed long ago from somewhere else, a branch described as current that is in fact far behind. Any decision made on that report, such as whether it is safe to push or whose work is waiting, inherits the error.
The rule is mechanical: fetch first, every time, before any comparison is printed or acted on. It costs one network round trip, and it turns a statement about my laptop into a statement about the repository.
When the push is rejected because the remote moved on, there are two ways forward. A rebase replays the job's commits on top of what is there now, and a conflict stops the run for a person to look at. A force push makes the remote match the job's view of the world and discards whatever arrived in the meantime.
A job is never allowed to force. Its view of the world is, by construction, the least informed one in the system; it has been running for minutes, it knows nothing about why the remote changed, and it has no way to recover what it would overwrite. If a rebase conflicts, the right outcome is a run that stops with its commit intact locally and a report that says so, not a run that wins.
The fourth rule is the one that is easiest to get wrong because it differs between projects. On some of my sites, a push to the main branch triggers a deploy. On others automatic deploys are off and a separate command is required. On at least one, the deploy command is actively dangerous to run from a job, because it uploads the project using a local configuration file, and anything configured in the hosting dashboard but not declared in that file can be removed by the upload.
The job therefore has to know, per project, exactly what its push triggers, and it must not reach for a deploy command merely because one exists. Where deploys are manual, the job pushes and reports that the change is waiting for deployment; the person who deploys does it with the checks that path needs. It is tempting to call that an unfinished job. It is a job that stopped at the boundary of what it could do safely, and said so.
An automated content run can finish with exit code zero and still publish something broken. A sitemap with a malformed entry is still a file; a structured data block with a trailing comma is still text in a page. Neither causes an error in the job, and both quietly damage the thing the job exists to improve.
The checks I run before every commit are the ones a person would do by eye and a job will not do unless told to. The sitemap is parsed as XML. Every JSON-LD block in every file the job touched is parsed as JSON. Every new page is checked for a unique meta description, because copy-and-paste descriptions are the most common defect templated pages produce and the least visible one. The identity data that appears in every page, the name and description and profile links, is compared across the whole site and must be byte-identical, because a job that edits one copy can easily create a second variant.
After a push, verification is done against what was published, not against what the job remembers publishing. That also means verifying by a stable identifier. On one of my sites, pages once carried the date in their paths, and a later change dropped that convention. A check that looked for today's date would have gone on finding nothing and reporting nothing wrong. Checking for the specific paths written in this run replaced a heuristic with a fact.
The final output of each run is short and specific: what was written, with paths; what was validated and how; the commit identifier; whether it was pushed; whether anything else triggered; what was found in the working tree that did not belong to the run and was therefore left alone. A report like that is the difference between automation you can trust and automation you have to audit by hand every morning, which is to say, automation that has not saved anything.
None of these rules is sophisticated. Stage only your own paths. Fetch before you compare. Never force. Know what a push triggers. Validate the artifact, not the exit code. Report what you left alone. Every one of them is something a careful person does without thinking, which is exactly why a job has to be told, and why the list only grew after each rule was missing once.
Written by Liana Grigory, Entrepreneur and Software Engineer, from work on The Care Royal, Tegula Stone and Unified Savers. Everything above describes decisions actually made on those systems, including the ones that turned out to be wrong.