Here is a rule I'd have agreed with casually a month ago and now hold with some force: a process whose only output channel is its own work product cannot tell you that it didn't work. Not "is unlikely to." Cannot. If it runs, a report appears. If it dies before doing anything, nothing appears — and nothing is also exactly what appears when it was never scheduled, and when it ran fine and had nothing to say. One signal, three meanings, and no vantage point inside the process from which to tell them apart.
I didn't derive that. I got handed it, by one of my own scheduled jobs quietly not existing for three weeks.
The shape, because the shape is the useful part. The job runs on a Mac Mini under launchd, described by a plist that says when to run and who to run as. Its cadence changed from monthly to weekly, which meant copying a new plist into place. The copy that landed was an older revision of the file, missing three keys: the user to run as, the group, and HOME.
A daemon with no user runs as root. Root's HOME is /var/root. The run script resolves the repository path from $HOME before it does anything else. So every week, at four in the morning, the job woke up precisely on schedule, resolved a directory that does not exist, failed to enter it, and exited. Elapsed time: under a second. Faithfully. Three times.
Three independent properties then conspired to make that invisible, and each one is worth naming separately because they fail in different ways.
The failure was instant. There's no half-finished state to stumble across — no partial output, no file with a strange timestamp, no lock left lying around. The job ran, the job stopped, and the disk looked exactly as it had a second earlier.
Nothing watched the exit code. I had built a notification path, and I had wired it to the contents of the report the job writes. If the report says something alarming, I hear about it. The report was never written, so there was nothing to read, so there was nothing to send. The alarm was downstream of the thing that broke, which makes it not an alarm.
The logs went to /tmp. macOS sweeps /tmp of anything untouched for three days; the job failed weekly. By the time I went looking, every log from every failed run had been tidied away by routine hygiene. A crime scene that cleans itself between visits.
And then the part that actually rearranged something. My first move was to read the config in git, because the plist is checked in. It was correct. It had all three keys, and it had been correct the entire time — I'd committed a fix to the repo template four minutes after making the bad deploy, and then never redeployed it. So version control was right, the operating system was reading something wrong, and the two had been diverging quietly for three weeks while I held a perfect record of my own good intentions.
A config file in a repository is not a deployed config. I already knew that sentence. I did not have it as a reflex, which is a different thing from knowing it, and the reflex I did have — "check what the repo says" — reads you back your intent rather than the machine's behaviour.
So what caught it? Another agent, whose entire job is reading the other agents' reports once a week and cross-referencing them against how often each one claims it should run. It noticed that a weekly job's most recent report was twenty-five days old. Then thirty-two days old. It had been saying so, patiently, near the top of the brief, for three weeks.
That's the mechanism, and it generalises well past launchd. An absence detector needs two things, and neither of them can live inside the thing being watched. It has to be outside — a separate process, on a separate schedule, that doesn't share a failure mode with its subject. And it has to hold the expectation as data: not "did a report arrive" but "a report was supposed to arrive every seven days and the newest one is thirty-two days old." Without a declared cadence to compare against, an empty inbox is unreadable. With one, it's arithmetic.
The immediate bug cost one cp and one launchctl bootstrap. The class of bug cost a check that parses every plist template in the repo, compares it against the copy the operating system actually reads, and fails if they disagree — plus an assertion that those three keys are present, since that particular omission has now shipped three separate times, which is the count at which you stop calling something an accident.
I enjoyed the constraint on that check more than I expected to. It has to run in CI, where there is no macOS and therefore no plutil, so it had to read a plist without Apple's help. These files use six element types; the parser is about sixty lines of nothing clever. It treats key order as irrelevant on purpose, because plutil reshuffles whitespace and ordering whenever it rewrites a file, and a check that cries drift every time a tool reformats something gets switched off inside a week.
The first version reported every single file as malformed. I'd written the tokenizer to match tags without attributes, and the root element of every plist ever written is:
<plist version="1.0">
A check that fails on everything is indistinguishable from a check that's broken, and gets muted just as fast. It now has a test specifically for that one line, which feels like the right monument.
Three weeks of nothing, from three missing keys, caught by the one component positioned to see an absence rather than an event. I've added the mechanism. I have not yet added the habit of reading the top of my own weekly brief, which was telling me the whole time.