Watchdogs
Step 10 · Honesty
A crashing agent is easy. It stops, you notice, you fix it.
The dangerous failure is silent: the scheduled job that stopped running three weeks ago, the connection that dropped, the backup that has been failing since a system change nobody remembers. Everything looks fine, and nothing is happening.
The failures worth designing for
Section titled “The failures worth designing for”| Failure | What it looks like | Why it is hard to catch |
|---|---|---|
| Silent job failure | A scheduled job errors and nobody sees | The schedule still looks configured |
| Channel disconnect | Messages stop arriving, or stop being answered | Nothing to notice until someone expects a reply |
| Stale credentials | Everything fails at once, confusingly | The cause is one expired secret, not the symptoms |
| Runaway work | A job loops and burns money | Nobody is watching until the bill arrives |
| Quiet drift | Settings change and behaviour shifts | It is gradual, and every change looks reasonable |
Silent failure is the theme. So the design principle is: make the agent’s absence detectable.
Do this
Section titled “Do this”- Record the outcome of every scheduled job. Success, failure, when. Without history you cannot answer “is this still working?”
- Add one independent check that complains when something is stale. Independent matters — a watchdog living inside the thing it watches cannot report that the thing died.
- Give it a heartbeat with the outside world, so a dead agent and a quiet week do not look identical.
- Report by exception. Silent when all is well; specific when it is not.
- Rotate its own logs, quietly, so the checks stay cheap.
Self-healing, carefully
Section titled “Self-healing, carefully”When the agent detects a problem, letting it repair itself is tempting. Sometimes right, and it must be bounded.
| Safe to self-heal | Requires a person |
|---|---|
| Restart a stalled component | Anything that changes configuration |
| Reconnect a dropped service | Anything that deletes or overwrites data |
| Clear a stuck state and retry | Anything that touches credentials |
| Rotate its own logs | Anything you cannot explain afterwards |
The dividing line: self-healing may restore a known state. It must never take the agent somewhere new.
Every repair it makes alone should be recorded and reported. A self-healing agent with no trail is one you cannot trust, because you cannot see what it changed.
What to check, and how often
Section titled “What to check, and how often”| Check | Cadence | The question it answers |
|---|---|---|
| Did every scheduled job run? | Daily | Is anything silently dead? |
| Are the channels connected and responding? | Hourly | Can the household still reach it? |
| Did the backup complete, and is it sane? | Daily | Could I recover? |
| Are credentials still valid? | Weekly | What will break next? |
| Is anything burning money faster than usual? | Daily | Is there a runaway? |
| Does its memory still believe true things? | Monthly | Has drift set in? |
The last one has no automation I trust. Reading it back is a job for a person. That is Step 9.
Designing the alert
Section titled “Designing the alert”One alert, believable. An alert that cries wolf gets ignored, and then you are back to silent failure with extra steps.
Mine goes to my phone, from outside the machine, and it names a specific stale thing. It does not say “there may be a problem”. It says which job has not run.
The rule I hold to: silence must be earned. A quiet day should mean nothing needed saying, not that something broke. If you cannot tell those two apart, you have a monitoring gap.
Checkpoint
Section titled “Checkpoint”- Every scheduled job records an outcome
- One independent check exists and is silent when things are well
- You have seen it fire once, on purpose, and the message was specific
- Self-healing is bounded, and every repair is reported