Skip to content

Watchdogs

Step 10 · Honesty

A crashing agent is easy. It stops, you notice, you fix it.

The dangerous failure is silent: the scheduled job that stopped running three weeks ago, the connection that dropped, the backup that has been failing since a system change nobody remembers. Everything looks fine, and nothing is happening.

FailureWhat it looks likeWhy it is hard to catch
Silent job failureA scheduled job errors and nobody seesThe schedule still looks configured
Channel disconnectMessages stop arriving, or stop being answeredNothing to notice until someone expects a reply
Stale credentialsEverything fails at once, confusinglyThe cause is one expired secret, not the symptoms
Runaway workA job loops and burns moneyNobody is watching until the bill arrives
Quiet driftSettings change and behaviour shiftsIt is gradual, and every change looks reasonable

Silent failure is the theme. So the design principle is: make the agent’s absence detectable.

  1. Record the outcome of every scheduled job. Success, failure, when. Without history you cannot answer “is this still working?”
  2. Add one independent check that complains when something is stale. Independent matters — a watchdog living inside the thing it watches cannot report that the thing died.
  3. Give it a heartbeat with the outside world, so a dead agent and a quiet week do not look identical.
  4. Report by exception. Silent when all is well; specific when it is not.
  5. Rotate its own logs, quietly, so the checks stay cheap.

When the agent detects a problem, letting it repair itself is tempting. Sometimes right, and it must be bounded.

Safe to self-healRequires a person
Restart a stalled componentAnything that changes configuration
Reconnect a dropped serviceAnything that deletes or overwrites data
Clear a stuck state and retryAnything that touches credentials
Rotate its own logsAnything you cannot explain afterwards

The dividing line: self-healing may restore a known state. It must never take the agent somewhere new.

Every repair it makes alone should be recorded and reported. A self-healing agent with no trail is one you cannot trust, because you cannot see what it changed.

CheckCadenceThe question it answers
Did every scheduled job run?DailyIs anything silently dead?
Are the channels connected and responding?HourlyCan the household still reach it?
Did the backup complete, and is it sane?DailyCould I recover?
Are credentials still valid?WeeklyWhat will break next?
Is anything burning money faster than usual?DailyIs there a runaway?
Does its memory still believe true things?MonthlyHas drift set in?

The last one has no automation I trust. Reading it back is a job for a person. That is Step 9.

One alert, believable. An alert that cries wolf gets ignored, and then you are back to silent failure with extra steps.

Mine goes to my phone, from outside the machine, and it names a specific stale thing. It does not say “there may be a problem”. It says which job has not run.

The rule I hold to: silence must be earned. A quiet day should mean nothing needed saying, not that something broke. If you cannot tell those two apart, you have a monitoring gap.

  • Every scheduled job records an outcome
  • One independent check exists and is silent when things are well
  • You have seen it fire once, on purpose, and the message was specific
  • Self-healing is bounded, and every repair is reported