How It Is All Auditable
Step 11 · A Team
Four agents doing work on their own is only acceptable if you can find out what they did. This page is what gets recorded, where, and how I actually check it.
The design principle: two records, for two readers. Machine-readable ledgers that are complete, and one plain-text log a person can read in a minute.
What is recorded
Section titled “What is recorded”| Record | What it captures | What it answers |
|---|---|---|
| Run history | Every scheduled job: when it started, when it finished, whether it succeeded, and its output | “Did last night’s jobs actually run?” |
| Incident log | Every failure, tracked as an incident with first seen, last seen, acknowledged and closed times | “How long has this been broken, and did anyone notice?” |
| Usage ledger | Every scheduled model call: job, model, tokens, duration, and whether the response was silent | “What did that cost, and where did it go?” |
| Per-agent session store | Each agent keeps its own conversations, and which model and provider served each one | “Which agent did this, and with what?” |
| Card history | Every handover: who wrote it, who owned it, when it moved, every comment, and every state change | “Why was this done, and by whom?” |
| The action log | A hand-written append-only log, one dated entry per piece of work | “What has happened lately?” |
Two of these deserve more words.
The incident log has a state machine. A failure is not just written down; it is tracked from first seen to acknowledged to closed. That is the difference between a log and an audit. A log tells you something broke. An incident record tells you nobody looked at it for two days.
Each agent keeps its own history. Because the agents are separate, their records are separate. I can ask what the health lane did last week without the household lane’s traffic in the way.
The one a person reads
Section titled “The one a person reads”All of the above is for machines and for troubleshooting. The record I actually read is the action log: a single plain-text file, append-only, one entry per piece of real work.
Each entry says what was asked, what was done, what was delivered, and what was left open. Here is the shape of one, from a form I needed completed:
Trigger: me, in the morning — asked for a registration form to be filled. The premise, corrected: the form was not where I thought. Filled from records: the details that already existed, and where they came from. Left alone: two fields that could not be verified — left blank rather than guessed. Deliverables: the completed document, and a draft ready for me to send. Still open: the follow-on form, offered as the next thing.
That entry is worth more than any dashboard. It says what happened, what it used, and — crucially — what it refused to guess.
Mine is at a little over a hundred entries. I read the recent ones weekly, and I can search the whole thing.
Why this matters more with several agents
Section titled “Why this matters more with several agents”With one agent, you can remember roughly what it has been doing. With four, across a household, a family, health and a business, you cannot.
The records are what let me say yes to four agents running unattended. Not trust — evidence.
How I actually check it
Section titled “How I actually check it”Three questions, and all three have real answers:
- “Did everything run overnight?” — the run history, and the watchdog that complains when something is stale.
- “Has anything been failing, and for how long?” — the incident log, with acknowledgement and closure times.
- “What has it been doing?” — the action log, read like a diary.
If any of those three has no answer, the setup is not auditable, however much is being logged.
What is not recorded
Section titled “What is not recorded”Worth being honest, because an audit section that claims completeness is lying.
- Its private reasoning. I can see inputs, outputs, tool calls and results. I cannot see why it chose one phrasing over another, and I do not want to read a transcript of it thinking.
- Everything from the family’s private lane. Deliberately. The family lane’s conversations are its own.
- Anything that was never written down. If a piece of work never produced a card, a log entry or a delivered output, it is invisible. That is the honest gap, and the reason the action log is a discipline rather than a feature.
An audit trail is not something you can turn on. It is what happens when every piece of work leaves a mark, including the work that failed.
Checkpoint
Section titled “Checkpoint”- Every scheduled run is recorded, and you can see the last success
- Failures are tracked as incidents with acknowledgement and closure
- Usage is metered per job, with the model recorded
- Every handover leaves a readable card with its full history
- There is one plain-text log a person can read
- You have read it recently and it was true