On July 27 I told my AI operator to open today’s notes. It launched Obsidian.
What I meant was create the Markdown files for the day and start filling them in. The clarification took ten seconds. It’s recorded in that day’s learning entry under “Human corrections required,” and it’s why I started questioning where responsible AI guidance has to show up.
My written rule says AI should carry execution, not judgment. The human owns noticing, the standard, and acceptance of truth claims. I wrote it on June 4 and revised it once.
It’s a good rule. It has never made me check whether I followed it, because a rule in a note can’t read itself back.
What forced the check was a required field in a daily log. It predates the rule by thirteen days.
What the field is
The log is part of a ninety-day exercise I started on May 22 to test whether AI tools were producing operational leverage or only activity. Each day records the task, outcome, friction, correction burden, failure modes, and reusability verdict.
The corrections heading is one blank line I have to fill before the day closes. It doesn’t accept “some.” It has to say what the correction was.
Through the August 3 entry, the log held 63 answered correction fields. I read all 63 expecting an error log.
That’s not what I found.
What I was actually correcting
The pattern I saw most often was me supplying direction the system had no way to derive.
On July 3, while building a Stoic quote app, I decided on quote-only cards and one prompted journal entry per day. The AI couldn’t infer that I wanted to reject reflection paragraphs or morning-and-evening pairs. On July 2, I decided that a calculator and a pomodoro timer were lane tests rather than committed products. On July 4, I chose which vault evidence belonged in the ebook and what had to stay out.
Those weren’t errors. They were the parts of the work that were never delegable.
Another pattern was facts only I had. On June 1 the entry notes that a blue envelope had already been sent on Friday May 29. On June 4 it corrects an assumption about my step count: a 14:00 to 23:00 shift still includes movement between meeting rooms, lunch, and breaks.
What I found less often was what I expected the whole field to contain: faults caught before delivery.
On May 24 a model claimed that Claude Code and Codex were CLI-only tools. That’s false, and the entry says it was caught before any writes. On July 24 automated visual QA flagged four problems in the generated ebook files: a blank verso page, duplicate title metadata, a two-line orphan, and a raw Obsidian todo marker. Each required a correction to the generator before delivery.
Five caught faults across two months of near-daily use isn’t nothing. It wasn’t the pattern I encountered most often.
What the field does that the rule can’t
The rule asks a question I can answer without looking: am I keeping judgment human? Yes. I believe that. I wrote it down.
The field asks what I had to correct today. I can’t answer that honestly from memory of my intentions. I have to reconstruct what happened.
That’s the mechanism. The field isn’t an active boundary during the work. It’s a forced reconstruction at the end of the day, at the cost of about one line.
Its by-product is a record of where my judgment was going. If you’d asked me in May, I would have said error-catching. The examples I kept recording were scope-setting and missing context. Those imply different failure modes. The main risk isn’t only that a wrong fact gets through. It’s that the work drifts toward whatever shape the tool finds easiest.
What the record can’t tell me
Seventeen of the 63 entries open with a negation: none, nothing material, no correction needed, or not applicable. Several then record a correction anyway, which says something about how quickly the first word gets chosen.
I can’t tell which days genuinely had nothing to correct and which had something I didn’t look hard enough to find. The field is completed by the same person who ran the session, usually when he’d like to be finished.
I also don’t have a known case where an entry marked “none” was later proved wrong. That doesn’t clear the method. It sets the limit of what the record supports.
A daily field is a prompt, not an audit. It makes the question unavoidable. It does nothing to make the answer honest, and I don’t have a second reader on it. The record also can’t tell me whether writing the correction changed what I did in the next session.
The version I’m keeping
I wouldn’t replace the policy. It still names the boundary. In my setup, it needs a mechanism that makes me account for what happened. Mine is a required field in a log I already complete.
A commit template or weekly review could test the same idea. The requirement is that the answer can’t be a restatement of my values.
Two months in, the field gave me one thing I can point to: a record of where my judgment was going. Whether the judgment was any good, or whether recording it changed the next session, are separate questions. The record can’t answer either one.