On September 11, I built and tested a negotiation-prep skill against two fictional situations. It became Day 090 when I closed the missing log entry two days later.
The Outcome field says Worked. The Reusable field says Needs refinement. I haven’t used it for a real negotiation.
That split is the clearest result of the ninety-day run. A task can finish cleanly without proving that it belongs in the way I work.
Two verdicts kept the record honest
The log records what happened that day and gives reusability its own verdict.
The distinction looks small until a build passes its checks. Day 090 produced a complete skill and a call-sheet template. Its rehearsal flow stops for the human response. The package was validated after one ambiguous rule was fixed. That supports Worked as the task outcome.
Its fictional exercises can’t show whether the preparation improves a real conversation. Needs refinement is the only reusable verdict the evidence supports.
Across the full log, 55 entries are Worked, 29 are Partial, and three are Failed. Three earlier entries use more specific outcome language. Those totals don’t form an improvement curve. They show what happened under a scoring system that became more explicit about the difference between delivery and adoption.
At Day 030, I wrote that useful AI was reducing repeated reconstruction. The next sixty entries added a harder test: the workflow had to work again.
Repeat use cut the list down
Nothing enters the workflow registry until it has been used at least three times.
Seven workflows crossed that line. Capture routing and interview-based log closing each reached four confirmed uses. Voice retrieval before drafting and reading-to-statement checking reached three. Audit-then-fix reached four, and the local sleep-review cycle reached five.
Evidence-first record reconciliation reached eight.
That workflow repairs missing or conflicting daily notes and learning entries from session records and file history, with direct confirmation when the files can’t settle it. It restores supported facts and preserves real skips. Private or unavailable details remain uncaptured.
The final use closed Day 090 from the September 11 session and installed package. It also kept September 12 and 13 outside the numbered run. Work happened on those dates, but the ninety-day sequence had finished.
This is maintenance work. It also solved a recurring operational problem: completed work could disappear from the record when one closeout failed or one AI surface couldn’t see another. Eight recoveries are stronger evidence than a clean first demonstration.
The ROI record blocked the bigger claim
The project was supposed to track operational impact. It often captured less than the work produced.
At the Day 060 review, the ROI tracker held four metrics. The latest had been added twelve entries earlier. During that gap I had built a vault backup path and reduced the paid AI stack. The ebook had advanced too. Most of the work had no defensible before-and-after measurement.
The tracker now holds seven evidence-backed patterns. Only Superwhisper-to-Cody capture routing is rated High confidence. The others remain Medium because the evidence is qualitative or the time saving wasn’t measured repeatedly.
The record can’t support a headline about total hours saved over ninety days. That restraint matters. Activity volume didn’t become ROI because I logged it every day.
The final week exposed another boundary. Enterprise Copilot and work-provided Claude supported real work, but the task details stayed inside employer systems. The personal log recorded those days as Partial because it couldn’t assess the result, correction burden, or operational impact from the evidence available outside work.
The Partial rating applies to the personal record’s evidence. It can’t evaluate what it can’t see inside employer systems.
The numbered run is finished
All ninety entries are closed. On September 13, after Day 090 closed, I archived the full 90-day learning path and its dedicated logging setup, including the plugin, dashboard, and daily logging skills.
The workflows that earned enough evidence to survive the experiment now continue as ordinary standing habits. They don’t need the numbered daily log to keep running. The three-use threshold did its job by keeping first demonstrations out of the registry. Evidence-first reconciliation finished with eight accepted uses and a High reliability rating. The ROI tracker prevented claims the available evidence couldn’t support.
The logging system itself is retired to the project archive. I don’t need to keep maintaining a separate apparatus just because it was useful during the experiment.
Negotiation-prep closes the numbered record as a completed build. Its next test is one real conversation. Until that happens, it stays outside the workflow registry.