The experiment started badly.
On Day 1, I built an Obsidian plugin for the AI Learning Path. The plugin worked. The logging habit didn’t. I missed the first entry and had to recover it the next day through a structured interview.
That was the first useful signal. The failure wasn’t model quality. It was system design. A logging workflow that depends on remembering to log is weak, even if the automation around it works.
Thirty days into that 90-day experiment, the strongest finding is not that I found better tools.
It’s that the useful tools earned their place by removing repeated reconstruction.
That sounds smaller than most AI claims. It is. But it’s also the part that held up.
From there, the test became more practical: which tools reduce real friction, and which ones just add another surface to maintain? Some only feel useful because they produce output.
The answer kept coming back to the same pattern: AI was most useful when it carried execution, not judgment.
Superwhisper became useful because it lowered capture friction. I could speak a messy thought while moving, then Cody could route it into the right place: daily note, effort note, task list, learning log, or nowhere. The important part wasn’t transcription. It was the whole capture-to-routing path.
Cody became useful as the daily operator because it reduced context rebuilding. It could scan open loops, return three focus items with source paths, update notes, close logs from evidence, and keep the vault state moving without me manually reconstructing the last few days.
Claude stayed useful, but in a narrower role. It earned quota when it changed the system: critique, diagnosis, structural fixes, editorial judgment. It wasn’t the right default surface for routine operations. That distinction mattered. A strong tool can still be too expensive for the wrong job.
The same filter applied to everything else.
Promptdeck earned a place because it made prompts available outside Obsidian without copy-paste friction. last30days earned a place as a scouting layer before deeper research. The weekly-note skill earned a place because weekly reconstruction repeats. The morning open-loop scan earned a place because it turned “what should I test today?” into three evidence-tied options.
Other things didn’t earn a place.
A productivity-tool decision was deferred four times. The useful conclusion wasn’t “try harder to decide.” It was that no real stake existed yet. Written rules and scheduled slots both failed because the decision didn’t matter enough. Carrying it honestly as an open loop was cleaner than pretending another forcing mechanism would solve it.
Fable 5 was stranger. It looked strong on real build work: two native macOS apps shipped, first-pass compiles, minor correction burden. Then it was temporarily pulled under a U.S. export-control order before I could even confirm which routing rule the adopt verdict should change. That changed the calculation. Capability alone wasn’t enough. Availability risk became part of the cost.
I’m not sure “AI carries execution, not judgment” survives that one intact. Deciding whether a tool might vanish next month for reasons that have nothing to do with its output isn’t a content judgment. It’s closer to a policy and availability risk read, and I don’t have a clean method for it yet. For now the honest version is: don’t make anything a default surface on capability alone, and accept that this category of judgment doesn’t fit the execution/judgment split I’ve been using.
The biggest shift was in how I judged workflows.
At the start, I was still tempted to count activity: tools tested, apps built, models compared, notes updated. By the end of the first month, the log had become stricter. A test only counted if it removed context reconstruction, reduced repeated steps, improved a decision, tightened a quality gate, or produced clearer close-out evidence.
That made the quiet failures easier to see.
Several planned tests never ran. The ebook editorial micro-loop carried for days without execution. On heavy work days, optional tests were the first thing to drop. The log eventually stopped treating that as pending work and named the real failure mode: carrying something forward keeps it visible, but it doesn’t create capacity.
That’s a better finding than pretending the plan worked.
The vault also changed role. It stopped being only a place where results were stored. It became the operating surface. AI could read prior decisions, retrieve voice patterns before drafting, connect source notes to statement notes, update workflow records, and turn daily evidence into weekly review material.
But the vault didn’t make the AI right.
It gave the AI better material to work from. The judgment still stayed mine: what to keep, what to reject, when a claim was too smooth, when a tool was adding maintenance, when an extra reviewer was producing noise.
That boundary is the part I trust most after 30 days.
AI is useful when it reduces the work around the work: retrieval, routing, formatting, checking, publishing, handoff, review. It’s less useful when it creates new surfaces that need care, or when it turns planning into a substitute for execution.
The month-one rule:
If a workflow removes repeated reconstruction without lowering quality, keep testing it.
If it only creates more output to manage, cut it.