Draft and Focus13.7563° N, 100.5018° EBangkok

Photography, Productivity and AI

A Streak Is Not a Record

On June 1, I picked a habit tracker and wrote myself a warning in the same note.

The tracker was Streaks. I already owned it, so the test cost me nothing. It also fit Health and Apple Watch better than the alternatives I was considering. The warning was this: I’d bought Streaks once before and it hadn’t stuck. If the reason was friction, an app would fix it. If I’d chosen habits that didn’t matter, or if the reflection loop was missing, no app would fix anything.

Today I can tell you the habit survived. I only know because this article made me go back through the Streaks history, seven weeks after the note went quiet. Nothing I built had asked for the verdict when it was due.

Earlier this year, I wrote that streaks don’t work: miss one day, the counter resets, and by day three you’ve quit the app. My own record disproves the collapse. This one broke twice and came back the next day both times. So a streak can protect continuity. It still can’t explain what happened after the break.

What I designed

The rules were good. The test covered three habits over two weeks. This article follows the exercise one. It also had a clear boundary: the app is the input surface, not the ledger. Definitions and reflection stay in Obsidian.

The exercise block started on June 2 with 30 push-ups and 30 squats. A plank moved into the first week because the core was the obvious gap. Glute bridges were due in week three. One new movement per week kept the friction of starting low.

Before I could argue with myself about it, I wrote down what “automatic” would mean: 12 of 14 days completed, including at least three completions on low-energy or inconvenient days.

That threshold isn’t a streak. It has two misses built into it, and it asks specifically for the ugly days, because a habit that only runs when the day is clean isn’t a habit yet.

What the record shows

Before this retrospective check, the last contemporaneous update to the exercise note was dated June 8. It defines the 12-of-14 threshold and the review prompts, but it contains no verdict.

The training log behind that effort stops on the same date. Its final entry sets a next check: if the block is still holding at the 14-day mark, make the exercises harder before adding volume. Two files in two systems, scheduling the same check. Neither of them ran it.

There’s no outcome recorded for the 14-day check or for week three. There’s no mention of push-ups, squats, planks, or glute bridges anywhere in my July notes. The steps target I flagged as too low, 4,429, was never revisited.

The app tells a different story. The routine was still active on July 28. Its only breaks were July 5 and July 22, and both lasted one day.

So the habit survived, and the first fourteen days ran unbroken. That looks like a clean pass on the completion half of the threshold. It isn’t quite one. Some days I tapped the habit done myself. Other days a Watch workout closed it. The history doesn’t say which was which, so a completed day proves something happened, not that the push-ups and squats did.

The other half is worse. Streaks records that a day was done, never whether it was a hard one, and I never wrote that down anywhere else. The three low-energy completions I asked for can’t be counted now or reconstructed later.

Waiting to check cost more than the verdict. It cost the evidence.

The same weeks, recorded differently

At the same time, I was running a 90-day AI learning path built on a different rule: skipped days don’t advance the count, and nothing planned gets logged as done.

The sharpest version of it sits in the weekly review for the week ending June 14: “Do not backfill an intentionally skipped AI Learning Log day to keep a streak. 2026-06-14 stays visible as skipped.”

I started that review on June 15 and finished it the next day. June 15 was day fourteen of the exercise block. So while I was writing down a rule about keeping a skipped day visible, the check that would have made the exercise result visible came due, and nothing fired.

When I drafted this on July 28, that log held 56 entries recorded on 55 of 67 calendar days. Twelve days were visibly blank. Of the 56 entries, 35 were marked Worked, 16 Partial, 2 Failed, and 3 used their own wording. The path was running at Day 056.

Both systems allowed me to miss. Only one of them asked what the miss meant.

The difference shows up when I go looking, though not as cleanly as I wanted. The longest gap is five days, July 17 to 21. Two of them are accounted for: I was sick and then resting, and the note that restarts the week says so. The other three have nothing behind them. The weekly review records the hole once per missing day and then writes down a rule instead of a reason: do not infer work for July 17 through July 19.

That’s a thinner answer than I expected to find. It’s still a different kind of record. Nearly a third of the entries are Partial or Failed, and the log shows which experiments produced nothing. Where it can’t explain a gap, it says so and blocks the guess. Streaks tells me I missed July 5. It doesn’t tell me why, it doesn’t mark the question as open, and the vault never asked.

Continuity isn’t enough

James Clear’s identity frame still holds: every action is a vote for the type of person you want to become. Gabe Bult’s Two Day Rule turns continuity into a simple instruction: never skip twice.

Those ideas organise around continuity, and continuity is worth protecting. But they have nothing to say once the chain is already broken. They treat the miss as the thing to avoid, not as the thing to read.

Anne-Laure Le Cunff’s Tiny Experiments gave me the missing frame. A goal ties identity to one outcome, so any deviation registers as failure rather than feedback. Run the same behaviour as a time-boxed experiment and there’s no wrong answer, only an observation. That makes an honest log survivable. Sixteen partials and two failures out of 56 would be a demoralising scoreboard. As experiment results, they’re the useful half of the data.

Ryder Carroll supplies the piece I actually missed. In the Bullet Journal method, incomplete tasks don’t carry forward on their own. At the end of each month, you move, schedule, or drop each one. If a task isn’t worth rewriting, it isn’t worth doing. The rewrite isn’t admin. It forces a verdict.

I had no migration step for the exercise habit. So it wasn’t killed, and it wasn’t renewed. It just stopped being mentioned.

The missing step

The threshold wasn’t the problem. Twelve of fourteen with three hard days is still the rule I would write today.

The problem is that a threshold isn’t a system until something runs it. I wrote a check but never gave it a fixed place. The tracker held the completion data, which was half of what the threshold asked for and not a clean half. The vault held the definition of success. Nothing brought them together at the 14-day mark, and nothing wrote down the part the tracker couldn’t see. A live habit and a dead one left the same gap in the ledger.

The AI log at least makes absence legible. It has a rule for what a skipped day means and a place where that meaning gets written, even when the meaning is that I don’t know. That’s the whole difference, and it costs about a line of text.

I’m not switching apps. The fix is to give the 12-of-14 check a fixed slot in the weekly review and write the verdict in the vault. That means recording what the app can’t see, and recording the answer even when the answer is that the habit stopped. A record that only holds the good weeks isn’t a record. It’s a highlight reel with a counter attached.