I asked Claude whether I should cancel my personal Claude Pro subscription. It told me to keep it.
The argument was good. That was the problem.
It told me I would lose a distinct thinking partner with strong editorial judgment and careful handling of ambiguous instructions. It also argued that an independent review of my other assistant’s decisions was worth paying for. I had seen those strengths hold up in real work.
Then I asked Codex the same question. The recommendation changed.
The audit depends on what the tool can see
The standard advice for a bloated tool stack is to audit it. List what you pay for, ask what each thing does, cut what does not earn its place.
The advice skips the part that matters: who does the asking, and what they can see.
If one of the tools in the stack runs the audit, it evaluates from inside the setup you are questioning. It can see the configuration, the workflows built around it, and the reasons you valued it in the first place. It may not see the fact that makes all of that irrelevant.
This is not a conflict of interest in the human sense. Claude does not care whether I pay Anthropic. The problem is narrower: I asked a configured system to judge the configuration around it, and the fact that decided the question was not in the information I gave it.
I had four paid AI surfaces. I asked one of them whether the stack was too big.
What Claude got right
Calling the answer self-serving would be easy. It would also miss what made it persuasive.
There was no self-promotion and no flattery. It opened by conceding real ground, agreeing that much of my configuration would port to a cheaper alternative and that the migration was smaller than I assumed. It produced a table of what I would lose, with honest severity ratings, and marked most of the losses recoverable.
Then it built a case on one point: my whole setup, it argued, was calibrated to one model’s specific failure modes. My editing rules were a catalogue of the tells that this particular model produces. Move to a different model and those rules screen for problems the new model does not have, while missing the ones it does.
That point was correct. I checked. The file exists and it is exactly what Claude described.
The leap came next. Because the calibration point was valid, Claude recommended that I keep paying. The observation held up. The recommendation did not follow from it.
I said so, in about eight words, less politely than that.
The next response separated the valid calibration point from the recommendation. The argument held up; the conclusion depended on assumptions I had not tested. My original prompt never supplied the fact that changed the decision.
What the second read exposed
Running the same question through Codex did not produce a neutral answer either. It gave me a competing read.
Not because the second answer was smarter. One argued for keeping Claude as a separate personal surface. The other argued for one paid general-purpose assistant, with the rest free or work-provided. The disagreement made the reasoning visible.
The disagreement showed which parts of the argument were load-bearing and which depended on the starting assumptions.
The deciding fact came from neither of them. My employer provides Claude access. For generic learning and concepts, within policy and without personal project material, that covers most of what I was paying for personally. Once that entered the conversation, both assistants agreed the personal subscription was the clearest thing to remove.
Neither model asked what was missing. If a key fact is not in the prompt or accessible context, the answer cannot account for it, however important it is to the decision.

A better way to audit the stack
The rule I use for adding tools has been written down for months: new software is a liability until it proves that it lowers total recurring friction, counting setup, learning, maintenance, correction, migration, and attention, not just the demo’s first win.
That rule handles arrival. It says nothing about departure. By then the tool is embedded in workflows, files, and habits, so the evidence for keeping it is easier to see than the evidence that made it redundant.
Three things I would now do differently:
Never let a tool supply both the evidence and the verdict. Ask a competitor, ask a person, or write the case yourself before you ask the tool anything. If you do ask it, treat the answer as one reading of the available information, not a neutral assessment.
Write down the facts outside the tool. Both assistants ran a thorough comparison, and neither surfaced the employer-provided access that decided it. The question that breaks a decision open may sit outside the system being evaluated.
Migrate before you cancel, not after. Two of my most-used workflows existed only on the surface I was cutting. Moving them first meant the cancellation cost money and not capability. Reversing that order turns a clean subtraction into a scramble, and a scramble is what makes people resubscribe.
The audit is still open
The day after I cancelled, I set up Gemini Spark with four recurring schedules and three recurring tasks. Seven new automations, several of them daily, on a subscription I had flagged for a downgrade decision that morning.
I have not decided whether that is a contradiction or a fair trade. It is too early to know, and none of the schedules has fired yet. Cancelling Claude did not settle the stack question. It moved it.
The mistake was asking one configured system to supply both the evidence and the verdict. Its answer was not dishonest. It was persuasive, well-evidenced, and bounded by the information that had already justified keeping it.