---
title: "Your CLAUDE.md is not magic"
date: 2026-10-08
description: "Why I want AI-written code checked independently, with shared standards and room for developers to choose their own workflows."
tags:
  - "ai-coding"
  - "personal-workflows"
  - "llm-agents"
---

In a recent Claude session, I was getting far more code comments than I wanted. I asked it to use fewer and added the same guidance to `CLAUDE.md`, but the code still had too many comments.

Excessive comments are a fairly small annoyance, but the experience sharpened something I have been thinking about for a while: I want the requirements I depend on to be checked independently of the agent writing the code. I am increasingly less interested in making everyone follow the same AI workflow, and more interested in agreeing what acceptable work looks like.

I have been using AI to help with development since the early Copilot days, then early ChatGPT. I started using Claude Code roughly a year and a half ago, and I have used Codex a lot in my own time. Those experiments have changed my opinions repeatedly. They will probably change again.

## Asking is not enforcement

The online conversation I remember from the earlier days was full of prompt engineering. Structure your prompt like this. Give the model this role. Phrase the request in this particular way. As models improved, some of that preparation became work the model could do itself.

When I say prompting matters less now, I mean that ritual around the wording. The information still matters enormously. I need to explain what I want, what I know, and what would make the result useful. I just spend less time trying to package it in the supposedly correct shape.

There is still a lot of enthusiasm for defining elaborate workflows in Markdown, whether that means instructions, skills or a larger plugin. Some of it is useful. But writing “always do this” does not mean it will always happen. Anthropic's own documentation describes [`CLAUDE.md` as context rather than enforced configuration](https://code.claude.com/docs/en/memory).

It also moves underneath you. A new model may not need instructions that helped the previous one. Those instructions might now get in its way. Or the old model might never have followed them consistently in the first place.

When a fresh model comes out, one of the first things I do is ask it to audit my existing instructions and Markdown files. I find that useful, but it is still the model helping me maintain guidance. It does not turn the guidance into a guarantee.

For something consequential, I would rather check the completed work.

> If a requirement matters enough to stop a change being accepted, the check should happen whether or not the coding agent remembers to run it.

## Review has less time than the output needs

This feels more pressing because the pace of building software has changed so much. We are expected to produce more, and in my experience the depth of human review has struggled to keep up.

I have seen reviews become hurried: “Hey, we've got to move. Can you take this?” Review has not disappeared, but it has become less stringent. There is only so much code someone can properly understand in the time available.

It is scary having accountability for a change you do not understand. Someone can get their own Claude session to review a change that your Claude session has already reviewed, but a human is still putting their name to it. Another model response does not make that discomfort go away.

I am not trying to settle whether this change in pace is good or bad. Telling people to review everything as thoroughly as before does not create more time.

An automatic model review may help, but I do not think merely adding a reviewer to a pull request answers the problem. What is it looking for? Does it know the requirements of this codebase? How much confidence should we put in its judgement?

I want a clear distinction between checks with fixed, mechanically evaluated conditions and reviews that need model judgement. Both can be required to run independently of the coding agent. Only the former has a verdict that does not depend on what a model thinks. Making a model review mandatory makes its execution dependable; it does not make its conclusions dependable in the same way.

A separate reviewer can still miss something, misunderstand a requirement, or behave differently after a model update. A review having run does not establish that the change is good.

## The criteria need human friction

This is where I think a team needs to be strict: the codebase-specific criteria against which work is accepted.

I would want qualified people to discuss changes to those criteria, with more than one person protecting them. Almost a council, although that sounds grander than I mean. People who know why a requirement exists and are prepared to question an addition, rather than accepting it because Claude suggested it.

“Oh, Claude said the thing. Now this is how the thing is done.” I recognise that habit. A convincing explanation can slide into becoming a rule without much consideration of whether it deserves to be one.

I want friction there. A good discussion should be necessary before something is added. Otherwise the criteria can grow endlessly, accumulating every plausible suggestion until the review is carrying a pile of rules nobody has properly defended.

Independently run checks are already increasingly part of how I work. For now, that seems to be working better for me and giving me more confidence. The group of humans protecting a shared set of criteria is how I would like that idea to work in a team.

The attraction is that it leaves room for people to work differently. One person might prefer Claude, another Codex. Someone might want a detailed workflow; someone else might work better through a long conversation. If we can agree on the requirements and check the output, we have less reason to prescribe every step they take to get there.

I expect that to travel better between models than a workflow built around steering one model through a particular sequence. It will still need attention, especially wherever a model supplies the judgement.

> Diagram: [Individual workflows, shared checks](interactives/shared-checks/). Open the native diagram on billiem.uk.

![Personal workflows produce candidate changes, followed by required checks and a decision. A separate support branch shows qualified humans maintaining model-review criteria.](media/shared-checks-static-f7f93cad69ad-27f71fb9a372-375.png)

*A static version of the proposed boundary. The checks run independently of the coding agent; model-review judgments remain fallible. Human stewardship of the criteria is a team proposal.*

## Useful knowledge, cheap text

While talking through this article, I initially said plugins were silly. That was too broad. I use skills and workflows myself, and [plugins can package tools and hooks as well as instructions](https://code.claude.com/docs/en/plugins).

There are absolute gold mines of human knowledge worth putting into these systems. Someone understands a problem, knows the awkward constraints, or has learned what matters through experience. A model can help turn that knowledge into a useful document or reusable skill. That can save me explaining the same important things repeatedly.

Generating a large document is easy now. What I question is the value of a workflow that contains no knowledge beyond what another person could generate by asking the same model. Its size and apparent completeness do not give me much reason to trust it.

And making it compulsory across a team can be damaging. People have different minds and work in different ways. A shared workflow that suits its creator may push somebody else away from the way they work best. I would rather protect the shared requirements and leave more freedom around the process.

That includes my own favourite way of talking to models: voice.

## Talk first, then read

This article is an example. I dictated my thoughts for as long as I had something to say, elaborated, and asked the model to interrogate me. Its questions helped me get more of the thinking out, including the correction about plugins and what I actually meant by prompting not mattering.

I read the responses. I do not want to listen to the model talk back when I can use my eyes to inspect what it has understood and decide what needs correcting. Once it has enough context, it can organise the work, prepare tasks for other sessions or delegate to agents.

My recommendation is blunt: you should be using voice-to-text. I use [Wispr Flow](https://wisprflow.ai/), and I have also written about [building a local Mac dictation app with Codex](/posts/codex-built-local-wispr-flow/). (By the way, my referral code [BILLIE67](https://wisprflow.ai/r?BILLIE67) gives new users a [free month of Pro](https://docs.wisprflow.ai/articles/6496688316-how-to-find-your-referral-link).) My preference for talking still does not make it the right interface for every person or task.

[Joseph Lapin describes a similar approach](https://josephlapin.com/blog/talk-to-your-claude-code/): speak the goal and its backstory, then let the model organise the material.

What I like is being able to give the model a lot of substance without first turning it into a polished prompt. I can ramble, contradict myself, answer questions and refine the direction. Organising that material is work the model can help with.

I do not think anyone has found the optimal way to use these tools. People are trying to achieve different things, and another model will come along and change what works. I want room to keep experimenting with how I work, while knowing that the checks I depend on will still happen when I finish talking.
