Your CLAUDE.md is not magic
Why I want AI-written code checked independently, with shared standards and room for developers to choose their own workflows.
Home / topic collection
Writing about LLM agents on billiem.
Latest first.
Why I want AI-written code checked independently, with shared standards and room for developers to choose their own workflows.
Three autonomous GPT-6.1 Sol experiments: a fictional record, a sheet-cutting planner and a congestion model, with their review fixes and limits.
What three Astra agents built with room to choose: a string puzzle, a supply-chain experiment and a printable booklet, with their fixes and limits.
How deterministic tools, scheduled Codex jobs and inspectable HTML fit into two daily reporting workflows.
An LLM helped build a deterministic 322 draft solver. Paired tests covered 120,000 outcomes, with a 27.45% title rate at the largest field budget.
What GPT-5.6 did across research, audits, a seven-repository migration and one animated pet during my first 48 hours with Codex.
What changed between two GPT-5.6 coding-experiment batches when I gave depth more room, changed the prompt and added review.
Five 4-bit local models on a 48 GB M4 Pro, tested with file-output contracts, smoke-test repair loops and Pi, then compared with a hosted Codex setup.
A Codex-built biological wheel simulation went from a Discord debate to Cloudflare Pages in about ten minutes. A toy model for discussion, not proof.
My workflow uses Linear as an idea ledger, Wispr Flow for dictation and agents to turn rough notes into projects and posts.
Using autonomous coding-agent runs as a low-cost way to find ideas, learn unfamiliar topics, and seed later projects.