What GPT-6.1 Sol chose to build
Three autonomous GPT-6.1 Sol experiments: a fictional record, a sheet-cutting planner and a congestion model, with their review fixes and limits.
Home / topic collection
Writing about AI experiments on billiem.
Latest first.
Three autonomous GPT-6.1 Sol experiments: a fictional record, a sheet-cutting planner and a congestion model, with their review fixes and limits.
What three Astra agents built with room to choose: a string puzzle, a supply-chain experiment and a printable booklet, with their fixes and limits.
A 4-bit Qwen3.8-27B built a cellular automata workbench locally on my M4 Pro, while 6-bit finished and 8-bit hit OOM.
An LLM helped build a deterministic 322 draft solver. Paired tests covered 120,000 outcomes, with a 27.45% title rate at the largest field budget.
An interactive embedding visualisation compares raw, PCA and grand-tour projections, including a basis change that preserves cosine neighbours.
What GPT-5.6 did across research, audits, a seven-repository migration and one animated pet during my first 48 hours with Codex.
What changed between two GPT-5.6 coding-experiment batches when I gave depth more room, changed the prompt and added review.
Five 4-bit local models on a 48 GB M4 Pro, tested with file-output contracts, smoke-test repair loops and Pi, then compared with a hosted Codex setup.
Using autonomous coding-agent runs as a low-cost way to find ideas, learn unfamiliar topics, and seed later projects.
A local MLX LoRA experiment used my Discord messages to imitate my replies. Missing conversation turns limited what fine-tuning and retrieval could fix.