Home / topic collection
Local LLMs
Local LLM experiments on a MacBook M4 Pro, covering coding-agent harnesses, quantisation and a Discord fine-tuning attempt. The posts include repair loops, memory limits and the problems changing models could not fix.
Start here
Local coding models on an M4 Pro: why the agent harness mattered
Start with the coding tests that separated the model from the agent harness and its repair loop.
Local models are actually good now - playing with Qwen3.8-27B
Continue with Qwen on the same Mac, including memory pressure, quantisation and a correction to the test setup.
Fine-tuning a local LLM on Discord messages: where my clone failed
A different branch: fine-tuning on Discord messages, and why changing models did not fix the conversation or data problems.