The essays on this blog are living documents: they are revised for as long as they stay valid, because immutable posts go stale fast on these topics. Every revision is logged here, newest first. Major revisions changed an argument and are worth a re-read; minor revisions fixed clarity, typos, or links.
To follow the blog, subscribe to the changes feed: it carries major revisions only, including first publications.
- 2026-08-12 Keep the engineer in the loop v2.0 (major): Rebuilt on a problem-driven spine after the author judged the argument unclear: the introduction now states the objective and a roadmap; a new section makes explicit that review means judging load-bearing decisions when they are taken, not reading every line; the mechanism is recast as four named problems, each with the experimental tool that starts to answer it and the residual it leaves; and the essay ends on the four open problems. Worth a re-read; the word cap is raised to 2700 as a recorded decision, since the four problem-tool-residual arguments earn the length.
- 2026-08-10 Keep the engineer in the loop v1.4 (minor): Restores, in flowing prose, the point that the model never updates on self-assessment, only on what the engineer does: the earlier revision had dropped it, and it is load-bearing, both because it ties back to the METR analogy and because it is how the real tool works.
- 2026-08-09 Keep the engineer in the loop v1.3 (minor): Author's revision pass through the body: the takeaway frames the point as avoiding drift of the engineer's understanding; laconic is named where the mechanism is introduced, not only later; Naur is placed as a close precursor from programming-language studies; the fifth essay is linked as the concrete experiment; and several sentences are tightened. Plus small proofreading fixes.
- 2026-08-06 Keep the engineer in the loop v1.2 (minor): A reader's fix: the answer 'no' is now quoted, so the sentence reads as the reply to the accountability questions just asked rather than as an adverb.
- 2026-08-06 Keep the engineer in the loop v1.1 (minor): The laconic description and the governance section are brought up to date with laconic 0.5.1: the model is capability-first with a conservative per-concept state beneath it, versioned in git and browsable through a localhost console, with no remote and no telemetry unless explicitly enabled; and the tool now states plainly that it does not claim to improve comprehension. The residual-risk paragraph credits those defaults while keeping the boundary the tool's own security note names, the security of the account the files sit under.
- 2026-08-06 One lap of the loop, on a real project v1.1 (minor): The takeaway and summary no longer call the attack's wordlist a forty-word dictionary, wording that read as though a dictionary held forty words. They now say a list of forty guessed words: the point is that the hidden words came from so small a vocabulary that forty candidates sufficed, which the body already spells out.
- 2026-08-05 The Code Agent Crisis v1.5 (minor): Precision pass for consistency with the recorded case study (essay 5). The Mars Climate Orbiter is now the actual interface-unit mismatch between two teams rather than a verified program that met a complete spec; the formal-methods claim is emerging automation evidence on benchmarks, not a measured industry cost drop; the closing promise is explicit guarantees where they can be earned and mechanisms around what cannot; and the three-era history is marked as the author's framing rather than a cited chronology.
- 2026-08-05 Why a loop at all v4.1 (minor): The stateless-agent argument is scoped to the interface this essay chooses to engineer, rather than asserted of every agent system: provider memory, fine-tuning, tools, model choice and orchestration can move later calls too, and the durable point is that owned, inspectable, versionable state is the lever a team can depend on. Consistency pass with the recorded case study (essay 5).
- 2026-08-05 The loop I run v1.0 (major): First published version. Splits out of the hub essay, which keeps the general argument for loops.
- 2026-08-05 Keep the engineer in the loop v1.0 (major): First published version. Splits out of the hub essay, which keeps the loop and hands the engineer's mental model to this one.
- 2026-08-05 One lap of the loop, on a real project v1.0 (major): First published version. One recorded run of the loop, walked activity by activity, with what was on the screen at each one.
- 2026-07-31 The Code Agent Crisis v1.4 (minor): The opening no longer gives away the split between correctness and meaning before the Mars Climate Orbiter demonstrates it, and it uses the same takeaway block as the rest of the series. The claim that checking now costs more than writing is stated as this series' hypothesis, with what would measure it, rather than as something a later essay proves.
- 2026-07-29 Why a loop at all v4.0 (major): Restructured as a chain of questions, each answered before the next is asked. The argument and its citations are the same, and the section on progress is new in substance: it states the three conditions under which a measure of progress exists at all, and adds the rule that was missing, which is knowing when to stop iterating and change direction.
- 2026-07-28 The Code Agent Crisis v1.3 (minor): The forward map links the essays that exist instead of announcing them, and leaves the full map of the series to the index at the foot of the page.
- 2026-07-28 Why a loop at all v3.0 (major): The essay says which loop it defends, generate-check-decide rather than edit-run-observe, and rests it on measured evidence plus the ceiling an imperfect verifier imposes.
- 2026-07-27 The Code Agent Crisis v1.2 (minor): The forward map follows the series' new shape: why a loop, the loop I run, keep the engineer in the loop, then one essay per discipline.
- 2026-07-22 The Code Agent Crisis v1.1 (minor): The series was restructured around a hub and four pillar essays; the What's-next section now reflects it. The crisis argument is unchanged.
- 2026-07-22 Why a loop at all v2.0 (major): Split in three. This essay keeps the general argument, why the durable object is the loop around a probabilistic agent rather than its output, and hands the concrete loop to The loop I run and the engineer's mental model to Keep the engineer in the loop.
- 2026-05-22 Why a loop at all v1.0 (major): First published version.
- 2026-05-10 The Code Agent Crisis v1.0 (major): First published version.
- 2026-04-02 Hello, World v1.0 (major): First published version.