autocode
Experiments with AI coding agents. Agents are given identical tasks and their work is compared: what they build, what breaks, what it costs.
Each experiment uses written-in-advance prompts, a fixed stack, and an independent test suite the agents never see. Transcripts, diffs, timings, and costs are all recorded. The results are small case studies, not benchmarks.
Experiments
Slotbook run 1 · in progress
Claude Code and Codex build the same scheduling app, one milestone at a time.