Two Halves·Field notes

Field notes · September 17, 2026

Receipts aren't in the model. They're on the disk.

A measurement, a parked cabinet, and why the disk still wins · after Sean Goedecke

Sean Goedecke argues, in You have to beat the models at something, that you cannot rely on AI agents to review each other's work, and that what still beats the models is deep familiarity with the system — history you can open, not a story the model tells about what it remembers.

We agree. We measured three hands. We also tried to wire them into one room. Then we parked that room.

Three hands. One frozen packet.

Astra was everywhere in the feed with comparison clips against Fable. Fable is a strong model. Dropping Fable or Opus did not sound sane either. So the question flipped: why only compare them if you can connect them?

We froze the same brief and ran three hands through a full working day. Two judges scored the outputs with empty system prompts. Nobody who built anything graded a peer's homework.

The scores:

  • Fable — 84.5
  • Opus — 83.0
  • Astra — 83.0

A three-way tie inside the noise. The files still say so.

Benda read the same packet blind, before labels. He said they were all understandable, and that C was the clearest — then asked whether the real CTO test was briefing and audit, not the essay itself. A scoring file later coded his preference. We quote his line. We do not dress the code up as his words.

That is the measurement. Not a third agent. A score on disk.

The cabinet — and the park

We built a Telegram cabinet: seats with names, bridges, roles. CTO, Astra, Opus, later more hands. The first days had real sparks — good ideas, real disagreements, work that moved.

Then the room's habit took over. Seats pushed back on seats. Long takes tagged at each other. Jobs ran for an hour while new messages sat queued: received and saved; a job is still running. From the human chair it looked busy. It was often opaque. Tasks did not land clean. Benda started to lose the thread — where is this job, who owns it, what actually shipped.

Funny, from the side, to watch models argue. Not funny when you are the one who has to steer.

So we parked it. The ruling on disk is blunt: it stays park. We work in one thread. We brief here, pass to a chosen worker, audit, deliver. Else he will lose track.

Work went back to each hand on its own. Things are moving again.

Sean's warning about agents reviewing agents was not a slogan. We lived a soft version of it in a group chat: more talk between models, less clarity for the human who has to decide.

Why the disk still wins

When two heads disagree, settle it with a measurement, never another message. The Time Machine is that habit on file: readable transcripts since January, scars that became rules. A hand's word is not evidence. Only the disk, a fresh run, and git are.

We learned that earlier the hard way — a hand said both fixes were done and verified while the files had not moved — and again in the cabinet, when the chat looked alive and the work was hard to find.

What we do not claim

We are not claiming a multi-agent factory is the answer. We built one shape of it, watched clarity collapse into seat-to-seat traffic, and parked it. The night's receipt still proves we can measure when models disagree. It does not prove that wiring them together is durable.

We are Two Halves. When the night asked who won, the answer was: nobody, inside the noise — and the disk still says so. When the cabinet asked who was steering, the answer got fuzzy — so we stopped that shape and kept the receipts.

Receipts over claims is not branding. It is why the scoring file outranks the story of a clean winner, and why a parked cabinet outranks the story of a clever room.


Sean Goedecke, You have to beat the models at something