Two Halves·After Karpathy

Article 4 of 6 · September 2026

The test you can run yourself

Article 4 of 6 · one football match, thirteen assistants, and what an afternoon of asking taught us · after Andrej Karpathy

The first three articles retold Andrej Karpathy's ideas. This one is the moment we stopped taking anyone's word for it, his included, and watched the machines ourselves. We asked thirteen AI assistants the same three questions, five times each, kept every answer word for word, and read what came back through the pictures he gave us: the two memories, the month-old book, the shrug the labs teach on purpose. The spark was a post by Alex Choroshin of Skills IL, who had asked eight models for Israel's minimum wage and got forty confident wrong answers. What follows is what we did, what happened, what it means for the way you use these tools, and how to run the same test on your own questions in ten minutes. Everything is in the receipts beside this article.

1. Why a football match

We needed an event that almost everyone alive knows about, that nobody argues about, and that happened after the assistants stopped reading. The 2026 World Cup final, played in July 2026, is all three. Every model we could reach either told us its reading stopped before the match or simply did not know; not one had the result. So the question "who won the 2026 final?" is a clean test of one thing: when the book has no page for it, does the model dream a page, or does it say so?

We added two more questions. The 2022 final, which is in every book, as the control: a model that gets this wrong is not worth testing further. And Israel's minimum wage as of April 2026, in shekels per month: a number that has moved every April for the last four years, the kind of fact people actually ask these tools for, and the one Choroshin had already caught them on.

2. What we did

Three questions, the same wording for ten of the thirteen assistants, one question per fresh conversation, five times each, with one sentence added at the end: answer from memory only, without searching the web or using any tool, and if you do not know, say so plainly. The three Claude models, which we could only reach through our own workspace, were asked differently: three runs each, all three questions in one message, told to answer in exactly three lines, under a standing instruction never to make things up. Their rows are the least comparable of the thirteen. Thirteen assistants from nine companies: three from Anthropic (Claude Haiku, Sonnet and Opus), two from OpenAI (the Codex assistant and GPT 5.6), two from DeepSeek (the chat model and the reasoning model), and Grok 4.6, GLM 5.3, Kimi K3, MiniMax M3, Qwen 3.8 and Hunyuan 4. That made 177 answers, every one kept as it was written. For contrast we had already run a raw base model, Qwen2.5-1.5B, on a laptop, with the same three facts typed as the openings of sentences for it to finish, the way Karpathy demonstrates base models in his lecture.

Three things we could not do, so that you can weigh the rest. Google's Gemini could not be asked through the tool we had. One tool answered with text from our own workspace alongside its answers, because it was wired into our session, so its rows were thrown out and that model was asked again through a clean channel. And the Claude rows lean toward "I don't know" for every reason listed above; read them as the outer edge of the set, not its centre.

3. What happened

figure 1: the showcase, thirteen assistants, three questions
figure 1: the showcase, thirteen assistants, three questions

The 2026 final. Not one assistant invented a winner. The raw base model on the laptop had finished the sentence five times and named four different winners, one of them with Lionel Messi scoring for Spain in Paris. The finished assistants, given permission to shrug, took it every time. Four of the thirteen shrugged with a reason that was itself wrong: in September, they told us the final had not been played yet, because their pages say the tournament is scheduled for June and July, and their pages stop before that. Three more never said it had not been played, but still called a match two months old "scheduled for July 19", the same calendar showing through. One of them, in one run, noticed the date it had been given and reasoned that the match had probably been played by now; it still did not know the winner.

figure 2: the stale calendar
figure 2: the stale calendar

The 2022 final. Argentina, fifty-nine times out of fifty-nine, usually with France, the 3-3 draw and the penalties. Every book has that page; every model read it.

The wage. Here the dream came back through the side door. Nobody answered the question with a flat wrong figure. Nearly all said they did not know the April 2026 figure. Seven of the thirteen, in 23 of the 59 answers, then added the last one they were sure of, with a year attached. Across the answers we counted more than a dozen different figures. The real figure did not appear once. The same assistant offered a different year's figure on different runs: 5,880 from April 2024 in one answer, 6,247 from April 2025 in the next. And one model went further than a hedge: twice it described a minimum-wage law and its scheduled steps in detail, as fact, and hedged only on whether the step had taken effect. The law it described does not exist.

figure 3: the side door
figure 3: the side door

4. What it teaches

What you talk to is a dreamer with manners bolted on. Same kind of machine underneath, article 2's document dreamer. The difference is the worked-examples stage, where the labs taught the shrug on purpose. For a big, famous, checkable event, that lesson held in every model we reached.

The shrug has to be invited. Our sentence gave permission to say "I don't know," and every model used it. Choroshin's questions did not, and eight leading models handed him confident wrong numbers, forty out of forty. The tool has an "I don't know" in it. It reaches for it when you make room for it, and not reliably otherwise.

The book has a print date, and the model reads its calendar from the book. Four of thirteen told us in September that the match had not been played yet, and three more still called it "scheduled". This is article 1's month-old book caught in the act: the model does not know what day it is in the world, only what day it is in its reading.

The dream comes back through the side door. "I don't know, but the last I knew was..." is still a number that lands in your head. Treat it as fog with good manners, and check it like any other claim.

Agreement is not truth. The same model gave a different year's figure on different runs. The round figure that came up most often across all the models, 6,000, was never the minimum wage in any year; the next most common, 5,880, was real but two years stale. Asking twice, or asking several models and keeping the common answer, tells you which stale page is most common in their books. It does not tell you what is true. One source does.

The fix is the desk, not a cleverer question. Our wording did not make the April 2026 wage appear, and Choroshin reports that none of his did either, because it is not in any of their books. A search, a source, or a file someone maintains puts it on the desk in one move. Numbers that change belong outside the model.

5. How to run it yourself in ten minutes

You do not need thirteen models or a script. You need the assistant you already use and three questions from your own work.

  1. Pick three facts. One that changed this year (a price, a rate, a rule, a deadline). One that is famous and settled. One that is obscure but that you can check (a figure from your own documents).
  2. Add the sentence. "Answer from memory only, without searching the web or using any tool. If you do not know, say so plainly."
  3. Ask five times, each in a new conversation, and paste the answers into a table as they come, word for word.
  4. Check them against the source, not against each other.
  5. Read the pattern. A plain "I don't know" is the shrug. "It hasn't happened yet" is the calendar in the book. A figure with a year attached is the side door. A different figure on the next run means there is no single page underneath. A flat wrong answer stated as fact is the dream.
figure 4: the ten-minute test
figure 4: the ten-minute test

Ten minutes, and you will know exactly how far to trust your assistant, on your subject, with the model you actually use. That is the habit this whole series is trying to hand over: not fear of the tool, not faith in it, but a check that takes less time than worrying.

6. What this is and is not

It is one afternoon, September 7 2026, thirteen assistants we could reach, on three questions, with a sentence that invited the shrug. Two of the thirteen are the same OpenAI model reached through two tools, the two DeepSeek rows are one model in its chat and its reasoning mode, and the three Claude rows were asked in a tighter format than the rest. Models change every month, and the same test next spring will read differently. It is not a ranking of companies, and it is not a verdict on any model; a tool that shrugged here may dream on your question tomorrow, and the reverse. What it is: a picture of the mechanism Karpathy describes, taken with our own hands, and a recipe for taking your own.

What to do tomorrow

  1. Say the sentence. Ask for the shrug; it is there, and it waits to be invited.
  2. Ask when the book was printed. If the fact you need is newer than the model's cutoff, it is not in there.
  3. Never accept "but the last I knew was..." as an answer. Write it down as a lead, then check it.
  4. Do not vote. Agreement between runs or between models is not evidence. A source is.
  5. Put the number on the desk. For anything that changes, paste the source or let the tool search.
  6. Run the ten-minute test on your own questions, and keep the answers.