The journal

Did the work come back marked.

A randomised trial with nearly a thousand students shows what separates work that built judgment from work that only produced an answer, and the test it yields fits on a single memo.

A board approving an efficiency is rarely told which kind of efficiency it is. Work that was forming judgment and work that was only consuming time look identical on a process map, and on a headcount line. When the machine absorbs either one, the saving books this quarter. Nothing in the paper separates a healthy simplification from a drawdown of the capacity the institution runs on.

The book’s blade for this comes from two psychologists who spent their careers disagreeing. In 2009 Daniel Kahneman and Gary Klein published “Conditions for Intuitive Expertise: A Failure to Disagree” in American Psychologist. They agreed on two conditions for trustworthy intuition: an environment regular enough to hold real patterns, and enough opportunity to learn them through prolonged practice with rapid and unequivocal feedback. Chapter Four turns that into a question a board can ask before funding anything called formation. Does the world give back honest, structured feedback, so reps compound into judgment — or is the feedback mush, so they compound only into confidence?

Chapter Four asks that of a domain. The move worth making now is to ask it of a task. The unit becomes one piece of work — the memo, the file, the reconciliation — and the question fits a single box on a workflow diagram. Did that piece of work come back marked?

A field experiment tests exactly that. In the autumn of 2023, Hamsa Bastani, Osbert Bastani, Alp Sungu and colleagues ran a randomised trial in a Turkish high school, published last year in PNAS. Nearly a thousand students in about fifty classes sat four ninety-minute maths sessions. Each ran in three parts: the teacher taught, the students worked practice problems, then sat a closed-book, closed-laptop exam on the same concepts. Only the middle part was varied.

Two AI tutors were built on the same model. One imitated the standard ChatGPT interface — GPT Base. The other, GPT Tutor, ran on a prompt written to safeguard the learning. It held the correct solution so it would not invent one, it knew the errors students make on that problem, and it was forbidden to hand over the full solution. One line in that prompt does the work of this whole entry: if the student offers an answer, tell them whether it is right.

With the tools open, both arms improved sharply. GPT Base students scored 48 per cent better on the practice problems than students with no AI at all. GPT Tutor students scored 127 per cent better. In a board pack both read as the same win, and the second as the larger one.

Then came the third part of the session, laptops shut. The GPT Base students scored 17 per cent worse than students who had never had access, in the same session in which they had looked like the improved group. The GPT Tutor students came out statistically indistinguishable from that control group, their point estimate an order of magnitude smaller. The authors are exact about it. The negative effect was “essentially eradicated” in the GPT Tutor arm, “though we still do not observe a positive effect.” The safeguard bought back the drawdown, and that is all it bought.

Same model. Same problems. Same school, same weeks, students randomly assigned. The one thing that differed was how the tool handed the work back. In one arm it came back marked: a hint that made the student produce the next step, and a verdict on whether the step was right. In the other it came back done.

The formation literature named this requirement long ago. Anders Ericsson, Ralf Krampe and Clemens Tesch-Römer’s 1993 account of deliberate practice sets out what a repetition must contain before it builds anything: a well-defined task at a difficulty the person can just about meet, informative feedback, and the chance to repeat it and correct the error. Kahneman and Klein set the condition at the level of the domain. Ericsson sets it at the level of the single rep.

An older warning sits underneath this. In 1983 Lisanne Bainbridge published five pages in Automatica called “Ironies of Automation.” Automate the routine operation of a plant, she argued, and you leave the operator the tasks nobody could reduce to a rule, using skills the rest of the job no longer lets them practise. Bainbridge wrote about a console and an operator. Carrying her mechanism up to how an institution forms judgment across a whole bench is our argument, and it stands or falls as one.

Chapter Four gives a board three demands to make of any place AI has been let near a judgment. Name the cue, “so the tacit skill stops being magic and becomes a target.” Make the feedback real and fast, “so each rep comes back marked.” Keep the reps near where the judgment forms, “in the hands of the people who will one day hold it.” The board’s work is to govern that structure, and the training itself belongs to someone else.

That makes it runnable this week, on a real piece of work. Take one thing the machine now does or assists: the first-pass credit memo, the contract review, the customer call. Find the person who would have done it two years ago. Ask what told them whether they had got it right, and how long that signal took. Then ask what tells them now.

Most organisations have an answer, usually short. Someone still reads the draft and says where it was wrong, within a week, and the writer will be there in ten years to hold that judgment — in which case the saving is a saving. Or the work goes out with nobody’s mark on it, and this quarter’s productivity has been bought with a decade of formation, on a bill that reaches the succession line long after the person who approved it has gone. The question fits in the margin of any paper proposing to automate something. Did the work come back marked?

The list

Entries arrive weekly. Leave an address and they come to you, along with the publication date when there is one.