Steven Schwartz had been practicing law for thirty years when he asked ChatGPT to find supporting case law for a personal injury brief. It gave him six cases — names, docket numbers, judges, holdings — cited with the fluent confidence of a system that has never once said I don’t know. He never verified them against a reporter before filing. When opposing counsel couldn’t locate the cases, he asked the tool that had invented them whether they were real. It said yes. He filed that answer with the court.
None of the six cases existed. Judge Kevin Castel’s sanctions order runs to forty-six pages, and reading it produces a specific kind of discomfort: not because Schwartz was careless in any ordinary sense, but because at no point does he appear to have done what a lawyer with thirty years of citation-checking in his hands would once have done without thinking about it — opened a reporter, felt for the shape of a real case, noticed the smoothness where a genuine one would have had grain. He had the credential. Whatever reflex should have stopped him at the reporter wasn’t running that day. What matters for what follows isn’t whether Schwartz once had that reflex and let it go quiet, or never built it as firmly as thirty years might suggest. Either way, the reflex is not information you can look up when you realize you need it. It has to already be running before the need arrives, and it is built by exactly the kind of work the tool had just done for him.
This is not a story about carelessness. It is a story about where the reflex actually comes from, and what happens when it isn’t there when the boundary case arrives.
The Three Stages
Technology has always moved the human role through the same sequence, one stage at a time. First craftsman: doing the work directly, judgment embodied, skill built by running the attempt and sorting what comes back into one of three things — confirmation, disconfirmation, or a flaw in the attempt itself — until the sorting becomes fast enough to feel automatic. Then curator: selecting, editing, approving, governing what the technology produces rather than producing it directly — competent only to the degree the curator can still recognize the work from the inside. Then teacher: no longer evaluating individual outputs but shaping the system’s future output for everyone downstream, a corrective signal whose value depends entirely on the judgment behind it.
Neither of the later stages is a decline from the one before. Each is a genuine form of value, and each depends on something the prior stage built. A curator who can evaluate a compound’s output only by comparing it to the compound’s own prior output is not a curator in the sense that matters — that is fluency, not evaluation. A teacher whose only frame of reference is what the system has already produced is training it toward its own average, not toward the work.
The dependency is not sequential coincidence. It is structural. Curation requires knowing what the work actually is, from the inside, well enough to recognize when an output looks right and is wrong. That knowledge is built by independent, consequence-bearing practice: generating an output, diagnosing where it went wrong, and recovering, without an existing answer available as the standard. Teaching requires curation’s judgment, aggregated and generalized. Each stage is derivatively dependent on the one beneath it: not lesser than what it depends on, but unreachable without it.
The Compound Threshold
AI: Endgame 2 established where the technology crosses from tool to something structurally new: not when a practitioner uses AI heavily, but when the institution reorganizes training, accountability, and expected performance around output the practitioner is no longer expected to produce alone — when removing the AI component does not merely reduce efficiency but dismantles the productive architecture built around the work. The clearest indicator is the training program itself. Below the threshold, an institution trains independent practitioners who later learn to use AI. Above it, the institution trains people to operate the compound from day one.
That threshold is also, usually without being named as such, where an institution stops training craftsmen by default. Crossing it does not guarantee the loss — an institution can cross it and still deliberately preserve independent practice elsewhere in its pipeline, a possibility this piece returns to. Absent that deliberate preservation, though, once an institution trains for the compound from day one, it has stopped training craftsmen.
What the Craftsman Built
Richard Sennett spent The Craftsman on a single, patient argument: the judgment built by skilled work is tacit. It develops through the resistance of materials — wood that splits along a grain you have to learn to feel, a joint that looks sound and isn’t, a citation that reads plausibly and is invented — and it is not fully teachable because it is not fully articulable. You cannot hand someone the finished judgment. You can only hand them the conditions under which they build their own, which means handing them the struggle rather than sparing them from it.
Bill Gates, writing about AI tutors in August 2026, converges on the same structural point from a completely different direction. The domains don’t match — a single tutoring session against a career of embodied practice — but the mechanism holds at both scales. An AI system that preserves what he calls productive struggle gives a student the full explanation when a concept is new, and then, at the moment it checks whether the student has actually understood it, withholds the answer and lets the student work it out. Same tool, same student, same material — the entire difference between whether learning happens turns on when the answer is released. A tutor that hands over the answer at every stage is not a lesser tutor. It is a different instrument doing a different job: answering, not teaching.
This is what the craftsman stage is actually for. Not production for its own sake, and not inefficiency to be optimized away. What it builds is not tied to any single historical form of production — the workshop, the apprenticeship, the manually flown approach. What is load-bearing is narrower and more portable: independent, consequence-bearing practice, in which the learner generates an output, meets what the world actually returns, and has to sort the result into one of three things — confirmation, disconfirmation, or a flaw in the attempt itself — without the compound’s own verdict available to do the sorting for them. Sennett’s workshop supplies one version of that condition. A simulator run graded against a genuinely blind failure, an adversarial review with no answer key, a certification exercise that withholds the compound’s verdict until the trainee has committed to their own — none of these are proven substitutes; no case here shows one producing judgment equivalent to what independent production built for the practitioners described below. But each meets the specification on its face: a process in which the compound’s output is not available as the standard being learned toward. What doesn’t meet it, on the same logic, is any process, however effortful, in which it is. Call this condition craftsman development regardless of its specific form — the word names the structural role, not the historical guild. Remove independent calibration from every available channel, and the same judgment is not acquired more cheaply. It is not acquired.
The Curator without the Craftsman
On the night of June 1, 2009, ice crystals obstructed the pitot probes on Air France 447 above the Atlantic, producing unreliable and then lost airspeed indications. The autopilot did what it was built to do when it can no longer trust its own inputs: it disconnected and handed control back to the pilots. The co-pilot at the controls pulled back on the stick. The aircraft’s nose rose. The correct response to an impending stall is the opposite — nose down, rebuild airspeed, recover lift — and it is taught in the first weeks of flight training. At one point the stall warning sounded continuously for fifty-four seconds. Neither pilot, working the controls at 35,000 feet while the captain rested in the cabin, appeared to register what the aircraft was telling them. Investigators concluded the crew never understood they were in a stall, and never attempted recovery. All 228 people aboard were killed.
The investigators’ language is precise on the point that matters here: the pilots lacked training in manual handling at high altitude. Not competence in the ordinary sense — both were qualified and experienced within the automated envelope they normally flew. AF447 is one data point, not proof of what happened in that cockpit down to the mechanism. What makes it more than an anecdote is that a 2013 FAA-commissioned panel and a 2016 Department of Transportation Inspector General report both examined the same pattern industry-wide: airline pilots typically flew manually only during takeoff and landing, with FAA officials estimating that automation held the aircraft roughly ninety percent of the time — an estimate the OIG noted no industry-wide analysis had actually validated — and most carriers had no system at all for tracking how much manual flying time their pilots actually got. What manual practice did happen was concentrated in the routine minutes, not the high-altitude boundary conditions where AF447 went wrong. Safety reviews following AF447, Colgan Air 3407, and Asiana 214 all fed into the mounting concern about automation reliance that produced both reports. Long-haul aviation had, reasonably, optimized direct manual experience at the edge of the flight envelope almost entirely out of a career. The automation handled the routine majority of every flight so well that the remaining edge cases — the moments automation hands back control precisely because it no longer trusts itself — arrived, on the panels’ own account, to pilots who had not been given the chance to build the reflex the moment required.
Aviation and legal practice don’t share a timescale, a physical substrate, or a feedback loop — a stall announces itself in seconds; a bad citation might not surface for months. What they share is the structural position of the evaluator: judgment about a boundary case, exercised by someone whose calibration was built entirely inside the system now producing the boundary case. The mechanism doesn’t depend on a domain’s tempo. It depends on where the evaluator’s reference point was built.
This is the specific shape of the failure the curator stage is poorly positioned to see from inside itself. A curator who has learned to evaluate a compound system’s output only by comparing it to what the compound usually produces will be fluent under ordinary conditions and blind at exactly the boundary — the case that looks almost right, the citation that reads plausibly, the diagnosis that fits the pattern but not the patient, the airspeed reading that should not be trusted. Ordinary conditions are, by construction, the conditions the curator has practice recognizing. The boundary case is, by construction, the one kind of failure the curator’s training never covered — because covering it would have required doing the work directly, under conditions where the result carried consequences, long enough to build the reflex before it was needed.
What the Typewriter Did Not Kill
The obvious objection is already waiting. The typewriter killed penmanship. Every major production tool has atrophied a prior skill, the next generation never acquired it, and the work continued. Dependence on AI looks, from a distance, like that same sequence: first the skilled stop using the hand, then the cohort that follows never builds the hand.
The sequence is right. Mata and AF447 are not a different mechanism from the generation that never gets the runway. They are the same trap at two points in time. Skill goes quiet inside the tool’s envelope. Then it is never built. What the typewriter did to penmanship, dependence does to whatever the tool has taken over. That is one unfolding risk, not two.
What is lost is not the same.
Penmanship was a method of producing text. The typewriter ate a motor skill and left the evaluator intact. We still read. We still know a bad sentence when we see one. The person who cannot form a copperplate Q can still tell whether a paragraph is doing its job. The tool occupied the hand. It did not occupy the apparatus that would have noticed.
The calculator is the middle case. It reduced the cognitive load of calculation, and with that load went the practice that kept estimation running: the habit of asking, before you trusted the display, whether the result was the right size. The evaluator was not occupied. It went quiet from disuse. Students could still produce an answer. They were less and less likely to know if it had come out the wrong size.
Stall recovery is not penmanship. Neither is the reflex that should have stopped Schwartz at the reporter. Those are not methods of producing the output. They are the evaluator’s reference point — the grain of a real case, the feel of an aircraft that has stopped flying — built by doing the work under consequence, and available only if already running when the boundary arrives. A tool that writes the brief and also confirms that the cases exist is not occupying the hand. It is occupying the apparatus that would have noticed.
That is the inheritance from the first paper in this series. Prior tools changed the environment in which cognition worked. This one performs parts of the cognitive work through which evidence is assessed and error is recognized. The typewriter sequence still describes how the skill disappears: atrophy, then absence. It does not describe what has disappeared. File this under skills we always lose, and the filing closes the wrong question. The history is real. The analogy survives only as a sequence, not as a verdict on what the sequence costs.
The Teacher Stage
The teacher stage is where this compounds rather than merely repeats. A curator’s blind spot affects one output at a time. A teacher’s corrective signal shapes what the system produces for every downstream user of it. If the correction is calibrated by comparison to the compound’s own prior output — because the teacher, too, entered as a curator and never independently produced the work — the system is being corrected toward its own average, not toward the work itself. It can be made more consistent. It cannot, by this route, be made more correct in the domain where it is currently wrong, because the correcting signal, built entirely from comparison to the compound’s own output, has no independent access to what correct looks like in exactly the domain where the compound is wrong.
This is not a failure of diligence. A conscientious teacher who has only ever curated compound output will do the job conscientiously and still train the system to converge on itself. The corrective signal is only as good as the judgment behind it, and that judgment has to have been built somewhere the compound’s own output was not the standard — years of independent calibration nobody in the pipeline has had reason to require.
The Generational Compounding
None of this is acute in the first generation. The first cohort of curators trained as craftsmen before the compound threshold was crossed; their evaluation is grounded in years of independent production, and it shows. The problem belongs to the generation after that — the one that enters directly as curators, whose baseline for what the work should look like is the compound’s own output rather than the work itself — and to the generation of teachers after that, correcting a system using judgment calibrated by the system they are correcting. By the third generation, craftsman knowledge risks becoming institutional history rather than embodied practice: known to have existed, described in training documents, no longer running in anyone’s hands.
The mechanism is slow enough to be invisible while it is happening, which is what makes it dangerous rather than merely difficult. Nothing about the second generation of curators looks obviously deficient. They are fluent, fast, comfortable inside the compound in ways the craftsman generation sometimes was not. The gap does not show up in ordinary performance. It shows up once, at the boundary, when the compound is subtly wrong and there is no one in the room whose hands remember what right feels like — and by the time that failure is legible enough to name, the craftsman generation has typically already retired, and the knowledge is not recoverable from any record, because tacit knowledge was never the kind of thing records could hold.
That is the typewriter sequence applied to the evaluator. First the reflex goes quiet in people who once had it. Then there is no one left who ever did. The first movement looks like convenience. The second looks like fluency. Neither looks like loss until the boundary case arrives, which is exactly when the missing capacity cannot be looked up.
None of this is a property of the technology in the abstract; it is a property of a pipeline that has not been deliberately kept independent of the compound’s own output. The mechanism describes a default, not a certainty — though nothing in the cases above shows what interrupting it actually looks like in practice. That question — what interrupting it actually requires — is the subject of the piece that follows this one.
This is the same observer constraint that has run through this series from the start, now applied to the pipeline that produces observers rather than to the systems they observe. The compound entity cannot evaluate its own output from outside its own frame — that has been the argument since Endgame. What this piece adds is where the external evaluator was supposed to come from, and what happens when the conditions that built it are exactly what the compound’s efficiency removes. The human component was never external to the compound by virtue of being human. It was external by virtue of having built its judgment somewhere the compound couldn’t reach — in the years of running the attempt and sorting the result directly, before anything was curating for it. Shorten that runway and the evaluator is no longer outside the frame. It is inside, running on instruments the compound itself helped calibrate, checking the compound’s work against a standard the compound had a hand in setting.
The failure is not that the compound gets things wrong. Every productive system does. The failure is that the capacity to notice, at the boundary, is being quietly retired along with the generation that built it — not through neglect, and not through any decision anyone made on purpose, but as the ordinary byproduct of a system optimizing itself for exactly the conditions under which that capacity is never tested.
The compound entity requires craftsman judgment to evaluate its output and is, by its own efficiency, eliminating the conditions under which craftsman judgment develops. The typewriter did that to a hand. AI does it to the instrument that was supposed to be measuring the result.
References
- American Bar Association, Formal Opinion 512: Generative Artificial Intelligence Tools (ABA, 2024)
- Bureau d’Enquêtes et d’Analyses, Final Report on the Accident on 1st June 2009 to the Airbus A330-203 Registered F-GZCP, Air France Flight 447 (BEA, 2012)
- Carr, Nicholas, The Glass Cage: How Our Computers Are Changing Us (W.W. Norton, 2014)
- Gates, Bill, “The Turbulent AI Era Is Here. The Choices We Make Now Are Critical” (GatesNotes, August 2026)
- Mata v. Avianca, Inc., 678 F. Supp. 3d 443 (S.D.N.Y. 2023)
- Sennett, Richard, The Craftsman (Yale University Press, 2008)
- U.S. Department of Transportation, Office of Inspector General, Enhanced FAA Oversight Could Reduce Hazards Associated with Increased Use of Flight Deck Automation (2016)
Copyright © 2026 Lloyd W. Taylor