Chapter 02 of 30

Never Stand in Front of the Steamroller

Concepts

CHAPTER 02 β€” Never Stand in Front of the Steamroller

STATUS

Full first draft

EDITORIAL PASS (2026-09-14)

  • Closure pass 2026-09-14: added C/T exposed-count scoring sketch (2 of 5 illustrative tasks exposed); reader can now compute exposure and name the boundary. No new studies cited.
  • “How to read the rest of this book” promised four MARKED claim strengths; no chapter uses those marks. Rewritten to describe the real convention: strength stated in prose early; from Ch19, explicit evidence tags (measured / demo / reported / source / owed).
  • Stale refs: “Part IV spends five chapters” -> Part 5; “failed to replicate” -> “did not survive replication” (Ch27 tie, signal not promoted). “Chapter 2’s four jobs” inside Ch2 -> “this chapter’s”.
  • Removed out-of-format “Research anchors” appendix (duplicated references, internal-note phrasing). Added missing reference entry for Autor et al. NBER w35720 (verified in Ch5 pass: 133 lawyers, +0.34/+0.38 SD, unaided +0.32 overall / +0.45 seniors, juniors bifurcated).
  • Spelling: travelled, labelled (x2) -> traveled, labeled.
  • Score ~900 -> ~950.

CENTRAL QUESTION

Given that the ground is moving, what can you actually stand on?

SECTION OUTLINE

  • Open with the book’s epistemic status as a dated position β€” demonstrated, not asserted, via METR’s own 2025 result and its 2026 self-correction.
  • State the steamroller claim and define it precisely: tacit β†’ codified β†’ automated β†’ commodity.
  • Evidence from payroll microdata, reported with the authors’ own caveats.
  • Why “be excellent at the codified part” fails: compression, not replacement.
  • Where the ground is solid: the jagged frontier, and the measured cost of not knowing where it is.
  • The four jobs the model cannot hold β€” intent, authority, verification, frontier judgment β€” read off Chapter 1’s architecture diagram.
  • The unification: the career argument and the architecture argument are the same argument, and the rest of the book is the engineering discipline for those four jobs.
  • The chapter’s own best counter-arguments.
  • How to read the book: claims typed as measured / inspected / argued / predicted.

LOAD-BEARING CLAIMS

  1. Being wrong responsibly means publishing the correction as loudly as the finding. METR is the worked example, not a hypothetical.
  2. The steamroller is the codification pipeline, not “AI”. The danger is holding a role whose value is the codified part.
  3. Getting better at the codified part does not escape it β€” AI compresses the skill distribution by redistributing the best workers’ codifiable expertise.
  4. The four jobs outside the model appreciate rather than depreciate as capability grows, because better models increase the volume of output needing governance without supplying any.
  5. Applied AI is the engineering discipline of the position you want to be standing in. This is what makes the chapter load-bearing rather than an op-ed.

PAPERS / EVIDENCE

  • Becker, Rush, Barnes & Rein (METR), arXiv:2507.09089, July 2025. 16 experienced OSS developers, 246 real tasks in own repos (>1M LOC, >22k stars), randomized AI-allowed/disallowed. Devs forecast βˆ’24% time beforehand, estimated βˆ’20% afterwards; measured +19% time. 39-point belief/measurement gap.
  • METR, “We are Changing our Developer Productivity Experiment Design”, 2026-02-24. Selection bias: developers declining the no-AI condition even at $50/hr; 30–50% avoided submitting tasks they expected AI to accelerate. New estimates βˆ’18% (CI βˆ’38% to +9%) returning, βˆ’4% (CI βˆ’15% to +9%) new; both cross zero. METR’s own words: “only very weak evidence”; developers likely more sped up in early 2026.
  • Brynjolfsson, Chandar & Chen, Stanford Digital Economy Lab, Aug 2026 revision (ADP payroll, millions of US workers, data through June 2026). Six facts. (1) No economy-wide displacement. (2) Ages 22–25 in AI-exposed occupations 19% below kept-pace; experienced workers no comparable gap. (3) Widened from 15% (July 2025 vintage) to 19% (June 2026). (4) Operates through reduced hiring, not separations. (5) Concentrated where AI substitutes; complementary usage flat or rising. (6) Adjustment via employment not base pay. Levels: two most-exposed quintiles βˆ’11% Nov 2022β†’June 2026; three least-exposed +10%. Mechanism: declines in occupations involving codified knowledge; tacit-knowledge occupations see faster growth for experienced workers.
  • Brynjolfsson, Li & Raymond, QJE 140(2) 2025, pp. 889–942. 5,179 customer support agents, staggered rollout. +14% issues/hour average; +34% novice/low-skill, minimal for experienced/high-skill. Mechanism: disseminates best practices of the most able workers.
  • Dell’Acqua et al., HBS WP 24-013, 2023. 758 BCG consultants (~7% of ICs), preregistered. Inside the frontier: +12.2% tasks completed, 25.1% faster, higher quality. Outside it: 19% less likely to be correct than non-users β€” unwarranted trust in confident wrong output. “Jagged technological frontier.”

CONTROLS / LIMITATIONS

Canaries: authors label these descriptive, not causal; attenuate under education controls; some divergent trends predate ChatGPT (COVID era); larger in ADP sample than national survey benchmarks; no economy-wide displacement. Generative AI at Work: one firm, one occupation, unusually codifiable expertise and unusually clean metrics β€” direction transfers, not the 34%. METR: superseded by its own redesign; cite both or neither. The steamroller claim is typed PREDICTED, the book’s weakest claim category, and is marked as such.

THE CHAPTER’S OWN COUNTER-ARGUMENTS (do not cut these)

  • The broken ladder: senior judgment was historically acquired via junior codified work; “reduced hiring rather than separations” removes the bottom rungs. No good answer offered; named as structural.
  • Composition: advice that works for one person is arithmetically impossible for a cohort.
  • Verification may automate further than this book bets. It is a bet, not a theorem.
  • The evidence is young and mostly descriptive; METR’s reversal is the live demonstration.
  • The metaphor invites fatalism; the point is that where you stand is a decision.

DEPENDENCIES

Chapter 1 β€” the obligation diagram and the human = intent + authority assignment, which this chapter re-reads as a map of where to stand.

FORWARD BRIDGE

Three of the four jobs are engineering problems the book builds. The fourth, frontier judgment, needs a rule before it can be built into anything β€” Chapter 3 supplies it.

Explain this chapter with AI

Copy this prompt into ChatGPT, Claude, Gemini, a local model, or another AI.

Apply this chapter with AI

Copy this prompt into ChatGPT, Claude, Gemini, a local model, or another AI.

Part 1 β€” Where You Stand

A dated position

This book will be wrong about things.

Not as a disclaimer. As a statement about what kind of document this is. Everything here is a view from one position on a hill that is still being climbed. From where we are standing in late 2026, certain shapes are visible and certain ones are not, and some of what currently looks like a permanent feature of the landscape will turn out to have been a cloud.

The useful question is not whether a book about AI will age badly. It is whether it ages badly in a way you can detect and correct, or in a way that quietly misleads you for two years.

So rather than open with a hedge, here is what being wrong responsibly actually looks like, from a lab that does this well.

In July 2025, METR published a randomized controlled trial on AI coding tools. Sixteen experienced open-source developers, 246 real tasks in their own repositories β€” mature projects averaging over a million lines of code β€” each task randomly assigned to allow or disallow AI assistance. The developers forecast beforehand that AI would cut their completion time by 24%. Afterwards, having done the work, they estimated it had cut their time by 20%. Measured, allowing AI increased completion time by 19% (METR, 2025).

That result traveled a long way, and deservedly: a 39-point gap between what skilled practitioners believed about their own productivity and what was measured is a serious finding.

Then, in February 2026, METR published an update saying their experiment design had a problem. Developers were increasingly declining to participate in the no-AI condition, even at $50/hour, and between 30% and 50% of participants were avoiding submitting the very tasks they expected AI to accelerate. That is selection bias pointing directly against the headline. Their newer estimates moved the other way β€” around βˆ’18% time for returning developers (confidence interval βˆ’38% to +9%) and βˆ’4% for newly recruited ones (βˆ’15% to +9%) β€” with both intervals crossing zero. METR’s own summary is that developers are likely more sped up in early 2026 than their early-2025 estimate suggested, and that the new data is only very weak evidence (METR, 2026).

Read what happened there. A careful lab ran a clean experiment, got a striking number, published it, kept looking, found the flaw themselves, and stated plainly how weak the replacement evidence is. Nothing was retracted and nothing was defended past its evidence.

That is the standard this book holds itself to, and it is why Part 5 spends five chapters on experiments, including one in which a result the author wanted to be true did not survive replication. The durable lesson from METR is not the 19%. It is the 39-point gap between belief and measurement β€” which survived the design correction, and which is exactly why this book will not let a model, or a person, grade its own homework.

Given that the ground is moving, what can you actually stand on?

The one thing that looks clear

Here is the claim this chapter exists to make.

Never stand in front of the steamroller.

The steamroller is not “AI.” That is too vague to act on. The steamroller is a specific, repeating process:

    flowchart LR
    T["tacit<br/><i>known by doing</i>"] --> C["codified<br/><i>written down, teachable</i>"]
    C --> A["automated<br/><i>executed by machine</i>"]
    A --> M["commodity<br/><i>priced near zero</i>"]
    style A fill:#00000000,stroke-dasharray: 4 4
  

Knowledge starts tacit. Over time it gets codified β€” written into documentation, patterns, Stack Overflow answers, training corpora. Once codified, it becomes automatable. Once automated, it is priced accordingly.

This process is old. What changed is the speed of the last two steps, and the fact that a model trained on the codified layer can execute it immediately across every domain at once, without anyone having to build a domain-specific tool.

To stand in front of the steamroller is to hold a role whose value is the codified part. If what you are paid for is knowing things that are written down and executing procedures that are described, you are standing on the section of road the machine is currently flattening.

This is not a prediction that programmers will be unemployed. It is narrower and better supported than that, and the narrow version is the one worth acting on.

What the payroll data says

Brynjolfsson, Chandar, and Chen have been tracking this in administrative payroll microdata from ADP covering millions of US workers. Their August 2026 revision, using data through June 2026, reports six facts. The relevant ones:

  • There is no evidence of widespread, economy-wide job displacement. That is their first finding, and it should be the first thing anyone quoting this work says.
  • Employment of workers aged 22–25 in the most AI-exposed occupations now stands about 19% below where it would be had it kept pace with similarly aged workers in less-exposed occupations. In levels: employment for that age group in the two most exposed quintiles fell about 11% between November 2022 and June 2026, while the same age group in the three least-exposed quintiles grew about 10%.
  • Experienced workers show no comparable gap.
  • The divergence has widened steadily β€” 15% at the July 2025 data vintage, 19% by June 2026.
  • It operates primarily through reduced hiring rather than increased separations.
  • Declines concentrate in occupations where AI usage substitutes for human tasks. Where usage complements workers, employment is flat or rising, especially for experienced workers.

And the mechanism finding, which is the one this whole chapter turns on:

Employment declines for young workers appear in occupations that involve codified knowledge. Occupations that involve tacit knowledge see faster employment growth for experienced workers (Brynjolfsson, Chandar & Chen, 2026).

That is the steamroller, measured, in the authors’ own vocabulary rather than mine.

Now bound it, because the authors do. They describe these as early, descriptive indicators β€” canaries in the coal mine β€” not causal estimates. The patterns attenuate when controlling for education. Some divergent trends predate generative AI, particularly around the pandemic. The effects are more pronounced in the ADP analysis sample than in national survey benchmarks. And again: no economy-wide displacement.

Note also what “reduced hiring rather than separations” means concretely. The steamroller is not running people over. It is declining to lay track where the next cohort was going to walk. That is a materially different problem, and I will come back to why it is the hardest part of this chapter’s advice.

Why “be excellent at the codified part” stopped working

The instinctive response to all this is to get better. Be a stronger engineer than the model. That response has been measured too, and it fails for a reason that is not obvious.

Brynjolfsson, Li, and Raymond studied the staggered rollout of a generative AI assistant across 5,179 customer support agents. Access raised productivity β€” issues resolved per hour β€” by 14% on average. But the average hides the finding: a 34% improvement for novice and low-skilled workers, and minimal impact on experienced, highly skilled workers. Their suggestive mechanism is that the model disseminates the best practices of the most able workers, moving newer workers down the experience curve faster (Brynjolfsson, Li & Raymond, 2025).

Read that from the perspective of the experienced worker. The tool took the codified portion of your expertise, the part that could be extracted from transcripts of your best work, and distributed it to everyone who did not have it. You gained little. The gap between you and a novice narrowed sharply.

That is not replacement. It is compression. If your market value came from being better than average at the part of the job that can be written down, the model did not take your job β€” it took your differential.

This is why the entry-level effect and the compression effect are the same story seen from two ends. The codified layer is being commoditized. Juniors are affected first because their roles are the most purely codified. Seniors are affected second and less visibly, through the erosion of whatever part of their premium was codifiable.

Bound this one too: one firm, one occupation, a support context with unusually clean productivity metrics and unusually codifiable expertise. Software engineering is not customer support. The direction is what transfers, not the 34%.

A newer result complicates the compression story without removing it. In a pre-registered three-month trial with 133 patent lawyers, AI assistance raised drafting quality (+0.34 SD at 10 days, +0.38 SD at 90 days, larger for juniors) β€” but durable unaided judgment afterwards concentrated in seniors (+0.45 SD), while juniors bifurcated rather than improving on average (Autor et al., NBER w35720, 2026). That is NBER working-paper evidence, one occupation, Google-funded: it supports, within those conditions, that foundational expertise may be a prerequisite for extracting lasting skill from AI-assisted practice. Compression of immediate output and differentiated learning can both be true β€” and the ladder problem below gets harder, not easier, if juniors get the output gain without the judgment gain.

Where the ground is solid: the jagged edge

If the codified part is being flattened, what is not?

Dell’Acqua and colleagues ran a preregistered field experiment with 758 BCG consultants, roughly 7% of the firm’s individual-contributor consultants. On 18 realistic tasks chosen to sit inside current AI capability, consultants using AI completed 12.2% more tasks, 25.1% faster, at significantly higher quality. On a complex task deliberately chosen to sit outside that capability, consultants using AI were 19% less likely to reach a correct solution than those without it β€” because they extended unwarranted trust to confidently-presented, substantively wrong output (Dell’Acqua et al., 2023).

The authors named the shape of the problem: a jagged technological frontier. Some tasks fall easily inside current capability. Others, apparently similar in difficulty, fall outside it. The boundary is irregular and it is not visible from the task description.

So here is the job that has value: knowing where the edge is. And notice its properties. It cannot be codified, because the frontier moves β€” which means it cannot be handed to the model, and it cannot be learned once and banked. It is a standing obligation to check. The same paper shows what happens to people who skip it: they do worse than people with no AI at all.

That is the first of four positions worth holding, and it is the one this book’s Chapter 3 turns into an engineering rule rather than an intuition.

The four jobs the model cannot take

Chapter 1 ended with an assignment of responsibilities: model = cognition, runtime = coordination, tools = action, verifiers = evidence, human = intent and authority. That was an architecture diagram. Read it again as a map of where to stand.

The job What it is Why the model cannot hold it
Intent Deciding what should be done, and what would count as done The model can propose objectives; it cannot be the thing that wants one. Someone must own the definition of success and be accountable for it.
Authority Deciding who is permitted to cause which effects Capability is not authority (Chapter 20). A system that authorizes its own actions has no authorization step, only a delay before one.
Verification Obtaining evidence about reality, not agreement from another model Verification requires contact with the world β€” a test run, a source checked, a result reproduced (Chapter 21). Consensus among generators is not evidence.
Frontier judgment Knowing when the stochastic component is the right tool at all The frontier is jagged and moving; the model is measurably confident on the wrong side of it (Dell’Acqua et al., 2023).

The property these four share is the important one. They do not get cheaper as models improve. They get more valuable β€” because a better model produces more output per hour, and every unit of that output needs intent behind it, authority over it, and verification after it. Capability growth increases the volume of work requiring governance while doing nothing to supply the governance.

That is what it means to not stand in front of the steamroller. Not “move into management,” which is advice about a job title. It means occupying the part of the process whose demand is created by the thing doing the flattening.

And this is the point at which the career argument and the architecture argument turn out to be one argument. Every remaining chapter of this book builds machinery for exactly these four jobs: explicit context and budgets so intent is legible (Chapters 14–15), a ledger and artifact store so what happened survives (Chapters 16–17), claims and evidence levels so support stays distinguishable from assertion (Chapter 18), capability kept separate from authority (Chapter 20), independent verification (Chapter 21), and a router that holds the judgment about when to spend a call at all (Chapter 28).

Applied AI is not a set of tricks for getting more out of a chat window. It is the engineering discipline of the position you want to be standing in.

The same four jobs decide something closer to home. Once an assistant can build software, deciding what your own tools should do, what they may change, and how you would know a change helped is intent, authority and verification applied to your working environment. Chapter 30 returns to that.

Where this argument is weakest

A chapter that opens by promising to be wrong owes you its own best counter-arguments. Here are the ones I find genuinely hard.

The ladder problem is real and this chapter does not solve it. The route to senior judgment has historically run through the junior codified work, where you learned where the frontier was by being wrong about it on small things for several years. If the steamroller removes the bottom rungs β€” and “reduced hiring rather than separations” says precisely that it does β€” then “acquire senior judgment” is not actionable advice for someone who cannot get hired to acquire it. I do not have a good answer, because this is a structural problem requiring a structural response, and individuals absorbing it as a personal failure are misreading it.

The composition problem. Advice that works for one person can be arithmetically impossible for a cohort, since if everyone moves to directing AI, the ratio of directors to directed work becomes absurd. “Get out of the way of the steamroller” describes a smaller destination than the road it is flattening, and anyone telling you otherwise is selling something.

Verification might automate further than I am betting. This book’s central wager is that verification resists automation because it requires contact with reality, which is a bet rather than a theorem. Automated test generation, formal methods, and simulation all push on it, and if verification collapses into the model, a good part of this chapter’s advice collapses with it.

The evidence is young and mostly descriptive. The strongest labor result here is explicitly labeled by its authors as descriptive rather than causal, attenuates under education controls, and runs larger in its sample than in national benchmarks, while the productivity results come from single firms or single occupations. The METR reversal in this chapter’s own opening is a live demonstration that confident readings of this literature can age badly within months.

The metaphor invites fatalism. “Steamroller” can be read as “nothing you do matters.” That is the opposite of the point. The point is that where you stand is a decision, it is available to you, and the evidence about which ground is solid is better than it was two years ago.

How to read the rest of this book

Claims in this book come in different strengths, and the book tries never to let a weaker one borrow the authority of a stronger one:

  • Measured β€” a number produced by an experiment, reported with its sample and its bounds.
  • Inspected β€” an observation about source code that was actually read, identified by file and symbol.
  • Argued β€” a position defended from stated premises, like the determinism rule in Chapter 3.
  • Predicted β€” a claim about how things will go. The steamroller is one of these. It is the weakest category, and everything in it should be held loosely.

In the early chapters the strength is stated in the prose. From Chapter 19 onward, where the book leans on its own experiments, claims about CodeAI also carry explicit tags that separate a pinned measurement from a demonstration that ran but was not pinned, a frozen historical report, code that was only read, and evidence that is still owed.

Where a number appears, it has a source and a stated limit. Where this book proposes a contract stronger than the code currently implements, it says so. Where an experiment failed, the failure is the chapter.

Do this now

Fifteen minutes, one page. Audit your own exposure.

  1. List the five things you were paid to do last month. Be concrete β€” tasks, not a job title.
  2. For each, mark whether its value is the codified part (knowing what is written down, executing a described procedure) or the tacit part (judgment about what should be done, or about whether it worked).
  3. For each codified one, write the sentence: if a model did this at 90% quality for a hundredth of the cost, what would still need me? If you cannot finish the sentence, that is the finding.
  4. Now mark which of this chapter’s four jobs β€” intent, authority, verification, frontier judgment β€” you actually hold, as opposed to being adjacent to.

To turn the audit into a number, score it as a constructed teaching sketch (illustrative tasks, not measured data). Mark each task C (value is the codified part) or T (value is tacit judgment), then count codified tasks where you hold none of the four jobs:

Example task C / T Four-job held? Exposed?
Reset passwords per runbook C none yes
Triage routine support tickets C none yes
Write weekly status summary C intent no
Judge an ambiguous outage escalation T frontier judgment no
Sign off a production migration T authority, verification no

Here 2 of 5 tasks are exposed (2 codified tasks with no job held). The reader can now do one new thing after this chapter: compute their own exposed count and name the boundary β€” any task scoring C with no job held is the section of road to move off first, before any career decision is made on vibes.

Keep the page. Chapter 7 asks you to measure things; this is the only measurement in the book where you are the instrument and the subject.

Failure modes

  • Reading “no economy-wide displacement” as “nothing is happening.” Both facts come from the same paper. The aggregate is calm; the 22–25 cohort in exposed occupations is not.
  • Reading the entry-level data as a reason to disparage juniors. The mechanism is reduced hiring in codified roles, not a judgment about the people who would have filled them.
  • Trusting confident output near the frontier. Measurably worse than not using the tool at all on the wrong side of the edge.
  • Believing your own productivity estimate. Skilled developers were off by 39 points about themselves, in the direction of optimism.
  • Treating “move up the stack” as a completed action. Frontier judgment is a standing obligation to re-check, because the frontier moves.
  • Quoting any number in this chapter without its bound. Every one of them is from a specific sample in a specific period.

What this chapter established

  • This is a dated position, and being wrong responsibly means publishing the correction as loudly as the finding β€” the standard METR set for itself between 2025 and 2026.
  • The steamroller is the tacit β†’ codified β†’ automated β†’ commodity pipeline, and the danger is holding a role whose value is the codified part.
  • Measured: no economy-wide displacement, but a 19% kept-pace shortfall for 22–25-year-olds in AI-exposed occupations, driven by reduced hiring, concentrated where AI substitutes rather than complements, and specifically in occupations involving codified knowledge β€” all labeled descriptive rather than causal by its authors.
  • Getting better at the codified part does not escape it: AI compressed the skill distribution by handing the best workers’ codifiable expertise to novices, with 34% gains at the bottom and minimal gains at the top.
  • The frontier is jagged and invisible, and on the wrong side of it AI users did worse than non-users.
  • Four jobs the model cannot hold: intent, authority, verification, frontier judgment. They appreciate rather than depreciate as capability grows.
  • The honest weaknesses: the broken ladder, the composition problem, the possibility that verification automates, and evidence that is young and descriptive.

Next

Three of those four jobs β€” intent, authority, verification β€” are engineering problems, and the rest of this book builds them. The fourth, frontier judgment, needs a rule before it can be built into anything, because “use AI where it’s good” is not something you can put in a file.

The next chapter turns it into one. It argues that a component belongs on the deterministic side unless it demonstrably cannot be, that doubt about which side is itself the answer, and that the stochastic box is more stochastic than its configuration claims β€” temperature=0 does not make a model deterministic, and the reason has nothing to do with sampling.

Continue with If There’s Any Doubt, It’s Deterministic.

References