Meat Proxy
Repeated apparent success can turn human review into a rubber stamp. Decades of automation research show that simple practice and instructions do not eliminate complacency or automation bias. Review therefore has to be engineered and evidenced rather than assumed.
Part 1 — Where You Stand
Review degrades because the system succeeds
Ten weeks
Week one, the model drafts a paragraph review and you read every line of it. You disagree with two of its four flags, check a source yourself, rewrite one sentence. The whole thing takes twenty minutes and you feel you have earned the output.
Week three, the reviews have been good. You read them properly but you stop re-checking the sources it says it checked. Nothing bad happens.
Week six, you skim. You read the first flag carefully and glance at the rest. You have a feel for what a wrong one looks like now — or you believe you do. Nothing bad happens.
Week ten, you approve.
Not because you got lazy. After a long run of acceptable reviews, checking the next one can begin to feel like a poor use of twenty minutes. Each decision to look a little less closely can feel reasonable in isolation.
This opening is a constructed scenario, not a measured ten-week decay curve. The mechanism it illustrates is the important part:
Repeated success can make review thinner precisely when review still matters.
A system that fails often keeps the need for scrutiny visible. A system that appears highly reliable can make close checking feel less valuable, which creates the conditions under which a rare failure may pass through signed but effectively unexamined.
Call the resulting role what it is. You are a meat proxy: a biological component whose function is to convert machine output into approved output, adding a signature and no information.
What has to be true of a process for a human review to still mean something after the hundredth time?
It has a name, and thirty years of evidence
This is not a new phenomenon and it is not specific to language models. Human factors research has studied it since long before any of this, mostly in aviation and process control, under two names.
Automation bias is the tendency to over-rely on automated recommendations. It produces two distinct error types: omission errors, where you fail to notice a problem because the automation did not flag it, and commission errors, where you act on an automated recommendation in spite of contradictory evidence available to you.
Automation complacency is the attentional counterpart: reduced monitoring of an automated system that has been performing well.
Parasuraman and Manzey’s review of this literature reaches three conclusions that ought to change how you plan your week (Parasuraman & Manzey, 2010):
- Complacency and automation bias are manifestations of overlapping, automation-induced phenomena in which attention plays the central role.
- Both appear in naive and expert participants in the studies they reviewed — expertise does not by itself protect the monitoring task.
- Automation complacency cannot be overcome with simple practice, and automation bias cannot simply be trained or instructed away.
That third point forecloses the easiest response. The literature does not show that careful people are helpless; it shows that individual resolve alone is not a reliable control. If your entire review policy is I will be diligent, the policy is missing an engineering mechanism.
The multiple-task result matters particularly for AI-assisted work. Parasuraman and Manzey report complacency under multiple-task load, when manual work competes with monitoring for attention. A system that increases the amount of output one person is expected to oversee can create exactly that competition for attention.
Related evidence in knowledge work
A survey of knowledge workers shows a related pattern, with an important evidential boundary.
Lee and colleagues surveyed 319 knowledge workers, collecting 936 first-hand examples of real tasks. Higher confidence in the AI was associated with less critical thinking, while higher task-specific self-confidence was associated with more (Lee et al., 2025).
Those are associations from self-reported work examples, not a demonstration that AI confidence caused reduced critical thinking or that repeated success produced either confidence state. The useful distinction is still clear: confidence in the tool and confidence in your own ability to judge the task are different variables.
Qualitatively, the same study found critical-thinking activity shifting toward information verification, response integration, and task stewardship. That echoes the responsibilities Chapter 2 keeps outside the model without establishing that the transition is automatically being done well.
And recall the bounded experimental result from Chapter 2: on one managerial task deliberately selected to sit outside the model’s capability frontier, consultants using AI were 19% less likely to produce a correct solution than those with no AI at all (Dell’Acqua et al., 2023). That result shows the cost of misplaced reliance in that task. It does not establish a general meat-proxy error rate, but it is enough to show why a human signature cannot be treated as evidence that the output was independently checked.
Why the process has to supply the missing part
A generator cannot be the sole source of both the work and the criteria by which that work counts as good.
A model can respond to objectives, constraints, examples, and criteria represented in its training, system instructions, prompt, or context. What it cannot do for your application is independently own the objective against which its own proposal will be judged. Chapter 2 therefore keeps intent outside the model: the process has to state what is wanted and what would count as done, even if a model helps propose or refine those criteria.
That gives this book a useful definition of slop:
Slop is not merely bad output. Slop is output that is locally plausible and globally unaccountable — nothing in it is wrong enough to catch, and nothing in it was chosen for an explicit reason.
The problem is not that every sentence must contain an error. A piece can be individually defensible sentence by sentence while contributing little to an explicit objective.
So do not look only for mistakes. Ask what the piece claims, why each important part is there, what evidence supports it, and what would make the claim fail. Clean prose without an accountable objective can still be the wrong artifact.
Why the failure is catastrophic rather than gradual
The user-facing danger of the meat proxy is not that a bad output ships. Bad outputs ship all the time and get caught.
There are two compounding mechanisms, and the second is the one that hurts.
The first is deskilling — but it is differentiated, not uniform. If you stop doing the work, you can lose the ability to evaluate the work. The review that was cursory because it was unnecessary can become cursory because it is no longer possible. Chapter 2 warned that the path to senior judgment historically ran through doing the junior work. This is the same mechanism, operating on someone who already has the judgment and may quietly spend it down.
But a 2026 field experiment should stop us stating that as a law. Autor and colleagues ran a pre-registered three-month randomized trial with 133 practicing patent lawyers at eleven US firms, using a custom AI drafting assistant, with all work scored by blinded expert attorneys. AI access raised drafting quality at 10 days (+0.34 SD) and 90 days (+0.38 SD), with larger immediate gains for juniors. After three months, everyone redlined an application without AI: treated lawyers outperformed controls by +0.32 SD, but the durable gain was concentrated entirely among seniors (+0.45 SD). Juniors showed no average gain — their scores bifurcated instead, with fewer mediocre outcomes offset by more poor and more good ones (Autor et al., NBER w35720, 2026).
Read carefully, with the authors’ limits: one occupation, Google-funded, a working paper rather than peer-reviewed publication. Measured: AI-assisted practice built lasting judgment for seniors (+0.45 SD on the unaided redline) while juniors bifurcated with no average gain. The authors’ reading: foundational expertise may be a prerequisite for extracting durable skill from AI-assisted practice. The book’s extension, not theirs: if so, the one-way door is realest where expertise was thinnest to begin with. That refines rather than removes the warning.
Field observation of developers points the same way. Experienced developers using coding agents retain architecture and design control, supply context, and repeatedly verify rather than delegate — control as the observed expert pattern, from field observations (N=13) and surveys (N=99) (Huang et al., 2025).
Across roughly 400,000 Claude Code sessions, Anthropic finds people disproportionately make planning decisions (what, ~70%) while the model makes execution decisions (how, ~80%), with verified success rising from ~15% for novice-rated sessions to ~28–33% for intermediate and above — vendor observational data, classifier-judged, not causal, but consistent: the human owns intent, and expertise amplifies what the agent can do (Anthropic, 2026).
The second is the missing trail. When the failure is finally detected — externally, by a customer, an auditor, a regulator, or a court — the question is never “is this one wrong?” The question is how far back does this go?
And you cannot answer it. Your record is a sequence of approvals. The log shows a timestamp and a signature on each one. Open any entry and you find the edited file, but not the bytes the model returned, not the checklist of what was verified, only the source link pointing at whatever version is current now. An auditor sampling across ten thousand such entries finds the same thin record repeated, and has to treat the ones you read and the ones you glanced at as the same thing.
That is the catastrophic failure the meat proxy is walking toward. An unbounded blast radius, because the process that produced ten thousand results retained no evidence about any of them.
This is why the runtime chapters of this book are not administrative overhead. A record of what was checked and accepted (Chapters 14 and 16), artifacts that preserve what was actually returned (Chapters 11 and 17), and claims carrying their evidence (Chapter 18) are precisely the machinery that converts “everything is suspect” into “these four hundred are suspect, and here is why.”
The loop that produces the proxy, and where it exits into catastrophe:
flowchart TD
S["repeated success<br/><i>forty good outputs in a row</i>"] --> T["trust rises<br/><i>each step locally rational</i>"]
T --> C["checking thins<br/><i>skim, then approve</i>"]
C --> U["failure passes unexamined<br/><i>signed by you</i>"]
U -.->|"undetected failure<br/>looks like more success"| S
U --> E["approvals lose evidential value<br/><i>no record of what was checked</i>"]
E --> B["unbounded blast radius<br/><i>how far back does this go?</i>"]
The institutional trap
The meat proxy is not merely tolerated but frequently rewarded, and it is worth being clear-eyed about why.
Ben Green surveyed 41 policies requiring oversight of government algorithms and identified two flaws. First, the evidence he reviewed indicates that people may be unable to perform the oversight functions those policies assume. Second, he argues that weak human-oversight requirements can legitimize faulty and controversial algorithms without addressing their underlying problems, provide a false sense of security, and allow vendors and agencies to shirk accountability (Green, 2022).
Green studied government policy, not commercial AI review, so the transfer from his result to this chapter is an inference. The transferable warning is structural: requiring a human somewhere in the process does not establish that meaningful oversight occurred.
The same failure mode can exist commercially when a signature satisfies the procedure while the underlying scrutiny remains thin. Whether a particular organization prefers signature to scrutiny is an empirical question this chapter has not measured. The architectural response is to make the review itself inspectable — what was checked, against which state, with what evidence — instead of treating the presence of a person as the control.
If you are going to do it anyway
Some readers will do this regardless, and some tasks genuinely do not justify a careful human pass. So: if your review is going to be shallow, here is how to make it as strong as a shallow review can be. Offered without sarcasm — a well-built shallow review beats a badly-built one by a large margin.
If you use model judges, prefer a panel to a single judge where the evidence supports it. Verga and colleagues evaluated a Panel of LLM Evaluators — smaller models drawn from disjoint model families — against a single large judge across three judge settings and six datasets. The panel outperformed the single large judge, exhibited less intra-model bias, and cost over seven times less (Verga et al., 2024). That makes a panel a useful evaluation mechanism in those settings. It does not make the panel independent evidence that an output is correct.
Do not treat generator self-evaluation as an independent check. Panickssery and colleagues found a linear relationship between a model’s ability to recognize its own output and the strength of its self-preference (Panickssery et al., 2024). A self-grade may still be useful as another model-produced signal; it is not the external evidence Chapter 21 will require.
Prefer mechanical checks where the property is mechanical. Does every cited source resolve? Does every quoted sentence appear verbatim in the source? Do the numbers sum? Does it compile? Are all flagged spans real offsets into the document? These questions are repeatable and exact about the property they test. They are also silent about everything they do not test, which is why Chapter 21 later shows a well-built deterministic checker accepting a grounded, verbatim-quoted, wrong answer.
Sample deliberately instead of skimming indiscriminately. A genuinely random reviewed sample can produce an observed error rate with uncertainty; an unstructured skim cannot. There is no universal useful percentage: the number of items reviewed determines how much precision the estimate has, and the sampling policy should follow the consequence of missing a defect.
Record the shallow review anyway. Record the version, checks run, sampled items, results, and approver. That record can help a later investigation distinguish what was actually checked from what merely received a signature.
And the honest limit on all of it: these mechanisms can improve filtering, measurement, and auditability. They do not establish correctness or transfer accountability. A model panel is still a model-produced judgment. Where the remaining decision requires human expertise, the process still needs a person who understands the work.
The real path: review is architecture, not willpower
Now the version for work you actually care about.
The automation literature says simple practice, training, and instructions do not reliably eliminate complacency and automation bias. So review cannot rest on virtue alone. Treat it as a design problem: what would have to be true of this process for careful review to be cheap, hard to bypass accidentally, and measurable?
Six mechanisms, each of which the rest of this book builds:
Make the model emit reviewable artifacts, not unsupported conclusions. A claim tied to a source span and quoted text gives the process several mechanical things to check: that the source exists, that the quote occurs where claimed, and that the attribution is intact. Whether the evidence actually supports the claim may still require judgment. Chapter 18 separates those layers rather than pretending one replaces the other.
Put adequate deterministic checks in front of human attention. Do not spend judgment on arithmetic, schema validity, source existence, or exact-span matching when software can check those properties first. Human attention should arrive at the residual questions the mechanical checks could not settle, with the limits of those checks visible (Chapters 3, 14 and 21).
Route attention by risk, not by volume. Do not review every effect as though it had the same consequence. Irreversible, novel, and high-consequence actions deserve different authority and review policy from cheap reversible ones. Chapter 20 makes permission explicit so the process can route attention before the effect occurs.
Preserve the raw output. Keep the bytes the model actually returned before later edits or interpretations replace them (Chapters 11 and 17). Preservation alone does not bound an incident, but it makes later re-evaluation possible when the task, state, checks, and decisions are also retained.
Make the review record an artifact with evidence attached. An approval should name what was checked, against which state, and what the checks returned. A signature without that record cannot later tell you what the reviewer actually established (Chapters 14 and 18).
Test the review process itself. Periodically seed known defects that the review is supposed to detect and record whether they are caught. Treat the result as evidence about those seeded defect types, not as a complete measure of review quality. From Chapter 15 onward, the book applies the same principle to its automated evidence verifiers: deliberately corrupted copies must be rejected, and a corruption that survives exposes a gap in that verifier.
Notice what these mechanisms have in common. None depends on asking the reviewer to concentrate harder forever. Each changes the structure of the work so that useful evidence reaches human attention before the approval decision.
The heart at the end of the machine
The model generates. The runtime coordinates. The tools act. The verifiers check what can be checked mechanically.
What is left for the human is not one generic approval at the end. It is the responsibilities the book has already separated: intent, authority, frontier judgment, and judgment wherever no adequate mechanical verifier can establish the property that matters. The model can help with each of them; it cannot be allowed to turn its own confidence into evidence that they were discharged.
That is the difference this chapter cares about between plausible output and applied AI. Not a heroic reviewer compensating for a weak process, but a process that preserves meaningful human judgment after week ten instead of merely preserving the signature.
Do this now
Thirty minutes. Run a seeded-defect smoke test on your review process.
- Take ten recent AI outputs you approved. Real ones.
- Give copies to someone else — or to a script — and have them introduce a small set of defects your normal review is supposed to catch: a wrong cited year, a quotation absent from its source, a total that no longer sums. Keep the locations and, if practical, the number of inserted defects hidden from the reviewer.
- Review the batch the way you normally would, at your normal speed.
- Reveal the answer key and record which seeded defects you caught and missed, including the defect type.
Write the result down with the date. One ten-item batch is a diagnostic, not a stable estimate of your general review accuracy. Repeat the same kind of test over time if you want to detect whether the review process is changing. The important move is that review quality has become observable rather than assumed.
Failure modes
- Assuming diligence is a plan. Complacency appears in experts and does not yield to practice.
- Reading success as safety. The better the system performs, the faster you stop checking, and the more your attention is needed for exactly the cases you are no longer seeing.
- Letting the generator judge itself. Self-preference scales with self-recognition; the bias points the wrong way.
- A single judge model. A panel of diverse smaller models beat it, with less intra-model bias, at a seventh of the cost.
- Asking “is this good” instead of “does this check out.” Judgment where a mechanical check was available.
- Skimming everything instead of sampling deeply. Produces no error estimate, so degradation is invisible until it is external.
- Approving without evidence. Converts a recoverable incident into an unbounded audit.
- Confusing compliance with oversight. A required signature is not a performed review, and the policy may exist to place blame rather than to catch errors.
What this chapter established
- The meat-proxy failure mode can emerge after repeated apparent success: close review starts to feel less valuable even though the remaining failures still require it. The ten-week opening is a constructed mechanism, not a measured decay curve.
- Automation bias includes omission and commission errors; automation complacency is tied to attention under multiple-task load. Both appear in naive and expert participants, and simple practice, training, or instructions do not reliably remove them.
- In the knowledge-worker survey, higher confidence in AI was associated with less critical thinking while higher task-specific self-confidence was associated with more. In one outside-frontier BCG task, AI-assisted participants were less likely to reach the correct answer than the no-AI group. Neither result establishes a universal review-failure rate.
- The application has to supply and own the objective and success criteria against which model output is judged. For this book, slop is locally plausible output that remains globally unaccountable to such an objective.
- A thin approval trail creates a large audit problem because later investigators cannot distinguish what was substantively checked from what merely received a signature. Preserved observations, criteria, state, checks, and decisions are what make that history reconstructable.
- The patent-law experiment complicates a simple deskilling story: seniors showed durable unaided gains while juniors bifurcated with no average gain. The authors suggest foundational expertise may matter; the book treats greater risk where expertise is thin as an inference, not a demonstrated one-way door.
- Green’s study shows that human-oversight policies can legitimize weak government algorithmic systems without establishing effective oversight. The commercial meat-proxy analogy is this book’s extension, not a measured prevalence claim.
- Model panels can be useful filters and can outperform a single judge in the evaluated settings, but another model’s agreement is still not independent evidence. Mechanical checks should establish the properties they actually cover, and random sampling should be treated as measurement with uncertainty rather than ritual inspection.
- Serious review uses reviewable artifacts, adequate deterministic checks ahead of attention, risk-based routing, preserved raw observations, evidence-bearing review records, and seeded defects that test whether the review process detects failures it claims to detect.
- Review is architecture, not willpower. Verification is one of the human responsibilities this book keeps separate from generation, alongside intent, authority, and frontier judgment.
Next
One question remains before we build, and it is the one that decides how much of any of this you can afford to do.
Every mechanism in this chapter costs something. Deeper review costs attention. Panels cost tokens. Preserved artifacts cost storage. And the thing being reviewed costs money on every single call, forever — which is a strange property for software to have, because software used to get finished. The next chapter works out what you are actually buying when you buy intelligence, why a verifier puts a ceiling on that price and its absence removes one, and how to take the cost back out.
Continue with The Price of Intelligence.
References
- Raja Parasuraman and Dietrich H. Manzey. Complacency and Bias in Human Use of Automation: An Attentional Integration. Human Factors, vol. 52, no. 3 (2010), pp. 381–410. https://doi.org/10.1177/0018720810376055
- Hao-Ping (Hank) Lee, Advait Sarkar, Lev Tankelevitch, Ian Drosos, Sean Rintel, Richard Banks, and Nicholas Wilson. The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers. Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ‘25). https://doi.org/10.1145/3706598.3713778
- Ben Green. The Flaws of Policies Requiring Human Oversight of Government Algorithms. Computer Law & Security Review, vol. 45 (2022), 105681. https://doi.org/10.1016/j.clsr.2022.105681
- Pat Verga, Sebastian Hofstatter, Sophia Althammer, Yixuan Su, Aleksandra Piktus, Arkady Arkhangorodsky, Minjie Xu, Naomi White, and Patrick Lewis. Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models. arXiv:2404.18796, 2024. https://arxiv.org/abs/2404.18796
- Arjun Panickssery, Samuel R. Bowman, and Shi Feng. LLM Evaluators Recognize and Favor Their Own Generations. Advances in Neural Information Processing Systems 37 (NeurIPS), 2024. https://arxiv.org/abs/2404.13076
- Fabrizio Dell’Acqua et al. Navigating the Jagged Technological Frontier. Harvard Business School Working Paper 24-013, 2023. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4573321
- David Autor, Tanya Rodchenko, Josh Martin, Zanna Iscenko, Scott Strand, David Pearl, and Melissa Ferere. Does AI Assistance Enhance or Erode Expertise? Evidence from a Three-Month Field Experiment in Patent Drafting. NBER Working Paper 35720, September 2026. https://www.nber.org/papers/w35720 — Pre-registered RCT, 133 patent lawyers; +0.34/0.38 SD assisted quality, +0.32 SD unaided redline concentrated in seniors (+0.45 SD), juniors bifurcated. Working paper, single occupation, Google-funded.
- Ruanqianqian Huang, Avery Reyna, Sorin Lerner, Haijun Xia, and Brian Hempel. Professional Software Developers Don’t Vibe, They Control. arXiv:2512.14012, 2025 (v2 Aug 2026). https://arxiv.org/abs/2512.14012 — Field observations N=13, survey N=99; experts retain design control, supply context, verify.
- Anthropic. Agentic coding and persistent returns to expertise. June 2026. https://www.anthropic.com/research/claude-code-expertise — ~400,000 Claude Code sessions; user owns ~70% planning, model ~80% execution; verified success 15% novice vs 28–33% intermediate+. Vendor observational data, classifier-judged success.