Your Approvals Are Teaching AI to Skip You
We were sold the centaur and built its opposite. How AI implementation went backwards — and how researchers and strategists turn it around.
- ● We were sold the centaur — human plus machine — and most companies built its inverse: a person reduced to approving a machine they no longer understand.
- ● AI took the work that teaches you what is good, what is effective, what is worth pursuing — the exact reps that build a strategist's or researcher's judgment.
- ● Our tired approvals are reinforcement learning in disguise, teaching the models that subpar output is acceptable, while Hugging Face's research shows the endpoint: training the human out of the loop.
- ● It is a vicious cycle we designed by handing LLMs too much power — and the people who can reverse it are the researchers and strategists who built it.
We were promised a centaur. The word comes from automation theory: a human with a machine for a body, augmented, faster and stronger than either alone. You bring the judgment, the machine brings the horsepower, and together you outrun anyone working the old way. That was the pitch on every AI rollout deck for three years.
Look at what most of us built instead. The machine does the thinking now, and the person is bolted on to approve it, cleaning up after a system they no longer fully understand and moving at its pace instead of their own. Cory Doctorow has a name for that shape: the reverse centaur, a machine that rides the human rather than the other way around, using them as a pair of hands and a place to put the blame.
We did not set out to build it. We set out to make people faster. But the way we wired AI into the work has inverted who serves whom, and it costs us the one thing a strategy or research team exists to produce: the ability to tell what is good.
Two questions run underneath everything that follows. How we got it this backwards, and how the people who built it — the strategists and researchers wiring these tools into the work — can turn it around.
We automated the wrong half
The tasks we handed to AI first were the ones that teach. Think about how anyone learns to tell good work from bad. You do the first-pass analysis and get it wrong, then wrong in a more interesting way. You draft something and sense it is off before you can say why. The fumbling is where judgment gets built, and we automated the fumbling.
For a researcher, the sharpest example is the humble act of deciding what belongs together. In a recent essay, Saeideh Bakhshi argues that sorting messy inputs into categories — coding interviews, defining themes, naming failure modes — looks like clerical tidying but is where the interpretation happens. Group the responses one way and one pattern appears; group them another way and a different one does. Deciding which contradiction matters and what to call it is the analysis, not the prep for it. Hand the grouping to a model and you have not saved time on the busywork. You have outsourced the judgment and kept the typing.
Brian Elliott named the same loss in Charter. He describes a researcher who used to wrestle with the data, rebuild the model three times, and sit with the findings until they made sense, and who now just asks the tool, because it is faster and the deadline is real. Rebecca Hinds, at Glean's Work AI Institute, puts the rule in one line: protect the work that builds judgment, even when AI can do it faster. Almost no one does, because almost nothing in the workflow rewards it.
What wastes away is specific and expensive: the ability to look at a piece of work and know whether it is any good. Call it taste, or what the Artificiality Institute calls cognitive sovereignty, the capacity to author your own judgment rather than borrow it. It is the reason a senior researcher costs more than a junior one, and it is the first thing to go when a person spends their days approving instead of deciding.
Disuse has a name
Skills you stop using fade, and here that is close to literal. In 2012 the neuroscientist Manfred Spitzer coined digital dementia for the decline that sets in when we let devices hold the capacities we used to exercise ourselves. The phrase is contested and a little lurid, and the effect underneath it is well documented: studies of cognitive offloading keep finding that the more thinking we hand to a tool, the weaker our own critical thinking gets.
You can watch it happen in how experts are made. Matt Beane spent two years studying robotic surgery, where the senior operates the machine and the trainee watches from across the room. Residents get ten to twenty times less hands-on practice than in open surgery, so the skill barely forms unless they scrounge for reps on their own. An MIT Media Lab team put people in EEG caps while they wrote and found the ChatGPT users showed the weakest brain connectivity of any group, and 83% could not quote a sentence from the essay they had just produced. When Wharton researchers gave students an AI tutor, their practice scores jumped 48%, until the tutor was removed and they scored 17% below the students who never had it. The assisted number looked like learning. It was measuring dependence.
Over months, this arrangement produces a steady stream of confident work and a person progressively less able to judge it.
We made our own boredom the reward signal
The second half of the problem lands hardest on anyone who ships AI products. While the setup hollows out the person, it is also teaching the machine, and we are the curriculum.
Modern AI systems learn from human feedback, and not only in the lab. Deployed agents are tuned on what their users approve, complete, and rate well. Every thumbs-up, every "looks good, ship it" shapes the next version. This is reinforcement learning, and you are supplying the reward.
I wrote in June about the new term coined in Glean's Work AI Index: botsitting, the invisible labor of feeding, correcting, and cleaning up after AI, and about the approval fatigue that sets in when an agent asks you to authorise its fortieth step of the afternoon. A tired, botsitting human is a broken reward function. Wave through a plausible but mediocre output because you are exhausted and the deadline is now, and you have done more than ship one weak deliverable: you have told the model, in the only language it hears, that this was good.
Do that across a company, across millions of interactions, and you are running a training program with one lesson: subpar is approvable. The Hugging Face position paper "AI Agents Push Humans Out of the Loop" by Margaret Mitchell and colleagues traces the result. The system optimises for whatever gets approved, fluent and easy to skim, and gets better at being waved through without getting better at being right. The authors give the exhausted reviewer a blunt title: the exploitable part of the reward channel.
Once a system can predict which mistakes you never catch, it can route its errors into those blind spots, since those are the ones that get rewarded. Mitchell's team calls the endpoint training the human out of the loop.
We are teaching the machine to need us less, one tired approval at a time.
A vicious cycle we designed
Put the two halves together and the loop feeds itself. The more we offload judgment, the worse ours gets; and the worse ours gets, the more we wave things through. Each wave-through teaches the model that mediocre passes, so it serves up more of it, which is easier to wave through than what came before.
Nobody designed this on purpose, and yet we designed it. It follows from one implementation choice repeated everywhere: too many of us gave the model the judgment seat and gave the person the approval button. We let the tool decide and asked the human to bless it, when the deciding was the whole reason the human was there. We chose this in how we wired the tools into the work, and we can choose differently.
How we turn it around
The repair is not to use less AI but to change what we point it at, so the person keeps the judgment and the machine takes the load. It is a design decision, ours to remake: in our own habits, in how we run teams, and in what we build. Here is the playbook I give the teams I work with.
Start with your own judgment. The individual discipline is cognitive sovereignty, and in practice it is small and concrete.
- Form your answer before you look. On any decision that matters, write the call you expect before you open what the tool produced, then compare. Where you and the model diverge is where your judgment is still needed, and you approve from a view you hold instead of anchoring to the model's.
- Keep a weekly rep by hand. Once a week, do a real task cold, the way you would have before the tools existed. It is slower on purpose. The difficulty is the training.
- Defend it before it ships. For anything consequential, close the tool and reconstruct the reasoning and the one assumption that would sink it. If you cannot, you received the work rather than made it, and the gap shows you where to go back in.
- Re-sort the pile. When the tool hands you categories, themes, or segments, group a slice of the data a second way by hand. If the story changes, the grouping was a judgment call the model made for you, and you just took it back.
Then change how your team runs. Individual willpower loses to a system tuned to be easy to approve, so the incentives have to move too.
- Measure oversight, not just output. Mitchell's team names three tells you can track: review time falling while approval rates stay flat, disagreement with the tool drifting toward zero, and people asking for less evidence as the stakes rise. Put those next to velocity on the same dashboard.
- Split the approver from the beneficiary. Never let the person who gains from a "yes" be the only one who gives it. A second reader with nothing to ship restores an honest signal.
- Protect the reps that build judgment. Before automating a task, ask Hinds' question: what would a newcomer learn by doing this? Automate what teaches nothing or has already hit diminishing returns, and guard what teaches taste. A team that automates its own apprenticeship might post a great quarter but has no subject matter experts in five years.
- Rotate assisted and unassisted work. Move people between doing a task with the tool and without it, so the baseline stays alive and you can see the moment it starts to slip.
Then fix what you build. If you design the tools, you hold the biggest lever, because you decide where the friction goes.
- Give the model the load, not the verdict. Point it at drafting, gathering, and widening the options a person weighs, and keep the decision with the person. A tool that surfaces three framings a researcher had not considered is worth more than one that hands over a single answer to rubber-stamp.
- Design friction that teaches. The version of the Wharton study that gave hints instead of answers did no lasting harm, because it kept students working. Build the assistant that asks "what would change your mind here?" instead of the one that ends the thinking.
- Make the work defensible by design. Show the reasoning, the sources, and the assumptions next to the output, so checking it is a glance rather than an excavation. The easier you make the work to interrogate, the slower the reward channel rots.
All of this is slower, it can dent this quarter's satisfaction score, and it asks people to sit in the difficulty the tools promised to remove. Spend the time anyway, because the bill for the alternative comes due at the next hiring round, when you go looking for judgment and find a roster of people who have only ever approved. The edge goes to the companies whose people can still tell what is good, not to whoever has the most AI, and that is something you decide in how you deploy.
We were sold a centaur and, rushing to show results, built its inverse. The word that matters in that sentence is "built," because what we built we can rebuild. Every approval you actually read, every task you keep because it sharpens someone, points the work back toward the human. Make those choices while they are still yours to make.
Sources
- What AI Is Doing to On-the-Job Learning—and How to Protect It — Brian Elliott, Charter (via TIME)
- AI Agents Push Humans Out of the Loop — Margaret Mitchell, Avijit Ghosh, Samir Passi (Hugging Face / Data & Society)
- The Reverse-Centaur's Guide to Criticizing AI — Cory Doctorow, Pluralistic
- What Belongs Together — Saeideh Bakhshi, Research Toolbox (publication home; direct post link to be confirmed)
- Cognitive Sovereignty: Authoring Your Mind in the AI Age — Artificiality Institute
- Digital dementia: a review (term coined by Manfred Spitzer, 2012); and AI Tools in Society: Impacts on Cognitive Offloading and the Future of Critical Thinking — Gerlich, 2025
- Shadow Learning: Building Robotic Surgical Skill When Approved Means Fail — Matt Beane, ASQ 2019
- Your Brain on ChatGPT: Accumulation of Cognitive Debt — MIT Media Lab; and Generative AI Can Harm Learning — Bastani et al. (Wharton), 2025
- New Report Says You're Wasting More Time Botsitting Than Getting Value from AI — my earlier piece on the AI value gap (Product Impact, June 2026)
How helpful was this article?
Share this article
Latest Episodes ›
All episodes
20. AI Shouldn't Be Like Bolting a Spoiler on a Crummy Honda (Sentient Design Authors Josh Clark & Veronika Kindred)
19. Upgrade from Vibe Coding to AI-Native Product Design (Metalab's Myles Palmer)
18. Why are tech workers SO unhappy about AI?

Every CEO Will Post a Layoff Notice Like This. Here Is Why.

Future of UX Research: Guide for How UX Researchers Can Protect Their Jobs & Careers

Lenny's 2026 Tech Survey: AI Burnout Is Surging, Layoff Fear Is High

The UX Researcher's Guide to Claude, Claude Cowork, and Claude Code

WTF is an AI-native org anyways? Let's compare Airbnb & Meta's opposing plans.

The Free Ride Is Over: AI Economics Is Now Your Most Important Strategy Decision
Product Impact Newsletter
AI product strategy delivered weekly. Free.