The queue that feels like work

It's Monday morning. Fourteen things are waiting for your review, and eleven of them were written by a machine. A strategy memo. Three user flows. A data cut with the analysis already drafted. A customer email that is, honestly, better than the one you would have written. Your job is to read each one, catch what's wrong, and approve. You're fast at it. By eleven the queue is clear. It felt like a productive morning — you made a lot of calls.

Here's what should bother you. You can't quite say why you approved the ones you approved. They looked right. Nothing snagged. And the two junior people on your team — the ones who used to bring you rough drafts full of the specific, instructive mistakes that told you exactly where their thinking broke — are doing the same thing you are now. Reading machine output. Nodding it through. Getting faster at it.

I spent years in rooms like this before anyone typed a prompt into a chatbot. Put a dozen capable people around a plan none of them wrote, ask if anyone has concerns, and watch how fast a room signs off on work it couldn't have produced. The nods are genuine. The scrutiny is mostly theater — not because the people are lazy, but because you can't really interrogate an answer you have no independent version of. You can only check it for plausibility. AI didn't invent that gap between approving and knowing. It industrialized it, and put it on every desk.

Reviewing spends judgment; generating builds it

Start with what your judgment actually is, mechanically. It runs on two moves that used to come welded together.

One is generating — producing a candidate answer out of your own head, committing to it, and living with the gap between what you intended and what happened. The other is recognizing — looking at an answer and knowing, fast, whether it's right.

That second move is not trivial, and I don't want to wave it away. Gary Klein spent decades showing that expert judgment mostly is recognition: the fireground commander who pulls his crew before he can say why, the NICU nurse who flags a baby going septic before the labs confirm it. Experts don't run option trees. They recognize. So "just reviewing" can absolutely be real judgment.

But here's the part that matters. That expert recognition was built. The commander's instant read was paid for on a thousand real fires. Recognition isn't the opposite of doing — it's the residue of it. You earn the eye by generating answers and living with how they turn out, again and again, until the pattern-library is deep enough to fire on sight.

Which is why stripping the generating out is not a neutral trade. When you review an AI's answer, you spend that hard-won recognition and replenish none of it. In the memory lab they call the raw version the generation effect: people remember what they generated far better than what they merely read (Slamecka and Graf, 1978). The same asymmetry seems to run through judgment, one level up. The answers you produce and own build the eye. The answers you wave through don't.

Reviewing spends judgment; generating builds it.

The two used to travel together. You couldn't sign off on a colleague's analysis without, somewhere in your head, running your own version to compare against — the generating rode along inside the reviewing, for free.

That's the answer to a fair objection: plenty of experts build judgment by evaluating — the radiologist reading films she didn't shoot, the editor who never wrote the novel. They do. But watch what a good one actually does. She forms her own read and commits to it before she rules, then compares. She's generating; the review is just where it lands. Evaluation sharpens you exactly when it smuggles generation inside it, and builds nothing when it doesn't. Approving a finished, confident answer you never tried to produce yourself is the second kind.

And that is the kind AI serves up. It breaks the bundle: the answer arrives fully formed, confident, well-formatted, and you never have to build your own to check it against. Under time pressure, you don't. You just recognize.

Sometimes checking for plausibility is enough — I'll get to when. But on the calls that carry real weight, it isn't, and the failure is invisible from the inside. You're still making calls all day. It still feels like judgment. You'll feel exactly as sharp as last year — right up until the case that needed the built version. The answer that's confident, plausible, and wrong. The one only generating your own would have caught. You wave it through, because waving through is what you've been practicing.

This isn't new — AI just industrialized it

None of the diagnosis is my discovery, and the first move of an honest argument is to say so.

In 1983 the human-factors researcher Lisanne Bainbridge wrote a short paper, "Ironies of Automation," that has been quietly correct for forty years. Automate the easy parts of a job and leave the person to monitor, she showed, and you set a trap: "By taking away the easy parts of his task, automation can make the difficult parts of the human operator's task more difficult." You're kept on to catch the machine in the rare moment it fails — using skills that are decaying precisely because the machine now does the thing that used to build them. "The human monitor," she wrote, "has been given an impossible task." She even saw the generational version: later operators "cannot be expected to have" the skills the first generation built by doing. Nicholas Carr carried this into professional work in his 2014 book The Glass Cage — the doctor leaning on decision support, the analyst trusting the model. If you want the fuller intellectual history, it's theirs, and it's worth your time.

What's different now is not the mechanism. It's the blast radius. Bainbridge's irony lived in a few control rooms; the reviewing-instead-of-doing pattern is being installed across every desk in the knowledge economy at once, in the space of about two years, and sold as a productivity strategy.

And that it degrades the expert is not speculation.

60–70%
how often AI-assisted consultants reached the right answer on a task past the tool's competence — against about 85% for those working without it. Dell'Acqua et al., Harvard/BCG, 2023 (N=758).

That Harvard–BCG experiment is the one to sit with. Its 758 consultants were meaningfully better with AI on tasks the tool was suited to — and worse on one it wasn't. On a problem built to fall just outside the tool's competence, where its confident answer was wrong, the AI-assisted group reached the right answer 60 to 70 percent of the time, against about 85 percent for the people working without it. They didn't catch the confident wrong answer. They recognized it as plausible and moved on.

It shows up wherever the study design lets you see it. Endoscopists who had been working with AI polyp-detection got measurably worse at the job without it — their unassisted detection rate fell from about 28 percent to 22 in a 2025 study in The Lancet Gastroenterology & Hepatology. Two decades earlier, the same shape appeared in radiology: computer aid lifted the weaker mammogram readers and dragged down the best ones on the hardest cancers (Povyakalo and colleagues, 2013). Airline pilots keep their motor skills under automation but lose the cognitive thread — where they are, what comes next — when they stop actively flying (Casner et al., 2014). And the obvious fix, just train people to stay vigilant, doesn't take: a 2010 review of decades of studies concluded automation complacency, in experts as much as novices, "cannot be overcome with simple practice."

When approving is fine

Here's where arguments like this one overreach, so let me draw the line honestly. Reviewing instead of doing is not the enemy. Most of the time it's exactly right, and there are two clean cases where you should stop worrying.

The first is when there's a cheap, trustworthy way to check the answer that doesn't lean on your own judgment — what engineers call an oracle. A compiler is an expert approving a machine's output all day, and no one loses sleep, because the check is mechanical: the code runs or it doesn't. Unit tests, a reconciliation, a lookup against a source of truth — where an oracle like that exists, let the machine produce and let the human approve. You're not spending your judgment there; the check is. It simply moves up to the parts no oracle covers. The same logic covers low-stakes, easily reversed work: if a wrong answer is cheap to catch and cheap to undo, you don't need your best judgment on every instance of it.

The second is when the oversight is actually designed. A surgical checklist strips away the surgeon's discretion at the moment of action and forces them to verify instead of assume; across eight hospitals it cut complications from 11 percent to 7 and deaths from 1.5 to 0.8 (Haynes et al., 2009). That can look like the opposite of my argument — recognition beating generation — but it's the reverse, and there's a simple test that tells them apart: does the design force the human to do something, or just to nod? The checklist forces an active verification the surgeon would otherwise skip. It isn't a person passively approving. It's a designed act that makes the human do the one thing that catches the error.

So the line isn't "automation bad." Two different things get blurred: who produces the work, and whether the check is real. Automating the production is fine when the check is real — an oracle, or a designed forcing-function. It goes wrong in one specific place: high-stakes work, no cheap oracle, a usually-right machine that fails rarely and plausibly, and a human in a chair whose only move is to approve. That's the cell where the person is the last line of defense, and the arrangement is quietly disarming them.

Judgment debt

Now move it up from the person to the organization, because that's where it gets expensive.

Most companies are, right now, describing their human experts as their edge — the judgment, the taste, the discernment the machine doesn't have. And in the same quarter they're reorganizing those same experts' work so that it mostly consists of approving machine output. They are spending the exact asset they're advertising. Because the spend shows up nowhere — no line item, no visible loss — they're spending it blind.

Call it judgment debt. The name borrows on purpose from technical debt and organizational debt: a cost you take on now, invisibly, that comes due later with interest. (There's an individual cousin — in 2025 an MIT team measured weaker brain connectivity in people who wrote essays with an AI — and found most of them couldn’t quote a single line of what they’d just written — and called it "cognitive debt." Judgment debt is the organizational version: not one brain offloading, but a whole firm quietly consuming the capability it sells.) Every time an expert approves instead of generating, the organization books a small productivity gain and an unrecorded liability — a little less judgment in the building than there was. It compounds silently, and it comes due on the worst possible day: the high-stakes, non-obvious call where the real thing was needed and an approval reflex answered instead.

The debt has two sides, and they pull in opposite directions — which is what makes it genuinely hard. Your veterans are drawing down: they built the judgment and are now spending it without replenishing. Your juniors are worse off, because they never build it at all — the entry-level production work that used to be their apprenticeship is the first thing the machine took. One group needs pushing back toward doing; the other needs pushing into it for the first time. Both draw on the same scarce thing — real, un-automated work with consequences attached — and there's less of it every quarter. That shrinking pool is now something you have to ration on purpose, between keeping your best people sharp and making your next ones. Almost no one is doing it on purpose. It's just evaporating.

Oversight you can't fake

If the problem is that approving got stripped of generating, the fix is to design the generating back in — not everywhere, which would throw away the speed that made AI worth adopting, but on the calls that carry weight.

The core move is to generate before you reveal. On a decision that matters, form your own answer before you look at the machine's — even a rough one, even just its shape — and then compare. The comparison is where both the real review and the retained skill come from. Read the AI's version first and anchoring is instant and invisible; you'll find its answer reasonable because it's already in your head. Which is why this can't be left to willpower. It has to be built into how the work is set up — the machine's output withheld until you've committed your own — because a discipline you can fake, you will. And you don't need it everywhere: start with the decisions that have the biggest blast radius, the ones that are expensive to get wrong, and require the independent answer there.

The second move is to spend attention by consequence, not by the calendar. A five-percent spot-check catches nothing that matters, because the errors that matter are rare and confident. Put the deep, regenerate-it-yourself review on the decisions with real downside, and let the reversible, low-stakes work flow through.

The third move is the hardest, and it's about the juniors. The apprenticeship used to be a free byproduct of the work; it isn't anymore, so you have to rebuild it deliberately — give junior people real decisions, with consequences, on purpose, even when the machine could do it faster. And here the honesty has to be total. This costs the exact speed the AI was bought to deliver; a company under quarterly pressure will resist it; and the one company that does it while its competitors don't ends up training talent for the whole market. Some of this probably can't be solved inside a single firm at all — it may need the equivalent of a residency, a professional norm nobody gets to skip. I don't have that part solved. I'm confident the first step is to stop pretending the apprenticeship still happens on its own.

Underneath all three is a test you can run on any "human in the loop" the moment it's offered as a safeguard: can that person form an independent judgment, do they have the time to, and do they have the authority to act on it? Miss any one and you don't have oversight. You have someone positioned to take the blame.

Key takeaways
  • Generate before you reveal: on decisions that matter, commit your own answer before you see the machine's, then compare. Build it into the workflow — a discipline you can fake, you will.
  • Spend oversight by consequence: deep, regenerate-it-yourself review where being wrong is costly and hard to reverse; let the rest flow.
  • Rebuild the apprenticeship on purpose: give juniors real, consequential decisions even when the machine is faster — the eye no longer forms for free.
  • Test any "human in the loop": independent judgment + time + authority to act. Miss one and it's decoration.

The move that keeps you sharp is the move that keeps you

There's a quieter fear under all of this, and it deserves a straight answer. If you get good at encoding your judgment into the system — the rubric, the guidelines, the examples the model learns from — aren't you just training your replacement, faster?

Partly, yes. The part of your expertise you can write down as a rule is the part that's going to be automated, and guarding it won't save it. But that was never the valuable part. The valuable part is the judgment that writes the next rule when the situation shifts — the capacity to notice the rubric no longer fits and make a better one. A snapshot can't hold that. And it stays alive only the way it was built: by continuing to generate.

Which means the move that keeps you sharp and the move that keeps you employed are the same move. Not approving faster. Not encoding yourself more completely into the tool. Staying in the generation — using the machine for the production it's genuinely good at, and deliberately keeping the part where you form the answer yourself, because that part is both the edge and the exercise that maintains the edge.

The people who come out of this era with their judgment intact won't be the ones who reviewed the most AI output. They'll be the ones who kept answering the question first, and built their work so they had to.

So here's the small, specific question, the kind you can check tomorrow. Of the calls you approved today, how many did you generate your own answer to before you looked? That number, not the size of your queue, is how much judging you actually did. And the harder version, for whoever runs the place: of the real work that's left, who decided how much goes to keeping your best people sharp, and how much to making the next ones? For most organizations right now, the honest answer is no one. It's just being spent.