AI Transformation
Why Human Judgment Still Matters in AI Work
By Robert Antolin · · 9 min read
AI made producing an answer cheap. Checking one did not get cheaper, so that is where the value moved. Give a team a printer that runs ten times faster and the proofreader is still the bottleneck. A team that can produce drafts far faster than it can check them has not gained capacity, it has gained a backlog.
People stay in the loop because three jobs still sit with them: framing the work (deciding what the AI should do), verifying the output (checking it against something known), and owning the result (being the named person who answers for it). Each job fails in a way researchers have measured, and each failure looks like a productivity gain until someone checks.
Why does human in the loop AI matter at work?
Because those three jobs decide whether AI output is worth anything, and the tool is least reliable exactly when it sounds most certain. Human in the loop is not a person watching a screen. It is a person holding three jobs:
- Framing. Deciding what the AI should do, and noticing when it has gone off course. This is a judgment about the business problem, not about the model.
- Verifying. Comparing the output to something known before it takes effect. That something, the reference, has to exist before the run. Otherwise the check is just a second opinion from the same source.
- Owning. One named person accountable for each output that leaves the building. Not the team, not the tool.
The second job has its own argument, worth reading separately: the gate that works is verification, not approval. Approval is a person looking at a summary and clicking yes. Verification is a person comparing the output against a reference chosen before the run, and nothing applies itself until that comparison is done.
Who decides what the AI should do?
The person closest to the work, not the most technical person in the building.
Anthropic studied roughly 400,000 sessions of Claude Code, its AI coding tool. About 235,000 people used it between October 2025 and April 2026. People made about 70% of planning decisions and about 20% of execution decisions: people decide what to build, the AI decides how to build it. This is a vendor watching users of its own tool, not a controlled experiment, and coding is not all office work. Take the shape of the split as the design point, not the exact percentages.
The same study shows who can hold the framing job.
- Verified success rose from 15% of sessions for beginners to between 28% and 33% for intermediate users and above. "Verified" means the session looked successful and left hard evidence: saved code, a passing test, or the user confirming it worked.
- 19% of beginner sessions were abandoned, against 5% to 7% for everyone else.
Most of that gain is the step from beginner to intermediate, not from good to expert. So what matters is knowing the work, not mastering the code: enough understanding to steer and to recover when the tool goes wrong. The same coding-only limits apply.
Why do people stop checking AI output?
Because smooth, confident output switches off the checking reflex, and it does so hardest on exactly the cases where the tool is weakest. This is measurable, and it runs in the wrong direction.
Steven Shaw and Gideon Nave at Wharton gave more than 1,300 participants problems that each had a tempting wrong answer and a slower correct one. When the AI was right, accuracy rose 25 points above the no-AI baseline (the score people got with no AI at all). When it was wrong, accuracy fell 15 points below it. Participants felt more confident either way. The authors call this cognitive surrender: handing your judgment to the tool without noticing. The paper is a 2026 working paper, meaning other researchers have not yet reviewed it. These figures also come through the Washington Post's report of 7 July 2026 rather than the paper itself. So treat the size of both swings as provisional. The problems were built to carry a tempting wrong answer, so read it as what happens on the cases that matter, not as an average error rate.
The weak spots are not marked. Dell'Acqua, Mollick and colleagues ran a field experiment (a controlled test inside a real company) at Boston Consulting Group, published in Organization Science in 2026. Consultants using GPT-4 completed 12% more tasks, 25% faster, on tasks the model handles well. On tasks it does not, they made more mistakes than the comparison group. The study reports no number for those mistakes, so the size of that downside is unknown. It is peer reviewed, but it is one firm and an earlier model generation. Ethan Mollick calls that boundary the jagged frontier: the edge of what a model does well is uneven, and you cannot see it from inside a task.
Together, those two findings are a large part of why AI pilots fail at mid-market companies. The review step gets designed for the demo, where the tasks are the ones the tool handles well, then meets daily volume and nobody runs it. Build the check into the workflow at the weak points rather than leaving it to a reviewer's attention.
Is more human review always better?
No. A 2024 meta-analysis, a study that pools the results of many earlier studies, covered 106 human-AI experiments in Nature Human Behaviour. A person working with AI beat either one alone only when the person was the better performer at the task. When the AI was better, the pair did worse than the AI by itself.
That is an average across many experiments, so use it to decide where to put a reviewer, not to conclude that review is pointless. Human review raises accuracy when the reviewer is the better judge of that output and has a reference to check it against. A reviewer who is guessing adds noise.
Who owns the result when AI gets it wrong?
One named person per output, and the system gets called a tool. How a company labels its AI changes how hard people check it.
Emma Wiles of Boston University showed people the same work labelled two ways: as coming from a chatbot, or from an "AI employee". With the "AI employee" label, people caught 18% fewer errors. They were also 44% more likely to pass questionable work up to a manager rather than fix it themselves, which cancels the time saved. In her survey of 1,261 managers, about a third said their company frames AI agents as employees, and 23% list them on org charts. These figures come through MIT Technology Review's account of 29 June 2026 rather than the study itself. A second account says the effect was concentrated among managers at companies that already put AI on the org chart. So treat it as a risk that "teammate" framing creates, not a law that applies everywhere.
The mechanism matters more than the percentages. People feel on the hook for a direct report's errors and for a tool's errors. They do not feel on the hook for an "AI employee's". So naming the system a teammate moves accountability off the only person who could catch the problem.
Ownership is also where why AI adoption stalls stops being a question about models. Nobody's job description changed when the AI arrived, so the output lands in the gap between whoever prompted it and whoever received it.
Can a machine check the work instead?
Partly, and only where "correct" is something a machine can test.
Picking the right answer is harder than producing one. In the LLM-as-a-Verifier work from Kwok and colleagues at Stanford, UC Berkeley and NVIDIA, researchers asked what would happen if a model made several attempts at each task and a perfect judge picked the best one. On the Terminal-Bench V2 coding benchmark (a standard set of test tasks used to compare AI systems), that perfect judge would have solved 98.9% of tasks. That is a ceiling assuming a judge that does not exist, not a score anyone hit. The lesson is narrower: on that benchmark, producing a right answer was close to solved and picking it out of the pile was the hard part.
A machine check that works looks like this. In "Design Docs Are All You Need", researchers at Google DeepMind, MIT, Stanford and Google write the acceptance check (the test that says whether the result is right) into the plan itself. Every design document that states a number ends with a small worked example whose expected results are written out exactly, and the tests that enforce it are generated from that same document. Their rebuilt software reproduced hand-audited reference models down to rounding. That is one working system described by its builders, with no controlled comparison against other methods. The authors are explicit about the limit: it works only where a machine can check correctness. Where the output is a judgment, you are back to a person or a separate checker.
That limit is also the useful way to see what an agentic workflow is: a loop where the model chooses the next step and the tool for it, with no person in between. Safe where each step has a result a machine can check, a liability where it does not. When nobody owns verification for a step, the first decision is when not to use an agentic workflow at all.
A worked example, from a travel company where I led product and commercial. A faulty email parser, the software that reads flight confirmation emails and pulls out the itinerary, was replaced with a call to a Llama language model. A second prompt handled the roughly 5% of unusual cases the first pass missed. The parsing error rate came in under 2%. The replacement saved 750+ hours a year, and that figure is part of the 5,920+ hours a year saved by the wider AI adoption I led there, not an addition to it. Both are internal measurements that were never independently verified. The structure is the transferable part: the second prompt is a machine check sitting inside a workflow a person still owned.
A machine check does not delete the human job. It moves it from checking every output to auditing the checker on a schedule. A person pulls a sample now and then and confirms the check is still catching what it was built to catch.
What should a team keep for people?
Five rules, in the order they bite.
- The person closest to the work frames it. Knowing the work, not technical seniority, is what lets someone tell a good result from a confident wrong one.
- The reference is named before the run. A check invented after the output exists tends to be a reaction to the output.
- Checks are built in where the tool is weakest, not spread evenly. The edge of what a model does well is uneven and hard to see from inside a task, so even coverage spends the review budget on easy cases.
- Every AI output has one named owner, and the system is called a tool. The name changes how hard people check; the owner is who gets asked when it is wrong.
- Machine checks get audited by a person on a schedule. A check nobody has looked at since it was built is an assumption, not a control.
Checkpoints hold up under real volume. One agency running a content workflow with a person at every checkpoint, rather than one person at the end, puts the saving at roughly 75% of employee time on that workflow, by the owner's own count rather than a measurement.
Deciding which workflows need a person at which step, and which can run on a machine check, is what an AI readiness assessment is for. It is also the first thing worth settling in AI transformation work with mid-market operators, because every later tooling decision depends on it.
If you want to map which steps in your workflows still need a person and which can run on a machine check, a 45-minute working session on your actual process list is the fastest way to settle it. Get in touch to book one.
Next step · AI working session
Put this to work in your operation
Bring the process you want to automate and the data it runs on, and we come prepared. First conversation, no fee.
Book an AI working sessionSources
More insights
- What Is an Agentic Workflow? Harness First
AI Transformation ·
- Human in the Loop AI: Verify, Not Approve
AI Transformation ·
- Agentic Workflows: When Not to Use Them
AI Transformation ·