Ten seconds of a pixel creature shuffling on the spot above the words "Loading messages." That's an agent running my daily market news. And it is, oddly, one of the most satisfying things on my screen this year.
The satisfaction is easy to misread. It looks like the pleasure of speed, a machine moving faster than I can. I don't think that's it. What you are actually watching is a handoff. The moment the run starts, the morning's reading stops being mine to hold in my head. Something else is carrying it, and for those ten seconds I'm free.
But that freedom has a price, and I pay it at both ends of the shuffle: before it, when I write the brief, and after, when I decide whether to trust what comes back. This issue is about those two ends, because I think that's where the economics of agent work are actually being settled.
The obvious reading: agents make work faster
The standard story goes like this. Models can now complete longer and longer tasks without help. METR, which tracks this most carefully, has found that the length of task frontier models can complete on their own has been growing exponentially for years. Take that curve, multiply it by the number of knowledge workers, and you get the productivity case that underwrites a large share of today's AI capex.
I don't dispute the curve. I dispute the step where its output is treated as finished work.
Why that reading is incomplete
Two pieces of evidence sit awkwardly next to the curve.
The first is METR's own randomised trial. In July 2025 it reported on 16 experienced open-source developers working through 246 real issues in codebases they knew well. With AI tools allowed, they took 19% longer. Beforehand they had predicted a 24% speed-up, and afterwards they still believed they had been about 20% faster. METR is careful to call it a snapshot of early-2025 tools in one setting, and suggests AI may do worse where quality standards are very high or requirements are implicit: documentation, test coverage, formatting. That caveat is the point, though. Implicit standards are exactly what a brief leaves out. The model couldn't see them, and the humans spent their time making up the difference.
The second is in Anthropic's January 2026 Economic Index. On its API traffic, where tasks tend to run with less human back-and-forth, estimated success fell as tasks got longer. On the consumer product, where people iterate in conversation, it declined much more gently. And when Anthropic adjusted its own productivity estimate for those success rates, the estimate came down.
Put those side by side and a pattern shows. Unattended, the longer the task, the less often it lands. With a human in the loop, quality holds up better, but the human is now spending time. The productivity number depends less on what the model can do than on who catches what it got wrong, and at what cost.
The bottleneck moves
In the book I kept coming back to one phrase: the bottleneck moves. In compute it moved from chips to interconnect; in physical AI it moves from perception to integration. Fix the binding constraint and the next one becomes visible. Agent work is following the same script, and the constraint it is exposing is human.
The way I read it, every delegated task has two human costs wrapped around the machine's run.
The brief. An agent executes what you specified, not what you meant. Whatever you left implicit (the test-coverage norm, the client who hates a certain chart) gets either guessed or ignored. A brief written in a hurry is often the most expensive part of the task, because it sets what the whole machine run actually aims at. The METR developers paid for implicit standards on the back end. The cheaper place to pay is the front.
The check. Then something comes back that you didn't produce, and you have to decide whether it's right. This is a different act from doing the work yourself, and in some ways a harder one: you lack the trail of small decisions that would tell you where the weak points are. The work I came up in, transaction due diligence, is essentially this skill done as a profession. You're handed numbers someone else built, under someone else's incentives, and your job is to find where they bend. It's slow on purpose, and it's slow because it doesn't scale the way production does.
The asymmetry is this. The machine's side of the task keeps getting cheaper and longer-running. The human side doesn't follow the same curve. A model that can run for a working day still produces a working day's output to be checked, and checking long, unfamiliar work is where people are slowest and least reliable — or rather, where they're tempted to stop checking at all. That's what the METR participants did with their own speed: they felt faster and weren't. Nobody audits the feeling.
What this means for builders and capital allocators
In The Intelligent Economy I argued that once agents run for minutes or hours, the meaningful unit of work becomes the completed task rather than the token. I'd now tighten that. The unit is the accepted task: delegated, run, checked, and kept. The difference between a completed task and an accepted one is human time, and that time doesn't show up on the model invoice.
That changes how I'd look at a few things.
For anyone measuring agent ROI, the honest denominator is machine cost plus review time, priced at the reviewer's rate. In "The Missing Layer in Enterprise AI" I wrote about Dylan Patel describing how SemiAnalysis's AI spend ballooned within a year as his staff put it to work. The question I'd add now is how much of the payroll sitting next to that bill is people checking what the tokens produced. My guess is it's material, and that almost nobody is measuring it.
For investors, the bet I'd make is that a meaningful share of value in the agent stack accrues to whoever compresses verification: evals tied to one business's actual standards, and tooling that shows a reviewer where a long output is most likely to be wrong. Faster generation without faster checking just moves the queue. A company that shortens the check raises the ceiling on everything upstream of it.
For operators, the skill to hire and train for is specification. Writing a brief an agent can execute against is closer to writing a scope of work than to writing a prompt. It's a management skill, and until now people mostly acquired it once they had a team. They need it on day one now, working alone.
There's a fair counter. Agents are starting to check their own work, to run the tests and read the errors, and every improvement there takes load off the human end. Anthropic's June 2026 Economic Index also found that people who delegate most heavily report learning at the same rate as everyone else and feel more optimistic about their jobs, not less. If self-verification keeps improving, the human check shrinks toward a spot-audit and my argument weakens. I'd take that seriously. But it's a claim about checking, and checking is still the thing I'm saying sets the pace.
The signal to watch
The number I'd want is the gap between attended and unattended success on long tasks — the distance, in Anthropic's data, between the gently sloping curve where humans stay in the conversation and the steeper one where they don't. If that gap closes, agents are learning to check themselves, and the bottleneck moves again. If it stays open while time horizons keep growing, the scarce resource in the agent economy is the person who knows what good looks like, and has the time to look.
Watch that gap. Then go check the work.
Sources
METR, "Measuring AI Ability to Complete Long Software Tasks," 19 March 2025 — https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/
METR, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity," 10 July 2025 — https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
Anthropic, "Anthropic Economic Index report: Economic primitives," 15 January 2026 — https://www.anthropic.com/research/anthropic-economic-index-january-2026-report
Anthropic, "Anthropic Economic Index report: Cadences," 26 June 2026 — https://www.anthropic.com/research/economic-index-june-2026-report
Dylan Patel, "The Infinite Demand for Tokens, Claude Mythos, and Supply Constraints," Invest Like the Best, 23 April 2026 — https://podcasts.apple.com/us/podcast/dylan-patel-the-infinite-demand-for-tokens-claude/id1154105909?i=1000763216172
Carol Chen
Founder, Compute Notes
Builder of AI-native businesses and investor in AI infrastructure
LinkedIn: linkedin.com/in/carol-c-76461498
Website: computenotes.co
Network: network.computenotes.co
Get a free AI spend teardown → watt.computenotes.co
#AIInfrastructure #FinOps #EnterpriseAI #AIEconomics #Compute
