The Agent Can Produce Faster Than You Can Verify
AI can make production cheap enough that assurance becomes the scarce capacity.
As AI lowers the cost and time required to produce analysis, code, decisions, and other work, the scarce resource can shift from production to verification.
A 2026 MIT Sloan analysis of the economics of more capable AI put the constraint plainly: AI makes it cheap to produce work, but not to judge whether that work is any good. 1 That is not yet a claim that every enterprise is already verification-bound. It is a claim about asymmetry. Where outputs are consequential and hard or expensive to check objectively, realized value can be capped by assurance capacity rather than by generation speed.
A longitudinal case study of one mid-sized, unusually AI-forward software company makes the mechanism concrete. As AI-authored pull requests grew to about 90 percent of the total, raw volume grew 3.1× over the early-2025 baseline while the pool of developers acting as reviewers grew only 1.5×, so demand outran review supply and per-reviewer load roughly doubled. 2 That ratio is mechanism evidence from a single company, not a universal enterprise throughput number.
The control problem is therefore not whether to slow production. It is how to design verification as part of the production system: tier outputs by risk and reversibility, automate objective checks where they are decision-grade, preserve independent evaluation for consequential work, escalate unresolved cases, and measure the cost and latency of verified outcomes rather than raw output alone.
The agent that produces fastest is not the one that creates the most value. The one that can be verified at the speed of the decision is.
- Production and verification are not the same cost curve. AI can make generating work cheaper and faster without making judgment equivalently cheap or easy. That asymmetry can cap realized economic value. 1
- Review capacity can become a production constraint. In one mid-sized, AI-forward software company, raw pull-request volume grew 3.1× while the reviewer pool grew 1.5×, and per-reviewer load roughly doubled. That is a single-company observation, not a sector-wide ratio. 2
- Automated checks are part of verification capacity, not a substitute for the problem. The same enterprise case shows automated review expanding as human review supply was outrun. Structured checks can absorb part of the gap; they do not erase the need to design verification. 2
- Deployed AI still requires ongoing assurance. NIST treats validity and reliability for deployed AI systems as often assessed through ongoing testing or monitoring, not a one-time predeployment check. 3
- Verification should be risk-tiered, not uniformly manual. Low-stakes, reversible, or objectively testable outputs may not create a binding constraint. Consequential, hard-to-check work does. Autonoma’s synthesis is a five-part control model: tier, automate, independently evaluate, escalate, measure. 13
The evidence supports a three-part interpretation. First, cheaper generation does not automatically produce cheaper judgment. Second, when production volume grows faster than review supply, verification can become the binding constraint on realized value. Third, the right response is not universal human review. It is a verification architecture that matches intensity to consequence, reversibility, and objective testability.
Those three claims sit on different kinds of evidence and should not be collapsed into one slogan. The economic literature establishes a cost asymmetry between production and judgment. 1 The enterprise software case shows what happens when generation volume outruns the human review pool that used to absorb it, including a partial automated response that still leaves assurance design necessary. 2 Government primary guidance on deployed AI treats validity and reliability as often ongoing rather than one-time, which moves assurance into operating capacity rather than a closed predeployment gate. 3 Together they justify designing verification as part of production—not as a generic demand to slow every agent, and not as a claim that every workflow is already verification-bound.
Cheap production does not make cheap judgment
The economic mechanism is the load-bearing claim. In a 2026 MIT Sloan account of work by Catalini, Hui, and Wu, the verification gap is treated as a constraint on how fully the benefits of more capable AI can be realized: AI makes it cheap to produce work, but not to judge whether that work is any good. 1
That finding is an economic argument, not a census of enterprise backlogs. It does not establish that every organization already has a verification queue, or that there is a fixed verification-to-generation ratio. What it does establish is a credible separation between the cost of producing an output and the cost of standing behind it.
That separation matters in agentic workflows because the agent can now produce analysis, code, recommendations, and decisions at a rate that outruns the organization’s ability to check them. If the enterprise later treats unverified output as decision-grade work, it has measured production and recorded it as value.
One enterprise makes the constraint visible
The strongest empirical illustration comes from AI Writes Faster Than Humans Can Review, a 2026 longitudinal case study of a mid-sized, unusually AI-forward software company operating under a documented “2x” throughput mandate. 2
In that setting, as AI-authored pull requests grew to about 90 percent of the total, raw volume grew 3.1× over the early-2025 baseline while the pool of developers acting as reviewers grew only 1.5×. Demand outran review supply, and per-reviewer load roughly doubled. 2
Those figures are study-bounded. Adoption intensity was not randomized. The company is not a typical enterprise, and software pull-request review is not every form of agentic work. The Brief does not claim that a 3.1× production / 1.5× reviewer pattern exists across sectors.
What the case does establish is a mechanism: production can scale faster than the human review supply that previously absorbed it. The same study also shows the organization responding by expanding automated review as human review became scarce, while merge and revert rates held steady. 2 That is important counterevidence against a “more humans, item by item” prescription. Automation absorbed part of the gap. It did not make verification free, and it did not remove the need to know which outputs still required independent evaluation.
Assurance is continuous, and it should be tiered
NIST’s AI Risk Management Framework treats validity and reliability for deployed AI systems as often assessed by ongoing testing or monitoring. 3 That is a government primary-guidance statement about deployed systems broadly. It is not a requirement that every output receive human review, and it is not an incidence study of enterprise agent deployments.
It does change the design question. If validity is not settled at go-live, then verification capacity is part of the operating system, not a predeployment gate that can be closed. Continuous monitoring, automated checks, and retained evidence are legitimate forms of that capacity.
Autonoma’s synthesis is a five-part control model: tier, automate, independently evaluate, escalate, measure.
Tier outputs by consequence, reversibility, and objective testability. Automate deterministic or structured checks where they are decision-grade. Independently evaluate consequential outputs where the generator and the evaluator may share failure modes. Escalate unresolved or high-consequence cases, and allow a hold rather than forcing throughput. Measure verification latency and cost per verified outcome alongside raw production.
The model is not a claim that one review ratio will work everywhere. It is a way to make verification capacity observable and testable.
Autonoma forecast: Within 12–24 months, mature enterprise AI programs will increasingly track verification capacity alongside model and agent throughput, including review latency, automated-check coverage, exception rates, and cost per verified outcome.
The timing is uncertain; the underlying operating need is easier to see: once generation is cheap, the scarce resource is the capacity to stand behind what was produced.
- Verification throughput is measured against generation throughput for consequential workflows, not only as a backlog after the fact.
- Risk tiers become explicit product primitives. Outputs are classified by consequence, reversibility, and objective testability.
- Automated-check coverage is reported by tier, alongside evaluator disagreement, exception rates, and review latency.
- Independent evaluation is reserved and visible for high-consequence work where generator and evaluator failure modes may correlate.
- Escalation and hold paths exist. Unresolved cases can stop or slow throughput rather than being forced through.
- Cost and latency per verified outcome sit beside raw output volume in operating reviews.
For CIOs and CAIOs.
Treat verification capacity as part of production architecture for consequential agentic workflows, not as an after-the-fact review queue. Increasing autonomous throughput without a matching assurance design is a capacity decision, not only a model decision.
For CFOs and AI portfolio owners.
Evaluate economics using verified outcomes and assurance cost and latency, not gross generated-output volume alone. A surge in production that cannot be stood behind is not the same as a surge in value.
For risk, audit, legal, and compliance.
Define which outputs require independent evaluation, continuous monitoring, escalation, or retained evidence based on consequence and reversibility. Ongoing testing is a form of control, not an admission that the system was never ready.
For engineering and process owners.
Design workflows for objective testability, review capacity, evaluator independence, exception handling, and rollback or hold paths before increasing autonomous throughput. If the only path is “generate, then hope someone looks,” verification is already the constraint.
Weight: Strong. The strongest counterargument is that the Brief risks treating verification as a permanent human bottleneck when automated review, structured checks, and continuous monitoring can absorb much of the gap—and when many outputs are low-stakes, reversible, or objectively testable.
The evidence supports that objection. In the enterprise software case, automated review expanded as demand outran human supply, and merge and revert rates held steady. 2 NIST’s ongoing-testing language is itself a form of non-item-by-item assurance. 3 Low-stakes or easily tested work may never create a binding verification constraint.
The central empirical illustration also comes from one unusually AI-forward software company, not a cross-sector sample. 2 That is a meaningful limit on generalization.
The appropriate conclusion is therefore not “slow the agents” or “review everything by hand.” It is narrower: do not treat production growth as realized value unless verification capacity—automated, independent, or continuous—can keep pace with the consequence of the output. If an organization can show decision-grade checks, independent evaluation where it matters, and measured cost per verified outcome, the central risk described here is materially reduced for that use case.
Enterprises have spent the last two years celebrating generation. The agent writes the analysis, the code, the recommendation, the decision memo. Throughput goes up. The dashboard looks like progress.
Then someone has to stand behind the work.
That is the part that does not get cheaper at the same rate. In some workflows it will not matter: the output is reversible, the test is objective, the cost of being wrong is small. In the ones that do matter, verification is not a quality afterthought. It is the binding constraint on whether the cheaper production was ever value.
The design objective should not be to produce less. It should be to know what the organization can actually certify.