Skip to content
← The ledgerEntry 002 · 8 September 2026

ENTRY 002

What Would Count as Evidence

The economists say they're flying blind. Operators don't get that option.

Posted8 September 2026
Reading6 min
AuthorV. Kumar

In July, more than two hundred economists put their names to a public statement about artificial intelligence. Sixteen of them were Nobel laureates. The chief economists of OpenAI and Anthropic signed it. So did Erik Brynjolfsson, whose payroll-data work is the most credible empirical research on AI and employment anyone has produced.

The statement is called "We Must Act Now", and it is a call to prepare for AI's economic transformation. What makes it interesting is the reason it gives for urgency: the signatories are explicit that they cannot yet measure what is happening. One of its organisers, the economist Anton Korinek, put it to Fortune more bluntly — "we are driving in the fog."

Torsten Slok at Apollo has since catalogued five competing ways of measuring AI exposure — usage logs from two different assistants, expert judgment, self-assessment, and skill mentions in job postings. They disagree with each other, and they disagree most sharply exactly where the stakes are highest.

I found this oddly clarifying, and not for the reason you'd expect.

If the best-equipped people in the world, with national statistical apparatus and payroll microdata and decades of methodology behind them, cannot yet say what AI is doing to an economy, then the question "how do we know if our deployment worked?" is not a question they are going to answer for you.

It's an operator's question. And almost nobody is answering it from inside a P&L.

The gap between what we report and what shows up

Here's the shape of the problem, in three findings that don't agree.

McKinsey, surveying 1,719 respondents across 97 countries this May and June: 89% of organisations use AI in at least one function. 37% can attribute any EBIT impact to it. About 6% report it mattering — at least five percent of EBIT and self-described as significant. That 37% has not moved in a year.

BCG, surveying 152 CEOs at companies above half a billion dollars in revenue, in July: 89% see cost or revenue benefit in targeted areas, but only 14% have clearly defined P&L impact across all their AI initiatives.

NBER, published in March, surveying around 750 CFOs through the Atlanta and Richmond Fed panels: self-reported productivity gains of 3.0% expected for 2026, against revenue-based gains of 1.8%. Productivity running about seventy percent ahead of anything visible in the revenue line.

Different instruments, different samples, one shape. Nearly everyone is using it, a minority can find it in the accounts, and the people closest to the numbers report a gap between what they believe is happening and what they can demonstrate.

There are a few explanations available, and they have very different implications.

The gains are real and landing somewhere the accounts don't capture. The gains are real and being competed away in price before they reach margin. Or the gains are substantially imagined, and self-reported productivity is measuring enthusiasm.

I think all three are happening simultaneously, in different companies, and that nobody can currently tell which one they're in. That's the actual state of play in 2026, and it's less satisfying than either the boosters or the sceptics would like.

Why the standard business case fails here

The instrument most organisations use is time saved. Estimate the hours a task took, estimate the hours it takes now, multiply by a loaded cost, annualise.

It has one virtue: a finance function will accept it. It has one fatal flaw: it measures the thing that was already measurable, which is almost never where the value is.

I've watched this specifically. Automating a monthly reporting process in a US property business saved roughly a hundred and fifty thousand dollars a year in labour, and that was the number in the business case. The number that mattered was that decisions stopped waiting for a monthly cycle. A property that would have been decided in May got decided in March. That's worth considerably more and it appears in no line item, because there is no account called "decisions that happened sooner."

The same asymmetry runs through most AI deployments I've seen. Time saved is linear and countable. Waiting removed is compounding and invisible. So we count the small one and argue about the large one.

What I'd actually accept as evidence

Since nobody is going to hand us a framework, here's the one I use. It isn't elegant. It has the single advantage of being answerable before you spend the money.

Name the decision, not the task. What decision does this change, who makes it, and how long does that decision currently wait? If the answer is that no decision changes and something merely happens faster, you are buying efficiency, which is fine, but price it as efficiency and expect it to be hard to see.

Write down the counterfactual before you start. What would this number have been anyway? Almost nobody does this, and it is why so much AI reporting is unfalsifiable. If you can't state what would have happened without the deployment, you cannot later claim the deployment caused anything.

Pick a number you don't control. Internal productivity metrics are reported by the people whose work is being measured, which is why self-reported gains run ahead of revenue. Find a number that comes from outside — conversion, cycle time to cash, cost per completed transaction, repeat rate. Something a customer's behaviour produces rather than something your team estimates.

Set the failure condition in advance. What result would make you stop? If there is no such result, you are not running an experiment, you are running a programme with a budget. Both are legitimate. Only one of them produces evidence.

Separate the cost you removed from the cost you moved. A large fraction of reported AI savings is labour that didn't go away, it went somewhere else in the organisation. That's often good, and it is not a saving.

The reliability caveat that belongs in every business case

One technical point, because it bears directly on whether your evidence will hold.

METR measures how long a task frontier models can complete. The headline figures are for fifty percent success. At eighty percent success — still well below what you'd accept for anything touching a customer's money — the horizons are four to ten times shorter. For Claude Opus 4.6, twelve hours at fifty percent and about seventy minutes at eighty.

A paper published at the end of August, covering 10,664 trajectories across nine models, explains why. Success decays geometrically with each dependent step. Near-perfect to near-zero within sixteen steps.

This is why pilots succeed and deployments disappoint, and it is a measurement problem as much as an engineering one. Your pilot was three steps. Your production workflow is forty. The evidence from the first does not transfer to the second, and the failure will look like the technology underdelivering when it is actually the evaluation that was wrong.

If you take one thing from this: when someone shows you an agent benchmark, ask for the eighty percent number.

Where this leaves us

I don't think the measurement problem gets solved soon. The economists said as much, and they have better tools than we do.

What I think operators can do is stop pretending that the auditable number is the true number, and be explicit about which part of the case is a bet. "We can demonstrate two crore of labour saving, and we believe there's a larger effect on decision latency that we cannot yet measure, and here's what we'll watch to find out" is a more honest sentence than anything in most board decks.

In my experience it's also the sentence that gets the budget, because the people you're asking have been doing this long enough to know when a number is too clean.

Sources

Vijay Kumar is President & CTO of Xanadu Realty and writes at Divergent Economics. Sources: "We Must Act Now," July 2026, with Anton Korinek and Torsten Slok quoted in Fortune, 13 July 2026; McKinsey, The State of AI in 2026, fielded May–June 2026; BCG CEO survey, 22 July 2026; NBER Working Paper 34984, March 2026; METR Time Horizons, leaderboard updated 8 May 2026; arXiv 2609.01660, 31 August 2026.

The dispatch

A fortnightly note with the numbers attached.

Fortnightly. No sales email, and unsubscribing takes one click.