We Are Measuring AI Productivity Wrong, and It Matters
Task-level benchmarks flatter the technology while firm-level statistics lag it, and policy built on either alone will misfire.
Oak Ridge National Laboratory via Wikimedia Commons · CC BY 2.0Two sets of numbers dominate the AI productivity debate, and both mislead. Task-level studies show dramatic gains, code produced faster, documents drafted in minutes, and imply an economy on the verge of transformation. Aggregate statistics show modest movement and imply the whole thing is hype. The truth lives in the gap, and the gap is where measurement fails.
Task gains overstate because tasks are not jobs. Accelerating the drafting of a document does not accelerate the meeting that decides what the document should say, the approval chain it must survive, or the process it feeds. Organizations metabolize speed slowly, and the bottleneck simply moves.
Aggregate statistics understate because they were built for a different economy. Quality improvements, avoided hires, and work that shifts from billable services to internal automation register poorly or not at all in output measures designed around counting things.
The measurement gap has a policy cost: it fuels both complacency and panic on schedule. What would help is unfashionable, firm-level studies that trace where task gains do and do not become organizational ones. Until then, the honest position on AI productivity is the one no headline wants: it is large, uneven, and largely unmeasured.