Map

Measuring AI Coding Productivity

Wiki conceptai-codingmeasurementbenchmarkingllm-skepticism โ†ณ show in map Markdown
title
Measuring AI Coding Productivity
type
concept
summary
The recurring methods failures that make most claims about whether AI coding tools help unreliable โ€” proxy metrics, missing controls, novelty, selection, systems confusion
tags
ai-coding, measurement, benchmarking, llm-skepticism
created
2026-07-21
updated
2026-07-22

Whether AI coding tools make developers more productive is an empirical question, and almost every published answer is compromised by how it was measured. The failures aren't exotic โ€” they're the standard problems the human sciences catalogued decades ago, showing up because software engineering mostly builds its own study designs without borrowing that discipline. twelve-ways-wrong-ai-coding (Greg Wilson) is the fullest enumeration; this page pulls the recurring failure types into one place and pins the strongest evidence to each.

Computing has had its own version of that discipline since 1991: Raj Jain's art-of-computer-systems-performance-analysis covers experimental design, workload selection, and the statistical treatment of measurement results, and the AI-coding literature largely rediscovers its chapters the hard way.

The failure types

Proxy-metric / Goodhart failures. Lines of code, commit counts, PRs closed, and suggestion-acceptance rate are all easy to count and all measure activity rather than value. A 40% rise in lines per developer measures verbosity; deleting tangled code and replacing it with less is progress that reads as a loss. Once any of these becomes a target it stops tracking the thing you cared about โ€” developers make more, smaller commits, or split tickets. Acceptance rate is the sharpest case: an enterprise study of 400 developers found a 33% acceptance rate with high satisfaction and no measure of whether accepted code was correct or secure.

Missing control group / no counterfactual. Before/after comparisons attribute every change over the period to the tool, when you also hired, refactored CI, and changed vendors. Comparing AI-assisted developers against a control using "nothing" is worse โ€” that baseline doesn't exist, because real developers without an LLM still have documentation, colleagues, and thinking time. Internal validity requires a credible counterfactual, which most industry reports lack.

Novelty and Hawthorne effects. New tools feel fast, and observed people work differently. A four-week study that finds a boost has found a four-week boost; the effects that decide whether the tool was worth it โ€” skill atrophy (dont-outsource-learning), debt accumulation, changed collaboration โ€” surface over months. Self-report surveys stack a third contaminant on top: social-desirability bias, strongest when management chose the tool.

Selection bias. Comparing volunteers against non-volunteers compares two populations, not two conditions. Early adopters are already more motivated and often higher-performing; one longitudinal study found Copilot users were more active than non-users before the tool was introduced. It's the most common flaw in industry reports because it's the cheapest study to run.

Individual vs system. Writing code 30% faster changes nothing if ticket-to-production time is flat โ€” writing code wasn't the bottleneck, and more generated code means more to review. This is where the local win becomes a global loss.

The strongest evidence

Four results carry most of the weight because they avoid the failures above:

  • METR RCT (+19% time). A randomized controlled trial with experienced open-source developers found AI tools increased completion time by 19% โ€” the opposite of what participants predicted for themselves. A real control group on real (not greenfield) tasks, which is why it outranks the "55% faster" Copilot figure from a ninety-minute from-scratch exercise.
  • Xu 2025 (senior โˆ’19% / +6.5% review). Junior contributors gained output while senior developers lost 19% of their own productivity absorbing a 6.5% increase in review load from AI-generated code โ€” the individual-vs-system failure measured directly.
  • He 2026 (807 repos). Across 807 open-source repos adopting Cursor, a large but transient velocity increase alongside a persistent rise in code complexity and static-analysis warnings โ€” the novelty effect and the deferred cost, separated by measuring over time.
  • Liu 2026 (300k+ commits, >15% flawed). Over 15% of more than 300,000 AI-authored commits introduced at least one quality issue, and nearly a quarter of those persisted long-term โ€” the "easy half" (generation) looking good while the hard half (correctness, maintenance) degrades.

Relation to other pages

danluu-ai-coding-testing argues the same point from the bench: single summary numbers hide run-to-run variance wide enough that a few runs support any conclusion, so "show me the distribution." no-silver-bullet-llms supplies the DORA/CircleCI instability data that the individual-vs-system failure predicts at the delivery level. reviewing-ai-code is the individual-vs-system failure restated as a hard throughput cap โ€” review tops out near 400 LOC/hour, so faster generation just moves the bottleneck. constraint-decay explains why greenfield-task timings mislead: agents look strong on loosely specified work and degrade as production constraints accumulate. dont-outsource-learning is the long-run cost that novelty-window studies structurally cannot see.