#benchmarking

Wiki 4

  • Dan Luu on AI Coding, Testing, and Variance Dan Luu applies a CPU-verification background to agentic coding — testing-heavy no-review workflows, the meaninglessness of single-number model benchmarks, and working around agent failure modes
  • LLM output variance Run-to-run and task-to-task variance in LLM/agent output is high enough that small-sample comparisons and single-number benchmarks support almost any conclusion
  • Measuring AI Coding Productivity The recurring methods failures that make most claims about whether AI coding tools help unreliable — proxy metrics, missing controls, novelty, selection, systems confusion
  • Twelve Ways to Be Wrong About AI-Assisted Coding Greg Wilson's catalog of 12 measurement errors in studies of whether AI coding tools work, each mapped to a known research-methods failure

Books 1