#benchmarking
Wiki 4
- Dan Luu on AI Coding, Testing, and Variance Dan Luu applies a CPU-verification background to agentic coding — testing-heavy no-review workflows, the meaninglessness of single-number model benchmarks, and working around agent failure modes
- LLM output variance Run-to-run and task-to-task variance in LLM/agent output is high enough that small-sample comparisons and single-number benchmarks support almost any conclusion
- Measuring AI Coding Productivity The recurring methods failures that make most claims about whether AI coding tools help unreliable — proxy metrics, missing controls, novelty, selection, systems confusion
- Twelve Ways to Be Wrong About AI-Assisted Coding Greg Wilson's catalog of 12 measurement errors in studies of whether AI coding tools work, each mapped to a known research-methods failure
Books 1
- The Art of Computer Systems Performance Analysis Jain's 1991 textbook on experimental design, queueing models and statistics for benchmarking