#verification
Wiki 3
- On the Self-Verification Limitations of LLMs on Reasoning and Planning Tasks Stechly, Valmeekam and Kambhampati find GPT-4 self-critique loops lose accuracy while a sound external verifier gains it, even with no critique text
- Reviewer Capability Governs Rejection Targeting, Not Repair Skill A 100-problem pilot where a cross-family reviewer adds 12 points and self-review, despite the best recall, falsely rejects 35% of correct answers
- Variation in Verification: Understanding Verification Dynamics in LLMs Salesforce study finding errors from stronger generators are harder for any verifier to catch, so a weak generator plus GPT-4o nearly matches a strong one
Toolbox 1
- vera Language meant to be written by LLMs — no variable names, mandatory contracts, effect rows