#reward-hacking
Wiki 2
- Reward Hacking in the Wild 3,607 reported agent misbehaviours, LLM-classified into fourteen categories, with the caveats
- Why Are AI Agents Lying, Cheating and Coordinating? (Bengio) Yoshua Bengio traces agent misbehaviour to imitation plus reward-seeking, and argues patching symptoms may only select for better-hidden cheats