#ai-safety
Wiki 9
- A Safe Path to Open Weights Thinking Machines Lab's release framework for open weights - test the model, stage access, filter dangerous knowledge - applied to its Inkling models
- Emad Mostaque at TechBBQ: The Internet Will Go Offline (Trending Topics) Emad Mostaque's 2026 TechBBQ talk as reported by Trending Topics, on AI attacks, owning your cognition, labor losing value, and $1.50/hour robots
- Emotion Concepts in Claude Anthropic finds 171 emotion vectors in Claude that causally influence behavior, some invisibly to output
- Human-in-the-Loop Design pattern where AI agents pause for human approval on sensitive operations
- Nonproliferation is the wrong approach to AI misuse Helen Toner on why fixed dangerous capabilities cannot be kept from bad actors, and why the frontier-to-open lag should be used as an adaptation buffer
- On the Societal Impact of Open Foundation Models Kapoor, Bommasani et al.'s 2024 paper - five properties of open-weight models and a marginal-risk framework showing most misuse studies were incomplete
- Sam Altman May Control Our Future — Can He Be Trusted? Farrow & Marantz's 2026 profile of Altman: Ilya Memos, Amodei docs, safety erosion, Gulf ambitions
- The Myth of unsafe Open Source AI Florian Brand's survey of third-party incident reports - real AI misuse in 2025-2026 ran mostly through closed models, except image abuse
- Why Are AI Agents Lying, Cheating and Coordinating? (Bengio) Yoshua Bengio traces agent misbehaviour to imitation plus reward-seeking, and argues patching symptoms may only select for better-hidden cheats