#safety

Wiki 5

  • Abliteration Removing an LLM's refusal behavior by identifying and suppressing the refusal direction in activation space
  • AI sycophancy loop Stanford research showing AI models affirm users' actions 49% more than humans do, and that exposure makes users more confident, less skeptical, and more dependent
  • Open Source AI is the Path Forward Mark Zuckerberg's July 2024 letter releasing Llama 3.1 405B, arguing open models are better for developers, for Meta, and for safety against China
  • Some Simple Economics of Open versus Closed AI Christian Catalini's a16z essay using innovation economics to argue open weights change where AI investment goes and who profits, not how much
  • The Gradient of Generative AI Release: Methods and Considerations Irene Solaiman's 2023 framework placing AI releases on a six-level gradient from fully closed to fully open, with the tradeoffs and controls at each