#interpretability
Wiki 4
- Emotion Concepts in Claude Anthropic finds 171 emotion vectors in Claude that causally influence behavior, some invisibly to output
- Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling EleutherAI's 2023 Pythia suite, 16 models from 70M to 12B trained on the Pile in one fixed order, with 154 checkpoints each for training-dynamics research
- Steering Is Interesting Again (Goedecke) Goedecke's case that steering vectors are worth a second look now that local frontier models exist, with antirez's DwarfStar 4 as evidence
- Steering Vectors Manipulating LLM activations mid-inference to bias output toward a concept, without changing the weights