#post-training
Wiki 8
- 2 OLMo 2 Furious AI2's OLMo 2 report, fully open 7B/13B/32B models on up to 6.6T tokens with a stability fix list, Dolmino mid-training and the Tülu 3 RLVR recipe
- Frontiers in synthetic data Lambert's 2024 notes on synthetic data in post-training, from SFT on GPT-4 outputs to Gemini Flash being distilled from Pro
- How distillation is used today and what performance uplift it gives to open models Lambert's July 2026 note that distilled data seeds SFT for Chinese labs but matters less as RL grows, written against a Stratechery claim
- How much does distillation really matter for Chinese LLMs? Lambert reads Anthropic's February 2026 disclosure against DeepSeek, Moonshot and MiniMax and argues distillation helps but is not decisive
- LLM Distillation Training one LLM on another's outputs, from SFT on generated text to on-policy KD, and the 2025-2026 fight over Chinese labs distilling US models
- Olmo 3 AI2's Olmo 3 report (Dec 2025), fully open 7B/32B Base, Think, Instruct and RL-Zero models with every stage's data, code and checkpoints released
- OLMo: Accelerating the Science of Language Models AI2's first OLMo release (Feb 2024), 1B and 7B models on 2T+ Dolma tokens with weights, data, code, logs and 500+ checkpoints under Apache 2.0
- Synthetic Data & Distillation | RLHF and Post-Training Book by Nathan Lambert Chapter of Lambert's RLHF book on synthetic data, from SFT distillation and on-policy KD to AI feedback, Constitutional AI and rubrics