#reinforcement-learning
Wiki 6
- A $500 RL Fine-Tune That Beat the Frontier A $500 GRPO fine-tune of a 9B open model beat every frontier config on catalog review, 68× cheaper
- GLM-5.3: How Chinese labs keep stride with the frontier Lambert on GLM-5.3 and why Chinese labs match US models without relying mainly on distillation, plus Z.ai's staged release for cyber capabilities
- How distillation is used today and what performance uplift it gives to open models Lambert's July 2026 note that distilled data seeds SFT for Chinese labs but matters less as RL grows, written against a Stratechery claim
- Olmo 3 AI2's Olmo 3 report (Dec 2025), fully open 7B/32B Base, Think, Instruct and RL-Zero models with every stage's data, code and checkpoints released
- SWE-1.7 — Cognition's Coding Model Cognition's SWE-1.7, RL-trained from Kimi K2.7, reportedly near GPT-5.5 on coding benchmarks — with the strongest numbers on its own FrontierCode
- Synthetic Data & Distillation | RLHF and Post-Training Book by Nathan Lambert Chapter of Lambert's RLHF book on synthetic data, from SFT distillation and on-policy KD to AI feedback, Constitutional AI and rubrics