#benchmark
Wiki 5
- Benchmarking Opus 5 on SlopCodeBench Dex Horthy runs three Claude models through SlopCodeBench; the best, Opus 5, clears 24% strict pass
- Evaluating Quantized Models for Deployment ByteShape on why perplexity, KLD, and BPW don't rank quantized models for deployment
- GLM5.2 on AMD MI355X Wafer's vendor benchmark claiming 2626 tok/s/node for GLM5.2 on MI355X, and how they got there
- Interfaze AI startup building a deterministic-by-design model aimed at structured output and parsing tasks; published the open Structured Output Benchmark in April 2026
- Structured Output Benchmark (SOB) Interfaze's open benchmark for LLM structured output across text/image/audio with seven metrics; the load-bearing finding is JSON-Pass beats Value-Accuracy by 15-30 points on every frontier model