github.com
Benchmarking Opus 5 on SlopCodeBench
Opus 5 wins—and still fails 76% of checkpoints.
About Benchmarking Opus 5 on SlopCodeBench
HumanLayer tests Opus 5, Sonnet 5, and Opus 4.8 across 17 checkpoints from three SlopCodeBench challenges. The benchmark measures whether coding agents can evolve a codebase as requirements arrive incrementally. Opus 5 leads with a 24% strict pass rate, but no model completes any challenge defect-free, while verbosity, code volume, and complexity grow over time.
No comments yet. Have you tried it?