Fable 5.1's published benchmarks: AutomationBench nearly doubles
One number is doing most of the work in this release: AutomationBench, which measures whether a model can reliably automate real business pipelines, moves from 17.1% on Fable 5 to 31.4% on Fable 5.1. Roughly doubling, in one generation, on the benchmark closest to what most people actually want an agent for. The other numbers as read out: Terminal-Bench Science 1.0 above 50% against 24.7–29% for Fable 5 and Opus 5; agentic coding 55.8% against 42% and 52.3%, with GPT-5.6 Sol at 37.3%; GDPVal-AA-v2 1853 against 1723, 1824 and 1711; OSWorld 2.0 computer use 77.9% against 72.9% and 75.4%; Cursor Bench 3.2.0 73.4% against 70.5% and 70%. Anthropic also published score against mean cost per task on a log scale rather than per-token pricing — the argument being that a cheaper model that needs far more tokens is not actually cheaper. The reviewer's own caution is the right note to end on: benchmarks are not the experience of using a model, and he expects the real signal from hands-on "taste" tests over the following day. For the pricing, context limits and the breaking tool_choice change that come with this release, see our September 1 changelog entry.





