Published: 2026-09-14
Analysis & perspective

Pick models with five benchmarks, not one index: Terminal Bench, Apex Agents, Automation Bench, Omniscience, Deep SWE

Chapters / key moments (click to jump — plays here on the page)

What you can take from this is a method for choosing models, plus current readings of it for September 2026. The method: choose benchmarks that match your work, look at performance, cost and speed together, and run a model stack (state-of-the-art, workhorse, lightweight) rather than one model. The readings below will go stale. The method won't.

Source video

"Agentic Engineering Benchmarks: How I RANK Astra, Fable 5.1, and Open-Weights" by IndyDevDanWatch on YouTube →

The five benchmarks and why each is there

  1. Terminal Bench (v4): pure agentic coding.

    60+ tasks, each run in a prepared container with a verifier. Reading: Astra leads, Fable 5.1 second, then a big drop around 40–42%. GLM 5.3 beats Kimi K3 here even though other rankings put Kimi higher. On cost, Astra is about 4× cheaper than Fable and about 2.4× cheaper than Fable 5.1, and uses about 2.5–2.7× fewer tokens.

  2. Apex Agents: knowledge work outside coding.

    Investment banking, consulting and legal tasks written by practitioners, with short, realistic prompts. Astra and Fable lead, with Muse Spark 1.1 a surprise third in consulting. It has no cost or time data.

  3. Automation Bench: work across business apps, with guardrail violations counted.

    600+ tasks across finance, HR, marketing, ops, sales and support. Completing the task while breaking a guardrail counts as a fail. Grok 4.6 and GLM 5.3 rank high. If you ignore violations, Fable and Opus jump up, which tells you they finish tasks but break rules more often. Per-app readings help too: for Gmail work, Grok 4.6, Qwen 3.8 and GLM 5.3 hold up, and DeepSeek does not.

  4. AA Omniscience: hallucination.

    Answers are graded correct, incorrect, partial, or not attempted, and "I don't know" costs nothing. The top labs lead. Gemini 3.8 Flash and Muse Spark 1.3 are good runners-up, and GLM 5.3 Flash is risky. His design lesson: give your own agents an explicit way to say "I can't do this".

  5. Deep SWE (1.1): long-horizon software engineering.

    Short prompts, long tasks. Astra is first (about 30K output tokens, 29 steps, low cost), and Gemini 3.8 Flash is second, slower but far cheaper. Fable 5.1 was missing from the board when he recorded.

How to apply it

  • Write down the three or four kinds of work your agents do, and pick one benchmark that matches each.
  • Pick one model as the control (he uses Astra) and read every chart relative to it.
  • Treat "useful agent output per hour of token spend" as the metric that matters. Score alone doesn't capture it.
  • Fill three slots: state-of-the-art, workhorse (Gemini 3.8 Flash, GLM 5.3), and lightweight (Luna, GLM 5.3 Flash). Route work to the slot it needs.
  • Ignore benchmark providers that leave out major models, and check they add new releases promptly.

We collect these leaderboards in one place: benchmarks. Cost-model a stack in the cost calculator.