Last updated: 2026-09-19

Code generation & program synthesis — Benchmark Sources & Consensus

Generating or reconstructing working programs from specs, binaries, or natural-language descriptions. Frontier-difficulty benchmarks where state-of-the-art is still well under 10%.

Platforms tracked: Claude Cowork · Kilocode · Openclaw · Chatgpt

Consensus across 4 sources

Across 4 sources, results depend heavily on task size. Framework-level Rails tasks reach 92% top accuracy, while feature-scale tickets (35% best, GPT-6 Astra) and private enterprise codebases (38.8% best, Fable 5.1 in Claude Code) fall below 40%. ProgramBench's binary-reconstruction ceiling sits near 0%. Sources disagree on the leader, and no single model leads all four.

All Sources

We aggregate published benchmarks; we never run our own tests and never pick winners. Each row links back to the original publication.

SourceDateFindingMethodologyQuality
ProgramBench 2026-05-09 Claude Opus 4.7 leads at 3% "almost resolved" on 200+ program-reconstruction tasks from compiled binaries; 0% fully solved by any model. Frontier-difficulty. Reconstruct working source from compiled binary; automated pass/fail on hidden test suite. 200+ tasks across 12 languages. · 200+ tasks high winner: cowork
Ruby on Rails 2026-08-13 21 Rails tasks on a realistic app across 8 models: top accuracy 92%, cheapest full run 91 cents; models rarely reached for existing framework APIs (recall 8-35%), and success was higher when they did (92% vs 87%) 21 tasks on Writebook, each targeting one Rails API without naming it; Claude Opus 5, Claude Fable 5, GPT-5.6 Sol, GPT-5.6 Luna, Muse Spark 1.2, Kimi K3, GLM 5.2, DeepSeek V4 Flash; accuracy plus API-recall and cost recorded per model high
Ruby on Rails 2026-09-09 Feature-scale Rails tickets: GPT-6 Astra best at 35% of runs ($2.51/task), Fable 5.1 30%, Gemini 3.8 Flash 28%; GPT-5.6 Luna 0% 20 feature tickets on the Fizzy app, 10 models, 3 runs each; follow-up to the 21-task stage 1 high
withspecific.com 2026-09-11 Private enterprise codebases: Fable 5.1 (Claude Code) 38.8%, GPT-6 Astra (Codex) 33.8%, Grok 4.6 32.5%, Gemini 3.8 Flash 31.2% 10 tasks from private production repos, Harbor format, verifiers from existing test suites · 10 tasks medium

How we work

OpenClawDatabase aggregates and links to published benchmarks. We don't run our own tests, and we don't pick winners. Our weekly benchmark-aggregator routine scans 7+ live leaderboards (OpenRouter, Aider, SWE-bench, GAIA, LMSYS, BigCodeBench, MMLU-Pro) plus relevant Reddit and Hacker News threads, then writes structured entries into /assets/benchmarks.json. Every row here links back to the original publication.

← Back to all benchmark tasks · See also: Decision guide · Cost calculator

📬 Weekly Digest — In Your Inbox

One email a week: top news, releases, and our deepest new guide. No spam. Same content via RSS if you prefer.