Summary
Harness Arena: blind-judge Claude Code, Codex, Hermes, OpenClaw and OpenCode on the same task and model
Most model benchmarks hold the harness constant and vary the model. Harness Arena does the reverse: same task, same model, different harness, with blind human judging. That makes it the first public attempt to measure how much the harness itself changes the result, which is what our comparison pages keep running into. Disclosure from the video: it is a sponsored walkthrough, and the sponsor (On Demand) has its own harness in the arena.
"Harness Arena (Fully Tested): This NEW Benchmark TESTED Every AGENT HARNESS (which is the best?)" by AICodeKing — Watch on YouTube →
Step-by-Step Breakdown
- Browse recorded runs.
Open
harness-arena.ai→ Battle log. Filter by status (queued, in progress, awaiting judgment), category, or outcome. Expanding a round shows one column per harness, plus the model configuration (the example used Qwen 3.8), scores, community rating and deliverables. - Judge a round (requires an account).
Click the task title. Read the task and rubric, inspect the anonymous outputs, and score every required output 1–10. Harness names are revealed only after you submit. His advice: build a short checklist from the rubric and apply it the same way to every output. Check that the requested change was made and that nothing unrelated broke. For the 3D globe task, that means confirming the overlay is gone and that rotation and zoom still work.
- Read the leaderboard carefully.
Leaderboard → pick a category (Code, Research, Operations). Columns: rating, win rate, votes, W/L, median completion time. When he recorded, several harnesses tied on one vote each, so small differences mean nothing yet.
- Run your own benchmark.
New benchmark → choose tasks (built-in, or upload your own with the downloadable Excel template, one task per row), pick at least two harnesses, pick a model, and submit. He suggests starting with one task and two harnesses. Completed tasks can be judged before the whole dataset finishes. Including On Demand requires its API key.
- Self-host if you prefer.
The back end is on GitHub under the MIT license. Running it yourself still costs model and hosting fees.
Gotchas & Caveats
- Sponsored, and the sponsor competes. Weight the leaderboard accordingly until vote counts are large.
- A harness that wins on a free open-weights model may not win on a frontier model. The ranking only holds for the model it was run with.
- Speed alone is misleading. A fast run that leaves broken code is still a loss.
Key Takeaways
- Holding the model constant is the only fair way to compare harnesses.
- The most useful feature is uploading your own recurring tasks, which beats any public leaderboard.
- Blind judging against a written rubric is a method you can reuse internally even without the site.





