DeepSeek V4 Pro 0813 driving Hermes agent: agentic benchmarks jump, cost stays low
DeepSeek has shipped the production build of V4 Pro (the 0813 build), and the agentic benchmark movement is the story: Terminal Bench 2.1 from 72 to 88, CyberGym from 53 to 83. Fahd Mirza points Hermes agent at the model over API and gives it two unaided tasks — repairing a deliberately broken three-service application, and a one-shot raw-canvas rendering test — to see whether the reported numbers translate into agentic work. They largely do, with the significant caveats that the benchmark figures are DeepSeek's own and that the current pricing is expected to rise.
"DeepSeek V4 Pro 0813 with Major Agent Upgrade: Tested Locally" by Fahd Mirza — Watch on YouTube →
Key Takeaways
- What changed in the release. The 0813 production build replaces the earlier preview. Terminal Bench 2.1 goes 72 → 88 and CyberGym 53 → 83, with gains across nearly every agentic row rather than a single spike. Against GLM 5.2, Gemini K3, Opus 4.8 and Fable 5 it trades blows, leading on some rows and trailing on knowledge-heavy ones.
- How to use it today. It is available through the API and works as the backing model for Hermes agent — which is how the whole video is run. Weights are stated as opening soon; local execution is out of reach for most hardware in the meantime.
- The debugging test is the useful signal. The task was a report-generation app with a FastAPI backend, a Redis pub/sub worker and a SQLite store, broken such that submitted jobs sat pending forever. The model was given the goal with no hints and no hand-holding, and fixed the pipeline end to end — a multi-service failure requiring the agent to reason across backend, worker and frontend at once.
- Token efficiency improved alongside accuracy. Mirza notes the fix consumed noticeably fewer tokens than he expected, which supports DeepSeek's claim that the agent behaviour is more efficient and not merely more capable.
- The creative one-shot test took 16 minutes. A single self-contained HTML file animating a spit roast with working knobs for spin speed, flame height, seasoning and basting — hand-built in raw canvas and JavaScript, no libraries, no build step. It worked on open, with the searing and charring state changes rendering correctly; Mirza rates the spin animation as the weak point.
- Price positioning is the real argument, and it is temporary. The model lands at the front of the pack while costing a fraction of the frontier closed models — but Mirza flags directly that DeepSeek is expected to raise prices significantly, so the current chart "might not last for long."
- Treat the benchmark numbers as vendor-reported. He says so explicitly before showing them, and frames the hands-on tasks as the check on whether they mean anything.
Also Mentioned
Qwen 3.8 Max, a 2.4-trillion-parameter model, landed on Hugging Face the same day.





