Deep dive
Opus 5.5 vs GPT-6 Astra: 12 tasks, 17 hours of agent time, and a per-task cost sheet
This is the rare comparison video that reports time and cost on every single task, which is what makes it useful rather than merely entertaining. Both models were run on high effort in their own harnesses, on identical prompts. The quality result and the cost result point in opposite directions, and that tension is the actual finding.
"I Tested Opus 5.5 vs. GPT-6 Astra on 12 Real Use Cases" by Nate Herk — Watch on YouTube →
How the test was run
The same prompt was fired at each model inside its own harness, both on high effort. The work was done on subscription plans, but every task is also reported at API-equivalent billing, so the numbers mean something regardless of which plan you are on.
One methodology note worth repeating, because he raises it himself: in the previous day's Sol comparison, two experiments were spoiled when the two agents noticed each other and started coordinating. Here they were kept apart. If you run your own bake-offs, that is the trap to avoid.
The pricing setup
At list API rates, GPT-6 Astra is roughly 2.5× the per-token price of Opus 5.5. The framing question for the whole video is therefore not "which is better" but whether Astra delivers 2.5× the value. The measured answer is that it does not need to, because it burns far less wall-clock time and far fewer tokens reaching a comparable result.
Current published rates for both models are in our cost calculator, which is the number to plan against rather than any figure quoted in a video.
The totals
| Measure | Claude Opus 5.5 | GPT-6 Astra |
|---|---|---|
| Tasks won (his judgement) | 8 | 4 |
| Total agent run time | ~11 hours | ~6 hours |
| Total cost at API billing | $214 | $132 |
Scoring is explicitly subjective — his taste, his use cases. The time and cost columns are not.
Where each one won
Opus took the creative and judgement work. Video editing (31 minutes / $10 against Astra's 39 minutes / $21), the Instagram reel, the branded deliverable set, the 3D learning world, and the Canva browser-use drawing test — where Astra's output was weak enough that he flagged it as a regression from earlier Astra runs.
Astra took the specified, bounded work. On the code review both models scored 100/100, but Astra got there in a fifth of the time at half the cost. It also won the trip itinerary ($11 against $14.47) and the Instagram carousel (about $7 against $10).
His summary: Claude models behave like a wise old owl — judgement and taste, and they cope with vague goals. GPT models behave like a very good obedient worker — give them an exact spec and they execute it efficiently, but vague goals frustrate them. That maps directly onto how you route work, not just which subscription you buy.
The budget argument he raises against himself
On the 3D-world task Opus ran 1 hour 44 minutes; Astra finished far sooner for about $12. His own question: if you gave Astra the same $60 of budget and let it iterate with feedback until it hit that ceiling, would it beat Opus's single pass? He thinks probably yes. A single-shot comparison systematically flatters the slower, more expensive model — worth remembering whenever you read a bake-off, including this one.
Key Takeaways
- Run competing agents in separate working directories — agents that can see each other's work will start coordinating and spoil the test.
- Route by task shape, not by leaderboard: vague and creative to Opus, tightly specified and bounded to Astra.
- The 2.5× per-token price gap did not produce a 2.5× bill. Per-token price is a poor proxy for cost per task.
- Both models cleared the bar on ordinary knowledge work; the differences showed up on taste and on long autonomous runs.
- Single-pass comparisons favour the slower model. Budget-matched comparisons would likely read differently.
The companion test against the cheaper tier: Opus 5.5 vs GPT-6 Sol. Current rates for every model named here: cost calculator.





