Analysis & perspective
Try local AI before buying the hardware: a rented RTX Pro 6000 running Qwen 3.8 27B against Opus 5.5
A practical way to answer "is local AI good enough for my work?" without buying anything. Rent the GPU class you’d buy, run the open-weight model you’d run, and test it on your own prompts. The takeaway: the local model was slower, not much worse, which makes it a better fit for background agents than for interactive coding.
"Don't Buy a $12,000 Mac Studio M5 Ultra for Local AI (Do This Instead)" by Craig Hewitt — Watch on YouTube →
The setup
- Hardware: one NVIDIA RTX Pro 6000 (96 GB) rented on Vast.ai at about $2–2.25/hour. You can pause or destroy the instance like any VPS. He loaded $10 of credit.
- Model: Qwen 3.8 27B, deployed from Vast's model library. It needs roughly 96 GB unquantised.
- Client: OpenCode with the rented endpoint added as a custom model. He had Codex do the wiring with computer use.
- Baseline: Opus 5.5 on medium in the Claude desktop app.
Three tests
- Customer-support reply with hard rules: a tie on quality. Qwen took 3 min 12 s.
- A Monday–Friday plan from messy inputs: a tie. Qwen added a useful contractor check-in, and both listed their open questions.
- A go/no-go marketing call with "unknowns stay unknown" rules: Opus won. Qwen reached a different recommendation and invented some numbers the prompt had forbidden.
A judging trick worth stealing: paste the prompt and both answers into a third model and ask it to compare adherence to constraints, maths, strength of inference and whether the recommendation follows from the evidence. He used Codex (GPT-5.6 Sol, high). He also likes Grok as an impartial judge.
Where local models fit
Interactive use is "probably the worst possible application for local models, because they're slow." Give them routine, non-urgent work on their own schedule instead: categorising expenses, channel audits, SEO analysis, month-end books. That's the pattern a scheduled agent like Hermes is built for (see Hermes cron jobs and free and local models for Hermes). Keep a fast frontier model for work you're actively shaping.
Buy vs rent
- The M5 Ultra's selling point is memory bandwidth (about 1.2 TB/s, double the M5 Max's 614 GB/s) in unified memory, at $7,300 for 96 GB up to about $12,000 for 256 GB. Delivery is months out.
- An RTX Pro 6000 costs about the same, but you still need a box and cooling. Three RTX 5090s (32 GB each) cost more in total, but you can add cards later. Clustering brings interconnect overhead.
- His whole test cost $2.82. Rent first, and buy only once your own prompts prove local is good enough.
Prices and model versions are as stated in the video on 2026-09-24. Check current pricing on the vendor sites. Compare running costs with our cost calculator.





