Published: 2026-09-24
Analysis & perspective

Try local AI before buying the hardware: a rented RTX Pro 6000 running Qwen 3.8 27B against Opus 5.5

Chapters / key moments (click to jump — plays here on the page)

A practical way to answer "is local AI good enough for my work?" without buying anything. Rent the GPU class you’d buy, run the open-weight model you’d run, and test it on your own prompts. The takeaway: the local model was slower, not much worse, which makes it a better fit for background agents than for interactive coding.

Source video

"Don't Buy a $12,000 Mac Studio M5 Ultra for Local AI (Do This Instead)" by Craig Hewitt — Watch on YouTube →

The setup

  • Hardware: one NVIDIA RTX Pro 6000 (96 GB) rented on Vast.ai at about $2–2.25/hour. You can pause or destroy the instance like any VPS. He loaded $10 of credit.
  • Model: Qwen 3.8 27B, deployed from Vast's model library. It needs roughly 96 GB unquantised.
  • Client: OpenCode with the rented endpoint added as a custom model. He had Codex do the wiring with computer use.
  • Baseline: Opus 5.5 on medium in the Claude desktop app.

Three tests

  1. Customer-support reply with hard rules: a tie on quality. Qwen took 3 min 12 s.
  2. A Monday–Friday plan from messy inputs: a tie. Qwen added a useful contractor check-in, and both listed their open questions.
  3. A go/no-go marketing call with "unknowns stay unknown" rules: Opus won. Qwen reached a different recommendation and invented some numbers the prompt had forbidden.

A judging trick worth stealing: paste the prompt and both answers into a third model and ask it to compare adherence to constraints, maths, strength of inference and whether the recommendation follows from the evidence. He used Codex (GPT-5.6 Sol, high). He also likes Grok as an impartial judge.

Where local models fit

Interactive use is "probably the worst possible application for local models, because they're slow." Give them routine, non-urgent work on their own schedule instead: categorising expenses, channel audits, SEO analysis, month-end books. That's the pattern a scheduled agent like Hermes is built for (see Hermes cron jobs and free and local models for Hermes). Keep a fast frontier model for work you're actively shaping.

Buy vs rent

  • The M5 Ultra's selling point is memory bandwidth (about 1.2 TB/s, double the M5 Max's 614 GB/s) in unified memory, at $7,300 for 96 GB up to about $12,000 for 256 GB. Delivery is months out.
  • An RTX Pro 6000 costs about the same, but you still need a box and cooling. Three RTX 5090s (32 GB each) cost more in total, but you can add cards later. Clustering brings interconnect overhead.
  • His whole test cost $2.82. Rent first, and buy only once your own prompts prove local is good enough.

Prices and model versions are as stated in the video on 2026-09-24. Check current pricing on the vendor sites. Compare running costs with our cost calculator.