Run a 24/7 Private Hermes Agent on Your NVIDIA DGX Spark
Alex Finn demos setting up Hermes Agent on an NVIDIA DGX Spark powered entirely by a local model—no cloud API, no subscription fees, completely private. The DGX Spark runs headless (no monitor needed) and connects to your main machine via Tailscale, letting Hermes manage the device, install models, and run agentic tasks around the clock. The entire setup—including Tailscale installation—is handled by giving Hermes a single plain-English prompt.
"Hermes Agent powered by local models on the DGX Spark is basically magic" by Alex Finn — Watch on YouTube →
Key Takeaways
- DGX Spark runs in headless mode—plug it in, and Hermes Agent on your main machine can control it via Tailscale from anywhere in the world.
- Prompt to get started: "I purchased a new DGX Spark and want to set it up. I want to run it headless and I want you to be able to control it. Walk through setup, then install Tailscale."
- Local models are free once the hardware is paid for—no per-token billing, fully private, all data stays on device.
- LoRA adapters let you customize the local model's voice and output style to match your own.
- Use Hermes connected to a cloud model first to set up the Spark, then switch to the local model once it's running.
What you can actually set up from this
Extracted from the video's own transcript — the specifics the original summary left out.
Commands & Code Shown
hermes -p Qwen
Purpose: Launch the Hermes profile named Qwen that is wired to the local model
When to use: After Hermes has created the profile for you
Reproducible steps
- Run the Spark headless and join its network
No monitor needed. Power it on and connect your main computer to the network printed in the manual.
- Let a cloud-backed Hermes set it up
Install Hermes on your main computer with a cloud model first, then prompt: set up the new DGX Spark headless so you can control it, and install Tailscale on it so any device on your tailnet can reach it.
- Install the local model
Prompt Hermes to download and install Qwen 3.6 27B on the Spark and load it into memory. It found the right build and ran it under a llama server; the download took about 20 minutes.
- Create a second Hermes profile on the local model
Prompt: set up a new Hermes profile plugged into the local model and name it Qwen. Hermes detects the llama server, asks permission, and creates the profile.
- Tell the new profile its name
A new profile doesn't know the name you gave it. Say 'your name is Qwen' and it saves that to memory.
- Schedule work that's free to run
Examples shown: a 9:00 a.m. daily research report (created as a cron job), YouTube transcript to newsletter, and a small to-do app. Because tokens cost nothing locally, he suggests hourly jobs that would be too expensive on a cloud model.
Gotchas
- Sponsored by NVIDIA (he says he bought the Spark himself months earlier).
- Running a model 24/7 costs electricity; the savings are in per-token and subscription fees.
- The same local model can also back Claude Code, Codex or OpenCode through their custom-endpoint options.





