Deep dive
Long-horizon agents need experiments, not prompts: a five-step auto-research loop with scorecards
Erina Karati (ex-Microsoft, Supercell AI Innovation Lab) built Project Paradox, a framework of stateful game agents. It worked in short sessions and drifted over long ones. Her fix is a Karpathy-style auto-research loop around the agents. It carries over to any agent that keeps state, including memory-heavy personal agents like Hermes.
"Long-Horizon Agents Need Experiments, Not Just Prompts — Erina Karati" by AI Engineer — Watch on YouTube →
The architecture that worked short-term
- Per-agent memory namespaces (RAG-backed), so memories don't bleed between agents.
- An emotion vector (joy, sadness, fear, anger, disgust) updated after events.
- Belief/trust scores toward other agents and the player, raised or lowered by the LLM after each interaction.
- Importance scores on memories: anything above a threshold goes to a separate cache for better retrieval later.
Where it broke
Over long horizons, social consistency decayed. Agents kept the topic but lost the source, rumours hardened into stated facts, and agents knew a fact but didn't use it when planning. More memory didn't fix this. Agents need provenance (firsthand vs secondhand, verified vs uncertain), and raw episodic memory should be kept separate from current beliefs.
The recipe
- Freeze the harness, the scenarios and the metrics. The optimiser may not touch them.
- Define controlled scenarios, for example: public fact diffusion ("the bakery closes tomorrow" — who learns it, do they remember who said it?), rumour uncertainty (does "might leave" become "is leaving"?), replanning (a route gets blocked — do agents tell each other?).
- Log structured traces: observations, conversations, memory writes, retrievals, belief updates.
- Score with a balanced scorecard, never one "agent quality" number. Measure diffusion reach, source retention, uncertainty preservation and false-certainty rate, action consistency and time to replan, and privacy containment. Optimising one metric alone games it: diffusion alone teaches oversharing, and recall alone surfaces stale memories.
- Expose a small policy surface and search over it: memory-write policy, retrieval policy, communication prompt, trust rules, source attribution, replanning triggers. Keep a change only if the scorecard improves and the guardrails hold. Otherwise revert (a ratchet).
Examples of the kind of change the loop proposes: preserve source in memory writes and summaries; store confidence, mark firsthand vs secondhand, require hedging when retelling; classify useful public facts so agents proactively share them.
Why it matters beyond games
Support agents need to know which policy update supersedes which. Personal assistants need to keep, and correct, commitments. Research agents need provenance and contradiction handling. Coding agents need context across issues and changing requirements. They all carry state that affects future actions. She is careful about the claim: this is the right kind of surface to optimise, not proof that the system improved in general. Related reading: how Hermes memory works, context window.





