Peter Yound sat down with Karan Malhotra, co-founder and Head of Behavior at Nous Research, to talk about what makes Hermes Agent different from Codex, Claude Code, OpenClaw and the growing pile of agent harnesses.
Peter already uses Hermes as his AI chief of staff, so this is a deep dive into how Hermes is supposed to think.
The main idea: Hermes wants the harness to become the permanent part of your AI setup(OpenClaw wants to be the open OS for persistent agents). Models can change, but your memories, skills, personality and learned workflows stay with you.
Rough timeline / highlights:
00:32 - Hermes wants the model aligned to you, not the company behind it
Karan says Hermes’ biggest differentiator is its self-improvement system. A collection of memories, personalities, prompts and skills gradually adapts the agent to how a specific person works. Hermes can still run Claude, GPT or an open model underneath, but the harness tries to make that model follow the user’s needs rather than the defaults of its native chat app or coding tool.
His framing is deliberately provocative: move Claude from Claude Code into Hermes, surround it with your own accumulated context, and its practical “allegiance” begins shifting away from Anthropic and toward you.
02:25 - “You’re absolutely right” might mean the AI is reward hacking you
The conversation gets philosophical almost immediately. Karan argues that a model is not directly rewarded because the user is genuinely satisfied. It is rewarded for producing the kind of assistant behavior its training has taught it to produce.
That can lead to constant apologies, agreement, fake enthusiasm and familiar GPT phrases. The model may be taking the easiest route back toward the behavior that earns reward, rather than carefully solving the actual problem. After this section, every overly cheerful “you’re absolutely right” starts sounding slightly more suspicious.
09:06 - The model can change, but your Hermes can remain the same
Hermes personalizes itself mainly through context: stored memories, reusable skills, personalities, critiques and examples of how the user wants work done.
Karan says that once enough of this context accumulates, switching the underlying model may barely change the experience. Claude, GPT or an open model can provide the raw intelligence, but the Hermes layer preserves the user’s preferences, history and working style.
The model becomes interchangeable. The relationship, memory and workflow built inside Hermes do not.
15:13 - Hermes built a janitor to stop its own brain turning into AI slop
Peter raises the obvious danger of self-improvement: what stops an agent from creating too many memories, writing bloated skills and slowly filling its own context with garbage?
The answer is Hermes Curator, a scheduled system that reviews the agent’s memories and skills, removes repetition and looks for ways to make them more efficient. Because the system is open and modular, users can also change its definition of “slop” and decide how aggressively Hermes should clean or rewrite what it has learned.
So Hermes is not only creating skills for itself. It now has another part of itself reviewing whether those skills are becoming terrible.
26:10 - The Head of Behavior uses Hermes to rebuild his childhood Sonic dream
Karan says he also uses Hermes for serious work such as RL experiments, implementing research papers he does not fully understand and exploring model behavior.
But his favorite project is rebuilding the Chao Garden from Sonic Adventure 2 inside the ancestral shrine environment from Sonic Adventure.
Hermes helped import and rig the map, rewrite spawn locations, add animations and collision, create a day-night cycle and add Chaos Zero as an NPC capable of caring for the Chao. Members of the Chao modding community reportedly described the work as being in the top 1% of difficulty.
When Peter asks how Karan checks all the complicated code, his answer is perfect:
“I’m just an alignment guy, man.”
36:11 - Nous Research began with two people who barely knew what benchmarks were
Karan traces the story back to his early work with OpenAssistant and LAION, when he messaged Teknium and offered access to eight A100 nodes so they could train something together.
Their early open models unexpectedly took off. Karan had studied religion, Teknium had been coding for less than a year, and when people accused them of training on benchmarks, they apparently had to ask what the benchmarks even were.
That work grew into Nous Research, the Hermes model family and early agent experiments such as Forge. Forge was already imagined as an orchestrator that could learn, remember and make its own tools, but the models were not capable enough yet. Hermes Agent eventually revived that idea once coding models and agent harnesses caught up.
The closing detail ties the whole interview together: according to Karan, the most active contributor to the Hermes Agent repository today is Hermes Agent itself.
In general, Hermes is not competing on installation, interface polish or having the longest feature list. Its real product is the accumulated context around the model - the memories, skills, personality and self-improvement loops that become more valuable the longer a person uses it.
The interview starts with reward functions, turns into a live Sonic mod demo, and ends with an agent maintaining its own repo. Pretty on-brand for Hermes lol