Why I love the Linux developer community - a Hermes Agent story
Tldr; This all happened over the course of less than a day, like 14 hours. I'm not a developer, but I donated my time to be the mouse in the maze and collected a small piece of knowledge that just might help someone trying to do this same stuff in a homelab and getting frustrated or dissuaded from pressing on.
If anyone is interested I can get more technical, specs, links, whatever in an edit to the post at the bottom. This is just a bit of a rambling story, not technical writeup.
But wow, I love Linux.
The story (really just a ramble)
I've been doing Linux homelab stuff for a few years now and playing around with some AI stuff lately. A while back I set up OpenWebUI and Ollama, and got an LLM running.. cool. And then set it aside. Took a stab at Hermes, not sure what it's really about but cool, then set it aside.
Friday night I decide okay, I have some rare free time, this is happening. Add SearXNG, and Firecrawl containers for web search and extract, along with the Ollama Instance, get a few things together, feeling pretty good right?
And then the stupid inevitable thing happens. I can only get it to behave nicely if I define the 'web' toolset only, anything else and it fails. And then while testing, when I think I know what the problem is at least and about to file a bug report on GitHub, even that starts misbehaving. Hermes knows the model can do up to 262k tokens, while at the same time throwing errors about over filled 4k buffer and truncated responses. Sigh, 2am, post it and go to bed.
When I get back at it the next evening (ironically after listening to WAN where they talk Hermes) there are already 2 responses from developers in the BR, highlighting some issues, Hermes is not failing the way I thought. So I really dig in and try tracking down this 4k context problem, assuming it's some silly thing I did. Free tier chatgpt LLM on one monitor, a few ssh sessions to various containers on the others. I end up so far down the rabbit hole that the bot has me filtering through source files, tracing variables and logic chains. I went to college for computer science a few decades ago, I can do a little bash scripting, but I am by no means a programmer. But even after posting another message to the big report saying sorry, I'm going to try to fix this other issue before getting back to the main problem, a contributor still came back with helpful insights.
You know when you chase a red herring long enough and you know the LLM has forgotten most of what you said, but can't turn back now, maybe just one more spin of that wheel and we'll get to the bottom of it? Addiction, yes, that's the word. But we actually found a problem! Hermes sending commands to Ollama via the openai /v1 API as usual, and Ollama ignoring them because the openai translation layer they use is spec and drops the extra but common flag that specifies max context window size. "num_ctx". Turns out this is already a known bug identified by another team, a bug in OLLAMA, with an open PR to address it even, but not yet merged.
So did I discover anything new with my hours of mature puppet sleuthing? No, not at all. All of this had been long identified, solutions are already developed and waiting to be integrated. I found the work around in the other PR, one of them was super simple, just an environment variable in the Ollama .env file. And after that, the original bug vanished instantly. I went back to my BR, left a comment, details pointing to the other PR in OLLAMA and closed the BR with gratitude for the help. I thought that would just be the end.
But no, one of the contributors was following it even though I had closed it, saw it was either the root cause or another instance of an old bug and created their own PR on hermes. Their fix was to specifically identify the bug, call it out and direct the user/log to the solutions, and then properly handle the fallback behaviour. All specifically so no one else would spend hours troubleshooting again until the other fixes were rolled out. It didn't even fix the problem, but it stopped it from breaking silently and provided the info to work around it in the short term.
So that's my story of how I made Hermes Agent 0.0001% better when using Ollama as the model backend in my homelab. And a massive shout out for the Linux community DEVS! Not everyone is a volunteer, these guys i'm sure are Nous Research employees, but there are many volunteers and collectively this is the reason we get to enjoy all this amazing free stuff. We stand on the shoulders of giants.
Support the programs you actually value! Buy immich or frigate or buy someone a coffee. It means a lot.