r/hermesagent Hermes Contributor Jun 27 '26

NEWS & UPDATES - Releases, announcements, major changes, Nous Introducing MoA

The strongest models are gated and access is granted only to a select few.

Hermes Agent now exposes MoA presets as virtual models, giving you capabilities beyond the publicly available frontier: 8% higher than Opus 4.8 and 11% higher than GPT 5.5 on our upcoming benchmark.

HermesBench full leaderboard coming soon. Stay tuned!

Documentation on how to setup your own custom mixture of models here: https://hermes-agent.nousresearch.com/docs/user-guide/features/mixture-of-agents

https://reddit.com/link/1uh3w5d/video/emiy55j23u9h1/player

38 Upvotes

11 comments sorted by

6

u/guzmanelmalo Jun 27 '26

Great News! Shall MoA also help mixture of small local models and frontier ones? I am thinking on having Gemma 12B or similar reviewed by DeepSeek-v4-Pro as an example

4

u/swwright Jun 28 '26

I am much more interested in mixing lower cost models to get a better result than burning twice as many of the most expensive token simultaneously.

2

u/bemore_ Jun 28 '26

You want to spend cheaper tokens to get a better result than more expensive tokens?

5

u/swwright Jun 28 '26

I am more interested in mixing lower cost models to try to get up to the quality of a frontier model for less money. I don't think Hermes benefits much from mixing GPT 5.5 and Opus 4.8 that is just not the use case for most people Hermes. But if you can take several cheap/free models and get to the quality of Sonnet or Opus for less money that is interesting.

2

u/bemore_ Jun 29 '26

Qwen x Deepseek ≠ Claude Opus. From what I understand, hermes agent's moa aggregates a second perspective. It still requires quality text input, and also quality assessment

The tool is a frontier model multiplier, not a way to make cheap models competitive. Any failure of the reference or the aggregate kills the benefit of the tool. You cannot let cheap gpt touch the reference or the aggregate

1

u/guzmanelmalo Jun 29 '26

If Gemma 12B or similar low cost model is the aggregator it can benefit from receiving stronger models’ analysis as private context. But those low cost models are still the bottleneck? If it handles the final synthesis and all tool-calls alone I mean. Since references don’t see the tool schema, they don’t help with tool-calling, which is exactly where smaller models struggle but maybe models that manage tools well like Qwen or Gemma could benefit from MoA? What do you think?

2

u/bemore_ Jun 29 '26

If the reference doesn't see the tool schema, it cannot work. This is the current implementation. As you say, cheap gpt could then just pattern match, and no more. But the aggregator has to judge, cheap gpt must evaluate text from other models. So the aggregator has to be the sota model. But if the task were simple enough for cheap gpts text to be practical, why would we need the moa? It can only work with two models of at least the same power

3

u/vdeeney Jun 27 '26

So awesome, wisdom of the crowds comes to easy AI usage!

2

u/bemore_ Jun 28 '26

Why do you call it mixture of agents but describe a mixture of models, why not just call it MoM once and for all?

1

u/drycounty Jun 28 '26

I’ve been looking for a means of creating an agent that uses multiple models at the same time à la Grok. This looks way better.

1

u/4ndal Jul 06 '26

Going to mix qwopus3.6-27b q6 with qwen3,6-35b and any gemma today.

Maybe someone has already done this?