SAN FRANCISCO – Nous Research, the open-source AI lab behind the Hermes Agent, has launched a leaderboard that ranks AI models by how well they perform inside its own agent. It is using the results to push model makers to build for Hermes.
Nous released the leaderboard, Hermes Index, along with a test suite called Hermes Bench, the day before it announced a $90 million fundraise. Sam Herring, Nous' machine-learning engineering lead, described both Wednesday on a panel at SF Tech Week in San Francisco. He said the project had taken him four months.
The index tells AI developers, "Here's how your model does in this harness," Herring said. "And then it's on them to sort of improve their models specifically targeting Hermes."
Most AI benchmarks grade a model on its own. Hermes Index turns that around. It holds the agent constant and grades the models plugged into it. That gives the maker of a widely used open-source agent a voice in what the big model builders optimize for, a lever usually held by the labs themselves.
See also: Hermes maker Nous Research raises $90M in breakout for DeAI
How it works
A harness is the software wrapped around an AI model that lets it act as an agent: calling tools, keeping memory, running multi-step tasks. Hermes Index keeps Hermes as the harness and runs a range of models through Nous' inference portal on the tasks people most often use Hermes for, Herring said.
Nous then takes the scores back to the companies that build the models.
Nous also runs its own reinforcement-learning training with Hermes Agent in the loop, he said.
The index is not independent. Nous builds the harness it holds fixed. It also sells access to models through Nous Portal, which offers more than 200 models on paid plans from $20 to $200 a month.
Why the harness matters
Akshay, who leads inference benchmarking at Artificial Analysis, said on the same panel that the choice of harness has a large effect on results. In his framing, the model sets the ceiling and the harness decides how much of it gets used.
Herring said nobody can put an exact number on the split between model and harness.
He called Artificial Analysis a pretty definitive source for model and agent benchmarks. He said Hermes Index aims to fill a gap. Many agent benchmarks focus on coding, he said, while much of what people do with agents every day is less objective and harder to score with a clear right answer.
An audience member who works on a rival agent benchmark thanked Herring for launching the index. She said picking a model to run with Hermes is always a hard call.
"Hermes is for the people"
Herring pitched Hermes as the open alternative to closed personal agents. He compared closed products to a MacBook, which runs without the user thinking about the internals.
"Hermes is like the Linux machine," he said. Users can see inside it, change it, fork it and choose where the model runs.
"I don't think a lot of people realize how important it is to be able to own your data in this day and age," Herring said. "Hermes is open source. Hermes is for the people."
Nous released Hermes Agent under an MIT license in February. CEO Dillon Rolnick wrote that it has since been cloned more than 24 million times. By Nous' internal estimates, it drives about 2.5% of global AI token usage.
Herring wouldn't share the roadmap. "Expect to see Hermes Agent a lot more places," he said.
HOW AI WAS USED IN THE PRODUCTION OF THIS PIECE: I drafted this story with Claude, following the structure of Distro's blog-post skill and the Distro house-style rules. It worked from a transcript of the SF Tech Week panel I attended, Nous Research's fundraising blog post and the Hermes Agent website. I checked every quote against the panel audio before publication.