Skip to content
HomeNewsHermes Index: a score and a cost per task for 14 models in Hermes Agent

Subscribe to our newsletter

One short email when we publish. No spam, unsubscribe anytime.

Hermes sits as judge before a row of marble busts with a bronze scale weighing gold coins against a laurel wreath

Hermes Index: a score and a cost per task for 14 models in Hermes Agent

By OneClickClaw EditorialOctober 8, 2026

The most expensive of the 14 models in the new Hermes Index costs 215 times as much per task as the cheapest, according to the index page on the Nous Research portal, read on the 8th of October. The cheapest task comes to about five cents. The most expensive comes to $11.61 (€10.39).

Nous Research, the company that makes Hermes Agent, announced the index on X on the 6th of October. In the post the company wrote that the index averages several benchmarks of how models perform inside Hermes Agent, that it exists to inform people choosing between models, and that it shows the model makers how their models do inside Hermes. The newsletter AlphaSignal reported the index the same day.

Hermes Agent is the open source agent software from Nous Research, documented on its own site. An agent is a program that takes an AI model and lets it carry out work in steps: read files, run tools, search the web, remember what it learned. The model is the part the user picks, and the index puts a price on that pick.

What Nous Research measured

Each model ran four sets of tasks inside Hermes Agent, the index page says. One set is Hermes Bench, built by Nous Research itself: 150 tasks in 25 categories, covering Hermes skills, research, diagrams and art, memory, tool use and safety. The other three are outside test sets: TerminalBench 4, run without its four tasks that need a graphics card, TerminalBench Science and SkillsBench.

A task is one job handed to the agent. The agent starts in a workspace with real files, the page says, and some tasks come with follow-up turns, the way a person adds a second request after the first answer.

The score is the mean reward across the tasks, so partial credit counts, the page says. In plain words: a task done in full earns full points, a task half done earns half, and the score is the average over every task the model received. The Hermes Index figure is the mean score across the four sets, and the cost per task is the mean amount in US dollars that one task cost in model usage across the same four sets. Makers bill by the token, a piece of a word, so a model that reads and writes more on a task costs more on it.

The 14 models, cheapest first

The index page sorts the table by score. The table below sorts it by cost, cheapest first. Scores and dollar costs are Nous Research's own figures from the index page. Euro figures use the European Central Bank reference rate of the 7th of October, 1 euro = 1.1177 US dollars, rounded to the cent.

Model Maker Nous Research score Cost per task
Ling 3.0 Flash inclusionAI 21.56 $0.054 (€0.05)
GPT 6 Luna OpenAI 33.89 $0.141 (€0.13)
DeepSeek V4.1 Flash DeepSeek 36.91 $0.259 (€0.23)
Hy4 Tencent 26.13 $0.482 (€0.43)
GLM 5.3 Flash Z.ai 34.95 $0.497 (€0.44)
GPT 6 Sol OpenAI 44.10 $2.23 (€2.00)
Claude Sonnet 5.5 Anthropic 53.14 $2.82 (€2.52)
Gemini Flash 3.8 Google 38.12 $3.17 (€2.84)
GLM 5.3 Z.ai 39.23 $3.78 (€3.38)
Claude Opus 5.5 Anthropic 63.31 $4.99 (€4.46)
Kimi K3 Moonshot 39.16 $6.50 (€5.82)
Qwen 3.8 Max Alibaba 36.17 $6.69 (€5.99)
Grok 4.7 xAI 39.32 $10.77 (€9.64)
GPT 6 Astra OpenAI 56.25 $11.61 (€10.39)

Where the costs fall

Pie chart of the 14 models in the Hermes Index by cost per task: 5 under 0.50 dollars, 5 between 2 and 5 dollars, 4 over 5 dollars

The chart counts the 14 models from the same page by cost per task. Two of the four in the top group cost more than ten dollars.

The score column and the cost column do not run in the same order. Sorted by price, the page gives one sequence; sorted by score, another. The page attaches no verdict to either column.

What a month would cost

Hermes counts on an abacus under a frieze of moon phases while stacks of gold coins grow taller beside him

The figures below are arithmetic on Nous Research's per-task costs, for one example: an agent doing 20 such tasks a day for 30 days, which is 600 tasks in the month.

At the cheapest per-task figure in the table, the month comes to $32.40 (€28.99). At the figure in the seventh row of the table, the month comes to $1,692.00 (€1,513.82). At the most expensive figure, the month comes to $6,966.00 (€6,232.44).

The tasks in the index start from a workspace with real files and some carry follow-up turns, the page says, so a user's own tasks may run shorter or longer. The arithmetic shows the spread between models on the same work. A user's own bill follows the user's own tasks.

An agent that remembers needs a machine that stays on

Memory is one of the 25 categories in Hermes Bench, and follow-up turns are part of the test, the index page says. Hermes Agent keeps its conversations, memories and skills from one session to the next, its documentation says, so the person who comes back tomorrow expects the agent to be where they left it, with the files and the memory in place.

OneClickClaw hosts Hermes Agent and OpenClaw on an always-on dedicated server in the EU, at the same price for both, according to its plans page. The customer picks the model on day one and brings a key from that model's maker, so the cost per task in the table above lands on the maker's bill and the server fee stays flat whichever row the customer chooses. The 7-day free trial needs no credit card, and plans start at EUR 14.99 per month (Starter).

Helpful documentation

Want your own Hermes agent running without the setup? We host it on a dedicated EU server. Free for 7 days, no credit card.

Start your free trial

How managed Hermes Agent hosting works

Get notified when we publish new articles

No spam, unsubscribe anytime.

By subscribing, you agree to our Privacy Policy.