Hermes Agent Cuts Tool-Definition Tokens by Half With Progressive Disclosure
- hermes
- mcp
- tokens
- optimization
Large MCP tool catalogs are expensive. Each server's full JSON schemas ship on every API call inside the system prompt. Anthropic's engineering data puts tool-definition overhead between 15,000 and 60,000 tokens per turn for typical multi-server deployments. A 34-tool setup can produce ~45,000-token prompts, with roughly half consumed by schemas alone.
Hermes Agent now ships a built-in answer: progressive tool disclosure via tool_search. When active, MCP and non-core plugin tools are removed from the model-visible tools array and replaced with three bridge tools. The model discovers tools by name, loads full schemas on demand, then invokes them. Core Hermes tools like read_file, terminal, and search_files are never deferred.
How it works
The bridge exposes three tools:
tool_search(queries, limit?)— search the deferred catalogtool_describe(names)— load full schemas for specific toolstool_call(name, arguments)— invoke a deferred tool
Discovery uses BM25 lexical retrieval over tokenized tool names, source server names, descriptions, and parameter names, with Snowball stemming and a literal substring fallback. An optional embedding reranker using nomic-embed-text-v2-moe adds semantic matching on top.
Auto mode
The feature runs in auto mode by default. It activates only when deferrable tool schemas would exceed 10% of the active model's context window. Below that threshold, the tools array passes through unchanged and you pay zero overhead.
tools:
tool_search:
enabled: auto # auto (default), on, or off
threshold_pct: 10 # % of context at which auto kicks in
search_default_limit: 5
max_search_limit: 20
Token savings
Internal benchmarks from the Hermes repo measured two configurations against a baseline with heavy tool schemas.
| Configuration | Tool tokens per turn | Reduction |
|---|---|---|
| Arm A: eager loading | ~15,000 | — |
| Arm B: progressive disclosure | ~7,200–8,600 | ~43–52% |
| Savings | ~6,400–7,800 tok | ~43–52% |
A separate internal test with 72 registered tools on a 32K-context model showed a sharper reduction: tool-definition overhead fell from ~19,210 tokens to ~2,198 tokens, an 89% decrease.
Accuracy gains
The same Hermes benchmarks and Anthropic's published evals both show that fewer tool definitions can improve task accuracy. Large tool catalogs create decision paralysis; the model picks irrelevant options or hallucinates capabilities that are not present.
| Model | Without tool_search | With tool_search | Delta |
|---|---|---|---|
| Claude Opus 4 | 49% | 74% | +25 pts |
| Claude Opus 4.5 | 79.5% | 88.1% | +8.6 pts |
The accuracy improvement is not free. Live testing against Claude Haiku 4.5 showed extra round trips on cold tool access.
| Scenario | ON: elapsed | OFF: elapsed | Extra round trips |
|---|---|---|---|
| Single obvious tool | 18.5s | 9.7s | +3 |
| Multi-tool chain | 20.3s | 14.1s | +4 |
| Core plus deferred | 33.1s | 9.8s | +3 |
| Pure knowledge | 8.2s | 2.8s | 0 |
Single-tool tasks ran roughly twice as slow with the bridge active. Multi-tool chains were about 1.4x slower. Pure-knowledge prompts had zero extra round trips because no tool search was triggered.
Tiers and listing modes
Tool Search uses tiered disclosure. The catalog listing budget is min(threshold_pct% of context, listing_max_tokens). At tier 1, a skills-style manifest of deferred tool names and short descriptions stays visible so the model can often skip tool_search and go straight to tool_describe. At tier 2, even names collapse to per-server summaries for catalogs too large to list.
Core tools are never deferred. Subagents, kanban workers, and gateway sessions see only the tools they were granted; the catalog is scoped to the session's own enabled toolsets.
Practical notes
- The feature is already merged onto Hermes main and shipped.
- It costs roughly 300 extra tokens per turn for the bridge tools themselves.
- Deferred schemas loaded via
tool_describedo not benefit from system-prompt cache prefix, though they do enter conversation history and cache on subsequent turns. - Adding or removing tools mid-session invalidates the prompt cache, same as any toolset edit.
- Smaller models may need prompt guidance to write good search queries; the published accuracy numbers are with Opus-class models.
Sources
- Hermes Agent docs: Tool Search
- NousResearch/hermes-agent PR #31163
- NousResearch/hermes-agent PR #35276
- NousResearch/hermes-agent PR #35457
- MarkTechPost, Hermes Agent Ships Tool Search for MCP
- KaiDev, Hermes Tool Search Cuts MCP Overhead, Boosts Accuracy