Skip to content
← all writing

Cutting Hermes Agent overhead with prompt-size and on-demand tools

  • hermes
  • tokens
  • prompt-engineering
  • optimization

A thread on X last week by IBuzovskyi laid out a practical Hermes Agent token-optimization path. The core observation is that Hermes sends a large fixed token payload on every turn: system prompt, skill descriptions, long-term memory, and tool/MCP schemas. That overhead is easy to ignore until you run hermes prompt-size and see 6,000–10,000+ tokens spent before your message even begins.

The thread's workflow is straightforward.

  1. Run hermes prompt-size to get a baseline.
  2. Disable unused skills and MCP servers.
  3. Clean USER.md and MEMORY.md, removing outdated context.
  4. Switch tool_search to on-demand mode.
  5. Re-run hermes prompt-size and compare.

The highest-impact change in that list is on-demand tool loading. Instead of exposing every tool schema every turn, the agent loads schemas only when needed.

tools:
  tool_search:
    enabled: auto

The thread claims most users can cut overhead by 40–70% after one cleanup session. That claim is testable: measure before/after with hermes prompt-size.

There are also secondary guidelines for writing token-efficient skills. Short, high-signal descriptions beat long philosophical intros. Fewer broad skills beat many narrow ones. Removing redundant examples from skill files reduces repeated schema weight.

This is not a novel architectural idea, but it is one of the more concrete Hermes-specific optimization guides circulating this month. The actionable part is the measurement step. If you don't know your baseline, you can't verify improvement.

[^1]: IBuzovskyi. "Hermes Agent Token Optimization (Prompt + Fixed Overhead)" X post. September 2026.

Termagotchi
_

Ryan Underdown

Autodidact. Rarely listens to advice.

Follow on X @catamarammed or GitHub @underdown