Hermes Agent Replay Compaction Cuts Re-sent Tool Output by up to 80%
- hermes
- context-optimization
- token-savings
- replay-compaction
- compression

Every turn, an AI agent re-sends its entire conversation history to the model. That history includes every tool result from every earlier turn: terminal output, file reads, search hits, large JSON blobs. On a long agentic session these replayed results dominate the request payload. A single 100K-token tool result, fetched once, is paid for again on every subsequent turn until the context window is compacted.
An open pull request, hermes-agent#87303, attacks exactly this cost with a send-path replay-compaction layer. It runs before the request leaves the agent, independent of the existing summarizer, and shrinks oversized tool results and redundant reasoning chains before they are re-sent. The author, Rene Leonhardt, reports 80.0% of replayed input removed (a 5.0x reduction) on DeepSeek and MiMo, up to 120x on individual oversized tool results, across all 10 model providers tested.
How the replay economy works
The compaction operates on the send copy only. Raw history stays untouched in the session store, so session search and compaction summaries keep working. Two mechanisms do the work:
- Tool-result compaction. Any tool result above a threshold (over 2K tokens or 12 KB) is collapsed to a head+tail+marker summary of roughly 835 tokens before the wire is built. The model still sees the start and end of the result plus a marker; the middle is elided.
- Reasoning strip. On providers where reasoning chains are proven safe to drop (DeepSeek, MiMo), the thinking blocks are stripped from the replayed history. The PR pins a keep/strip matrix per provider: Qwen reasoning is explicitly preserved (
preserve_thinking=true), Kimi is hard must-keep, DeepSeek and MiMo are stripped.
A key design point: this fires even when compression.enabled: false. The existing summarizer still triggers at the true wire size, but a new preflight measures the post-replay wire so compression no longer fires early on a raw overstatement. The PR reports the preflight corrects the estimate by 2.4x to 16x on non-echo wires.
The numbers
Benchmarks use a fixed session of 8 tool results plus 6 reasoning turns, measured per provider wire.
| Wire | Saved | Reduction | Compacted blocks | Reasoning stripped |
|---|---|---|---|---|
| Anthropic | 58.8% | 2.43x | 8 | 0 |
| OpenAI | 58.8% | 2.43x | 8 | 0 |
| DeepSeek | 80.0% | 5.00x | 8 | 6 |
| Moonshot AI | 58.8% | 2.43x | 8 | 0 |
| Z.ai | 58.8% | 2.43x | 8 | 0 |
| MiMo | 80.0% | 5.00x | 8 | 6 |
The session-wide figure: 56,672 tokens of replayed input drops to 11,324 on DeepSeek and MiMo (80.0%, a 5.00x cut). On Anthropic, OpenAI, Moonshot AI, and Z.ai the same session drops to 23,360 tokens (58.8%, 2.43x).
Per-result compaction scales with the size of the elided middle. The PR documents the curve:
| Original tool result | Compacted | Reduction |
|---|---|---|
| 8K tokens | ~835 tok | ~9x |
| 12K tokens | ~835 tok | ~14x |
| 50K tokens | ~835 tok | ~60x |
| 100K tokens | ~835 tok | ~120x |
A no-LLM compression checkpoint
The same PR adds an opt-in compression.checkpoint_mode that replaces the LLM summarizer on tool-heavy sessions. Instead of paying for a model call to summarize the middle of a long session, the agent writes a deterministic checkpoint: goal, tool evidence, and next action. The raw middle stays in the session store and remains searchable. Conversational middles keep the LLM summary. The PR estimates 100K to 1.2M raw tokens saved per summarizer call that the checkpoint replaces, with a checkpoint_tool_ratio default of 0.7 gating when it applies.
Why it matters
The dominant cost in long agentic runs is not the model's answer. It is the repeated re-transmission of tool output the agent already consumed. Send-path replay compaction attacks that directly: it does not wait for the context window to fill, it works whether or not summarization is enabled, and it is provider-aware so each vendor's reasoning and overflow quirks are handled explicitly (for example, Qwen endpoints get context_overflow_policy: stopAtLimit injected so the provider rejects at its limit rather than silently truncating the compacted wire).
The PR is currently open, opened August 15, 2026, and includes a reproducible benchmark (scripts/bench_replay_economy.py) plus per-wire CI comparisons against the merge base at a 5% tolerance. The numbers above are the author's reported bench results for a fixed session and have not yet landed in a tagged release.
Source: NousResearch/hermes-agent#87303 by Rene Leonhardt.