Post

AI

In my rerun, Headroom saved 0 tokens on its GSM8K eval prompts

A Source Audit of Headroom v0.39.1: its seeded savings table reproduces, but in my rerun its own GSM8K and TruthfulQA eval requests passed through unchanged.

Original bar chart: three benchmark-shaped requests keep the same token count through the default Headroom proxy, while a JSON tool-result control drops from 4,844 to 2,863 tokens
Original bar chart: three benchmark-shaped requests keep the same token count through the default Headroom proxy, while a JSON tool-result control drops from 4,844 to 2,863 tokens

Editorial method: This Source Audit was researched and drafted with AI assistance under a policy-bound evidence harness; the reruns used a local stub instead of OpenAI, and I have not run Headroom on production agent traffic.

🤔 Curiosity: What does Headroom’s accuracy table actually test?

Headroom calls itself the context compression layer for AI agents, and it ships as a library, a proxy, and an MCP server. It is an Apache-2.0 repository created on January 7, 2026, and it had 74,273 GitHub stars when I pulled its metadata for this audit. Its GitHub description promises 20% fewer tokens for coding agents, 60-95% fewer for JSON, and “same answers.” In my TypeSafe routing audit Headroom appeared only as an appendix tool, and I wrote that its savings claims had not been rerun. This post does the rerun.

A compressor makes two promises. The first is that it removes tokens. The second, and the one that matters before you put it in front of an agent, is that removing them does not change the answers. The README backs the second promise with an accuracy table headed by one command, python -m headroom.evals suite --tier 1, and four rows: GSM8K, TruthfulQA, SQuAD v2, and BFCL.

So my question was narrow:

When Headroom reports “no accuracy loss” on GSM8K, how many tokens did it remove from those prompts?

The answer from my rerun is zero. The token-savings table reproduces to the last token. But when I replayed the GSM8K and TruthfulQA requests that Headroom’s own eval command generates, the default proxy forwarded them unchanged.

📚 Retrieve: What the pinned repository and three reruns show

I pinned the audit to release tag v0.39.1, commit d13e1966, committed on September 26, 2026. Its pyproject.toml declares version 0.39.1, the same version as the headroom-ai 0.39.1 wheel I installed from PyPI. I read the README, the evaluation package, the evaluation workflow, the Benchmarks docs page, and the wiki at that tag, then ran three local experiments on macOS (Apple silicon) with Python 3.11.

Rerun 1: the savings table reproduces exactly

The README’s Proof section says four scenarios were built from real MCP server output formats and measured with the provider tokenizer and the shipped compress(). It adds that the run is seeded and offline, “so you get the same numbers we did”, and gives the command index_proof_table.py --seed 20260902.

I ran that script from the tag against the installed wheel. Every number matched the README and the committed results file:

ScenarioBeforeAfterSaved
Code search (100 results)17,19913,59721%
SRE incident debugging55,95724,34057%
Codebase exploration58,80133,89542%
GitHub issue triage46,06732,42930%
Total178,024104,26141%

That is a good result for the project. The scenarios are synthetic generators, so this proves reproducibility of the token arithmetic, not answer quality. The README is also candid about where savings come from: it says savings scale with how repetitive the payload is, and that prose and already-dense output compress very little. Keep that sentence in mind, because it explains everything that follows.

Headroom v0.3.0 dashboard session showing 143.0k tokens saved, a Compression Quality card reading High, and a What Headroom Removed panel split into JSON Bloat and Repetition
The dashboard attributes removed tokens to JSON bloat and repetition, and its Compression Quality card says 100% of removed tokens were identified waste; it is a waste measure, not an answer check. Source: https://github.com/headroomlabs-ai/headroom/blob/d13e1966f820220b482a33c30bde1e926743939a/wiki/screenshots/cache-ttl-dashboard-live.png. Publisher/creator: headroomlabs-ai/headroom contributors. License: https://raw.githubusercontent.com/headroomlabs-ai/headroom/d13e1966f820220b482a33c30bde1e926743939a/LICENSE. Attribution: headroomlabs-ai/headroom contributors, wiki/screenshots/cache-ttl-dashboard-live.png, Apache License 2.0, pinned at tag v0.39.1 (d13e1966).

The accuracy table has two kinds of rows

The README’s accuracy table reads:

BenchmarkNBaselineHeadroomDelta
GSM8K1000.8700.870±0.000
TruthfulQA1000.5300.560+0.030
SQuAD v2100—97%at 19% compression
BFCL100—97%at 32% compression

Two rows carry a compression figure. GSM8K and TruthfulQA do not. The evals README goes further and places those two rows under the heading Standard Benchmarks — “No Accuracy Loss”, with the model stated as gpt-4o-mini. To its credit, the main README says the TruthfulQA delta falls inside the confidence interval at N=100, so it does not claim an improvement.

The code shows why the columns differ. The tier-1 suite in suite_runner.py defines nine benchmarks: GSM8K, TruthfulQA, MMLU, ARC-Challenge, and HumanEval go through an lm_eval runner; SQuAD v2, BFCL, and Tool Outputs go through a before_after runner; and CCR Round-trip is compression-only. The README shows four of the nine. The lm_eval path returns only a baseline score, a Headroom score, a delta, a pass flag, and a sample count. The before_after path also returns avg_compression_ratio and tokens saved. In other words, the suite records no compression measurement for GSM8K or TruthfulQA at all. Even the evals README’s own tier table describes SQuAD v2 as “reading comprehension with compression” and BFCL as “function calling with compressed schemas”, while GSM8K is just “math reasoning accuracy”.

The pass rule for the lm_eval rows is that the Headroom score is within 0.02 of baseline. The Headroom side runs lm-eval with --apply_chat_template as a local-chat-completions model pointed at the proxy’s /v1/chat/completions, and when the suite starts its own proxy it launches python -m headroom.proxy.server --port with no profile arguments.

That raised a testable question: what does that default proxy do to those requests?

Rerun 2: lm-eval with Headroom’s own flags, through the default proxy

I did not want to spend an API key or guess at prompt shapes, so I put a small recording stub where OpenAI would be. Then I ran lm-eval 0.4.13 twice with Headroom’s flags, once straight to the stub and once through a default Headroom 0.39.1 proxy:

1
2
3
4
5
6
# Same flags Headroom's comprehensive_benchmark.py passes to lm-eval.
# BASE_URL is either the stub directly or the Headroom proxy in front of it.
# Run once with --tasks gsm8k and once with --tasks truthfulqa_gen.
python -m lm_eval --model local-chat-completions --tasks gsm8k \
  --batch_size 1 --apply_chat_template --log_samples --limit 5 \
  --model_args "model=gpt-4o-mini,tokenizer_backend=tiktoken,base_url=$BASE_URL"

All 10 of 10 requests (five GSM8K, five TruthfulQA) reached the upstream with message arrays byte-identical to the direct run. The GSM8K requests were 11-message, five-shot, multi-turn chats; the TruthfulQA requests were single user messages. The only change the proxy made to the body was renaming max_tokens to max_completion_tokens.

Rerun 3: the proxy’s own counters, with a positive control

Identical bytes could also mean my setup was broken, so I replayed real GSM8K rows from Hugging Face as one long user message and as a multi-turn chat, a real TruthfulQA row as one user message, and a control the README says Headroom is good at: a tool result holding a 100-item JSON array. I read the proxy’s x-headroom-* response headers.

Original bar chart: three benchmark-shaped requests keep the same token count through the default Headroom proxy, while a JSON tool-result control drops from 4,844 to 2,863 tokens

The GSM8K requests reported 643 and 679 tokens before and after, with 0 saved. The TruthfulQA request reported 59 before and after. The control went from 4,844 to 2,863 tokens, 1,981 saved, through router:smart_crusher:0.65. The compressor works. It just had nothing to do on these prompts, which is exactly what “prose compresses very little” predicts.

Headroom v0.6.0 dashboard after 42 processed requests showing 0 tokens saved and Compression Quality reading No compression yet
Headroom's own docs screenshot of a proxy that processed 42 requests and saved 0 tokens. Passing traffic through the proxy is not the same as compressing it. Source: https://github.com/headroomlabs-ai/headroom/blob/d13e1966f820220b482a33c30bde1e926743939a/docs/screenshots/subscription_window_inactive.png. Publisher/creator: headroomlabs-ai/headroom contributors. License: https://raw.githubusercontent.com/headroomlabs-ai/headroom/d13e1966f820220b482a33c30bde1e926743939a/LICENSE. Attribution: headroomlabs-ai/headroom contributors, docs/screenshots/subscription_window_inactive.png, Apache License 2.0, pinned at tag v0.39.1 (d13e1966).

Headroom already has the right rule; it just is not applied to the README

The most useful line in the repository is in .github/workflows/eval.yml. Its HotpotQA recall step warns that when compression did not engage (about 0%), recall is not a meaningful fidelity signal. That is precisely the condition my reruns found for GSM8K and TruthfulQA.

The same workflow shows that the public repository’s default CI does not regenerate the table while that secret is unset. A comment says OPENAI_API_KEY is intentionally not set in the public OSS repo, and the weekly Tier 1 job writes a skipped report and exits when the key is absent. The Benchmarks page that the README links as “Methodology” contains no mention of GSM8K or TruthfulQA, even though it states that every number on it is measured locally and reproducible. It also says plainly that no committed result artifact exists for its LLM-judged QA comparison.

One more drift is worth knowing about. At the same tag, wiki/index.md still shows an older table for the same four scenario names with different figures, for example code search at 17,765 to 1,408 tokens (92%) instead of 21%, and it puts “0.000 delta” for GSM8K in a column labelled Compression. If you cite Headroom numbers, cite the README Proof table, which is the one that reproduces.

Headroom v0.6.0 dashboard showing 3.6M tokens saved at 41.7 percent and a Compression Quality card reading No waste signals detected
In this screenshot, savings are reported in tokens and dollars, and the Compression Quality card reports waste signals, not answers. Source: https://github.com/headroomlabs-ai/headroom/blob/d13e1966f820220b482a33c30bde1e926743939a/docs/screenshots/subscription_window_active.png. Publisher/creator: headroomlabs-ai/headroom contributors. License: https://raw.githubusercontent.com/headroomlabs-ai/headroom/d13e1966f820220b482a33c30bde1e926743939a/LICENSE. Attribution: headroomlabs-ai/headroom contributors, docs/screenshots/subscription_window_active.png, Apache License 2.0, pinned at tag v0.39.1 (d13e1966).

💡 Innovation: How to read a compression benchmark

My inference from the three reruns is this: in the configuration I reproduced, the GSM8K and TruthfulQA rows compare two runs that send the model the same messages. Their deltas therefore measure the run-to-run variation of the upstream model, not the effect of compression on answers. That is consistent with the README’s own reading of the TruthfulQA +0.030 as noise. It does not mean the published run was misreported. I do not know the model version, lm-eval version, proxy profile, or environment behind the published table, and a different setup could compress these prompts.

The practical rule I would take away is Headroom’s own, applied everywhere:

Row typeExampleWhat it showsUse it to decide?
Accuracy with a compression ratioSQuAD v2, BFCLAnswers held while context shrankYes, if the payload resembles yours
Accuracy without a compression ratioGSM8K, TruthfulQAScores through the proxy, with no record of how much was compressedNo: not evidence of compression fidelity
Token savings, seededProof tableArithmetic on synthetic scenariosYes for token math, not for answers
Project counterscommunity savingsVolume of tokens the fleet removedNo, it is unverified
Headroom community savings screenshot reading 59.0B tokens saved, with cost saved, requests optimized, active instances, active days, and a 48-hour tokens-saved chart
The project-reported community counter (59.0B tokens saved in this screenshot) measures tokens removed across instances; I did not verify it, and it says nothing about answers. Source: https://github.com/headroomlabs-ai/headroom/blob/d13e1966f820220b482a33c30bde1e926743939a/headroom-savings.png. Publisher/creator: headroomlabs-ai/headroom contributors. License: https://raw.githubusercontent.com/headroomlabs-ai/headroom/d13e1966f820220b482a33c30bde1e926743939a/LICENSE. Attribution: headroomlabs-ai/headroom contributors, headroom-savings.png, Apache License 2.0, pinned at tag v0.39.1 (d13e1966).

For a team evaluating Headroom on agent traffic, I would do three things before trusting it:

  1. Measure compression on your own payloads first. The README points readers to headroom savings for their own traffic. If your context is mostly prose, expect little; if it is search results, logs, or JSON tool output, the control above shows real reductions.
  2. Pair every accuracy number with its compression ratio. A fidelity check where nothing was compressed is not a fidelity result; Headroom’s own HotpotQA step says the same about ~0% compression.
  3. Keep the per-request evidence. The proxy already emits x-headroom-tokens-before, -after, and -saved headers, and the history view exports JSON and CSV. Logging those alongside task outcomes is a lightweight first step.
Headroom Historical Proxy Compression view with 143.0k lifetime tokens saved, two recorded checkpoints, and Export JSON and Export CSV buttons
The history view exports savings as JSON or CSV, which makes it easy to join token data with your own task-success logs. Source: https://github.com/headroomlabs-ai/headroom/blob/d13e1966f820220b482a33c30bde1e926743939a/wiki/screenshots/cache-ttl-dashboard-history.png. Publisher/creator: headroomlabs-ai/headroom contributors. License: https://raw.githubusercontent.com/headroomlabs-ai/headroom/d13e1966f820220b482a33c30bde1e926743939a/LICENSE. Attribution: headroomlabs-ai/headroom contributors, wiki/screenshots/cache-ttl-dashboard-history.png, Apache License 2.0, pinned at tag v0.39.1 (d13e1966).

This is the same lesson I keep finding in agent benchmarks. In the Moli audit the question was whether one efficiency chart joined two separate runs; in the GameDevBench audit it was which rows ran under the same confinement. Here it is whether a row exercised the thing being sold. For game-production pipelines, where agents read a lot of build logs and asset manifests, my expectation, which I have not measured, is that Headroom’s strengths line up with the payload. The answer-quality evidence for that kind of payload is the part still to collect.

Limitations of this audit

  • I ran lm-eval with --limit 5, not N=100, against a stub upstream, so I measured what the proxy forwards, not model scores.
  • I used the default proxy started the way the suite starts it. Other savings profiles or environment variables may behave differently, and I did not trace which internal gate declined these prompts.
  • The published table’s environment is undocumented, so I cannot say the same thing happened in that run. The reruns were also taken on one machine, on one day.

🎯 Key Takeaways

InsightImplicationNext step
The seeded Proof table reproduced exactly (178,024 to 104,261 tokens)Headroom’s token arithmetic is trustworthy and repeatableRun headroom savings on your own traffic
In my default-proxy rerun, GSM8K and TruthfulQA eval requests passed through with 0 tokens savedThe published rows report no compression ratio, so they do not establish compression fidelityDo not treat them as compression-fidelity evidence unless a ratio is reported
The suite records no compression ratio for lm-eval rows“No Accuracy Loss” there is not evidence about compressed contextAsk for accuracy paired with compression ratio
Headroom’s CI already flags ~0% compression as meaningless for recallThe right rule exists in the repoApply the same guard to the README table

🤔 New Questions

  • Which savings profile and proxy environment produced the published GSM8K and TruthfulQA numbers?
  • Would MMLU, ARC-Challenge, or HumanEval prompts trigger compression, given that they are also short and mostly prose?
  • How do SQuAD v2 and BFCL accuracy hold up at the 41% total reduction the Proof table shows for its agent scenarios, rather than at 19% and 32%?
  • What does an accuracy-versus-compression curve look like on real game-build and asset-pipeline logs?

References

Code and implementation (pinned to v0.39.1, d13e1966):

Documentation:

Datasets and tools used in the reruns:

Related on this blog:

Working on something like this?

I take a small number of paid, scoped reviews: AI agent/RAG architecture diagnosis, Unity CI & build-automation audits, and multimodal QA design review. Each one ends in a written findings document.

Work with me
This post is licensed under CC BY 4.0 by the author.

Search article titles, categories, and tags. Full text is searched when needed.

Type to search articles.