Token Optimizer Mcp
ooples/token-optimizer-mcp · 405 stars · JavaScript · MIT
MCP server Measure token savings per AI coding agent, optimize context, and share a live local knowledge graph across 16 CLI clients.
Install
The repo has no one-line install. Follow its README.
Files
and ships the benchmark so you can check it.
Why it wins
Providers cache the prompt prefix: cached tokens re-read at 0.1x, rewritten ones bill at 1.25x. Most compressors optimise bytes removed and ignore that multiplier. This one optimises the bill.
It wins both columns, against their real implementation, on their own fixtures. Not a reimplementation and not our fixtures: their harness compresses the payload, dumps it, their resolver redeems their own markers, and ours is handed the identical bytes.
Recorded 2026-09-25 at 18a2d085, by the command in the record's regenerate field. That second arm runs HeadRoom itself, so CI does not re-derive it the way it re-derives the table below -- it checks this provenance and these figures against bench/compression/headroom/results/head-to-head.json instead.
Over the corpus the shipped default takes 89.4% of the characters and 80.9% of the tokens; theirs takes 51.3% and 62.8%. Of the 9,919 identifiers planted in the corpus, 45 end up unrecoverable on our side and 8 on theirs.
Read the rows, though, because the ones we lose are not compression results. On grep-output, codebase-exploration and raw-build-log they report ~99.7% and we report 57-71%. Their number there is a content-cache reference: the block is not made smaller, it is taken out of the request, put in a store, and replaced by a 24-character <<ccr:...>> marker. The same is true of the four rows where they edge us out in the nineties.
So the last column is the one that decides a turn. A marker is not the content: to read what it stands for, the agent spends a request. zero-turn ids counts the identifiers planted in each workload that need no such request -- still in the text, or rebuildable from the text alone. On the three rows we lose on reduction we take that column outright, 509-5, 1046-4 and 430-3: their marker leaves almost none of it behind, our skeleton leaves all of it. We lose it on eight rows -- agent-loop, agent-loop-logs, code-search, human-authored-json, issue-triage, relevance-probe, repeated-reads and sre-debugging -- because our reduction there comes from spilling too, and a spill costs the same turn theirs does; browser-session ties at 336-336. Over the corpus it is 3,068 of 9,919 for us against 5,238 for them: this column goes to them, and the two columns have to be read together or each one flatters somebody.
The unit count moved with the instrument, not with the product. Every rule that found a retention unit keyed on a digit-bearing token, a markdown heading or a declaration, so issue-triage ("number": 3000) and relevance-probe ("id": "evt_0", the needle that workload exists to find) each scored ZERO units and reported a tie on an empty set. Counting a string value under an object key and a quoted substring inside a longer string -- on both arms, and excluding multi-line values, which are not literal substrings of the block they came from -- takes the corpus from 6,098 units to 9,919 and reverses this column, which read 2,717 against 2,278 before the fix.
browser-session used to be the one genuine engine loss on this corpus, at 6.0% against their 21.8%. It is now 93.7%, from folding exact long repeats inside a block rather than at a boundary someone else drew — a serialised message list with inline images is a single 780,000-character line, which every other pass here reads as one unit.
Most of that 93.7% is the fixture, and the honest number is lower. This payload holds four images, two of them distinct, and the base64 in them is generated rather than photographic: one distinct image alone folds from 160,032 characters to 5,572, which no real PNG would do. What carries over to a real session is the duplication — half the image bytes here are a second copy of an image already in the request, which is 44.3% of the whole payload, and an agent re-sending a screenshot it has already sent is ordinary. Folding only that is a ~44% reduction, still ahead of their 21.8% on the same row, and the generated base64 is worth about 49 points on top that we would not claim twice.
The run the marker names is still in the output above it, so the reader rebuilds it without asking for anything, which is why the characters fall much further than the tokens (28.9%) — the image tokens are counted from pixels on both arms either way.
ours, dial on is that like-for-like, and it is substitution, not reduction. Set spillWholeBlockBelow and a block our engines could not compress is moved out of the request whole, leaving [... n bytes, moved whole -> path]. It wins all twelve rows on both denominators, but nothing there was compressed: the bytes are on disk, at 1.00x the input, and the ratio is a measurement of a move. Any quote of that column that omits this sentence is a misquote.
Two things make it the better version of their trade, which is the only reason it exists. It is gated on the saving our engines actually reached, not on block size, so a block we compressed well stays in the request where the reader still has it — a content cache moves it regardless. And the marker carries a path the agent already has, so following it is a Read it issues itself, where a cache reference costs a retrieval round trip and degrades to [unresolved: entry not found] once the store has moved on.
It is off by default, because the trade is real: every one of the 2,334 identifiers a reader can rebuild from our output with no extra turn sits in exactly the blocks it would move -- the five rows where the dial fires are the five rows that reconstructible column lives on, and nowhere else. On by default, this would be their product with a better marker.
ours, preset is the same dial at the setting a caller would actually run -- spillWholeBlockBelow: 0.9, so a block is moved only where the engines could not take 90% off it. On seven of the twelve rows it changes nothing at all: the figure is the shipped one, the block stays in the request, and the zero-turn count is untouched. On the other five it matches their headline -- 100.0% on codebase-exploration, 99.9% on grep-output, 100.0% on raw-build-log -- and it buys that the same way they do. Those five rows are exactly where our zero-turn wins live, and the column takes all of them: 509, 1046, 430, 361 and 336 go to 0. Over the corpus it is 98.1% of the characters against 89.4%, and 386 zero-turn identifiers against 3,068. It also puts 46 identifiers beyond
Facts
- Kind
- MCP server
- Repo
- ooples/token-optimizer-mcp
- Group
- Uncategorized
- Stars
- 405
- License
- MIT
- Language
- JavaScript
- Last push
- 2026-10-09
- Forks
- 67
- Topics
- ai, caching, claude, compression, gemini-cli-extension, llm, mcp, mcp-server, token-optimization
- 1Everythingmodelcontextprotocol/serversThis MCP server attempts to exercise all the features of the MCP protocol. It is not intended to be a useful server, but rather a test server for builders of MCP clients. It implements prompts, tools, resources, sampling, and more to showcase MCP capabilities.85.8k
- 2Fetchmodelcontextprotocol/serversA Model Context Protocol server that provides web content fetching capabilities. This server enables LLMs to retrieve and process content from web pages, converting HTML to markdown for easier consumption.85.8k
- 3Gitmodelcontextprotocol/serversA Model Context Protocol server for Git repository interaction and automation. This server provides tools to read, search, and manipulate Git repositories via Large Language Models.85.8k
- 4Memorymodelcontextprotocol/serversA basic implementation of persistent memory using a local knowledge graph. This lets Claude remember information about the user across chats.85.8k
- 5Sequential Thinkingmodelcontextprotocol/serversAn MCP server implementation that provides a tool for dynamic and reflective problem-solving through a structured thinking process.85.8k
- 6Timemodelcontextprotocol/serversA Model Context Protocol server that provides time and timezone conversion capabilities. This server enables LLMs to get current time information and perform timezone conversions using IANA timezone names, with automatic system timezone detection.85.8k