ClaudeCodeMod

All shelves / MCP servers

Token Savior

mibayy/token-savior · 977 stars · Python · MIT

MCP server MCP server that gets Claude to 97.9% (188/192) on a real coding benchmark at -80% active tokens and -83% wall time, vs 78.3% plain. Structural code navigation + persistent memory engine. Works with every MCP client.

Install

In your shell
pip install "token-savior-recall[mcp]"

These repos do not share one command. When an entry shows a command, it was copied as published. Check the repo's README before you run it.

Open the repo

Files

README.md

Token Savior

One MCP server. One profile. 97.9% on tsbench at -80% tokens. Structural code navigation, persistent memory, and Bash command rewriting for AI coding agents.

-brightgreen)

mibayy.github.io/token-savior -- project site + benchmark landing Benchmark source + fixtures: not currently published (see Reproducing the score below)

Benchmark -- 96 real coding tasks (Claude Opus 4.7, May 2026)

Reproduces with the optimized profile (single env var). The harness that produced these numbers is described below; its repository is not public at the moment, so take the figures as reported rather than as independently verifiable.

A re-measurement was published here on 2026-08-09 and has been withdrawn on 2026-08-10. It reported new-token savings from a small replacement harness. Those numbers did not measure this server at all: across 143 benchmark sessions, exactly one called a Token Savior tool. The client running the harness had MCP deferred-tool loading enabled, so all 18 tools sat behind a ToolSearch lookup instead of appearing in the model's manifest. The model never saw them and fell back to Grep and Read — 66 greps, 30 reads, one MCP call. What varied between the "profiles" was the size of the cached prefix, not what the agent did.

The lesson is worth more than the numbers were: a benchmark of an MCP server must assert that its tools were actually called. Ours did not, so it happily compared two identical agents. That assertion now exists in the harness.

The headline figures above therefore stand as reported and unverified, as stated in the previous paragraph.

Re-measured on 2026-10-07, with the tools actually called. Same 8 tasks on a frozen copy of this repo (commit 9d15a1e, no CLAUDE.md), Claude Haiku 4.5, plain Claude Code against the auto profile with alwaysLoad. The harness checks that the Token Savior sessions really called the server.

All 48 sessions succeeded. v4.22.0, measured once, saved 14% and used the tools in 5 sessions out of 8. Editing, measured on 2026-10-08 on the same repository (3 tasks in files of 1 900 to 3 700 lines, graded by running the code and the module's tests, 4 passes): every session of both arms succeeded, and cost is too noisy to call a gain (from -54% to +16% per pass). Native Edit only needs a partial Read, so there is little left to remove there.

Wall time is worse because each benchmark session starts a fresh server that indexes the repo first; a long-lived session pays that once. What moved from v4.22.0 to v4.23.0 is mostly wording: search_codebase now names the function holding each hit, and the tool descriptions say plainly which native call they replace.

Who starred this repo?

On June 30, 2026 GitHub restricted stargazer and watcher lists to repo admins and collaborators, which broke every "who starred my repo" tool at once. I rebuilt one that still works, precisely because it only reads repos you own or can push to: starscope ranks the people who starred or forked your repo by influence, and surfaces their social accounts when their GitHub profile declares them.

Numbers on this very repo, computed with it: 1,147 people, 27% with a public social account, and the most followed carries 18,922 followers. The named list is visible to the repo owner and to nobody else — the public page shows aggregates only.

What's new

Release notes live where they can't drift out of sync with the code:

Quick start

pip install "token-savior-recall[mcp]"

Add to your MCP config (e.g. Claude Code):

{
  "mcpServers": {
    "token-savior-recall": {
      "command": "/path/to/venv/bin/token-savior",
      "alwaysLoad": true,
      "env": {
        "WORKSPACE_ROOTS": "/path/to/project1,/path/to/project2",
        "TOKEN_SAVIOR_CLIENT": "claude-code",
        "TOKEN_SAVIOR_PROFILE": "optimized"
      }
    }
  }
}

"alwaysLoad": true matters on Claude Code (2.1.28x and later): without it, every tool of the server sits behind ToolSearch and the model reaches for Read and Grep, which are already loaded. Measured on this project's benchmark before the option existed: one session in 143 called a Token Savior tool. With it, the manifest costs a few thousand cached tokens per session.

That's it. TOKEN_SAVIOR_PROFILE=optimized ships the Pareto-optimum config that wins tsbench. It bundles:

  • tiny_plus (15 hot tools manifest)
  • thin inputSchema (-44% manifest)
  • capture sandbox disabled
  • memory hooks gated for cross-project safety

No other tuning needed.

Activation (Bash compaction + rewriting)

Bash compaction and the PreToolUse rewriter are opt-in. Two env vars and one CLI call:

export TS_BASH_COMPACT=1       # PostToolUse output compactors (34 of them)
export TS_BASH_REWRITE=1       # PreToolUse command rewriter (10 rules)

ts init --agent claude --yes   # auto-merge hooks into ~/.claude/settings.json

ts init is idempotent. It detects existing hook entries, dedups by (matcher, command), prints a unified diff, and backs up settings.json to .bak-YYYYMMDD-HHMMSS (UTC) before writing. Supported agents: claude, cursor, gemini, codex, openclaw. Pass --dry-run to preview, or --global to write the user-level config.

Optional audit log of every rewrite:

export TS_BASH_REWRITE_LOG=$HOME/.local/state/token-savior/rewrites.jsonl

Compactor catalog (34)

Each compactor is a pure function (no I/O, no globals) returning a token-efficient rendering. The dispatcher returns None when no matcher fires, leaving the existing sandbox path untouched. Compound commands (cd ... && cmd) fall through to the last meaningful segment.

These run in PostToolUse, so they do not shrink the current turn. The hook fires after the tool has returned; it can add context, not remove it. The compact rendering is appended below the raw output, which stays. What you gain is persistence: the full output goes to the capture sandbox and outlives a context compaction. For an actual reduction of what reaches the model, use the PreToolUse rewriter (TS_BASH_REWRITE=1) — it edits the command before it runs.

ts_discover -- find missed TS opportunities

New MCP tool that scans your Claude Code transcripts for patterns where TS tools would have been cheaper than what the agent actually did.

ts_discover()                       # active project, last 30 days
ts_discover(project=None)           # ALL transcript projects
ts_discover(format="adoption")      # TS vs native ratio per session
ts_discover(format="adoption_json") # same, JSON

Findings: Read->Grep->Read chains, sequential find_symbol, edits without get_edit_context, memory_search without memory_index,

Facts

Kind
MCP server
Repo
mibayy/token-savior
Group
Uncategorized
Stars
977
License
MIT
Language
Python
Last push
2026-10-08
Forks
103
Homepage
mibayy.github.io/token-savior

More on this shelf

  1. 1Everythingmodelcontextprotocol/serversThis MCP server attempts to exercise all the features of the MCP protocol. It is not intended to be a useful server, but rather a test server for builders of MCP clients. It implements prompts, tools, resources, sampling, and more to showcase MCP capabilities.85.8k
  2. 2Fetchmodelcontextprotocol/serversA Model Context Protocol server that provides web content fetching capabilities. This server enables LLMs to retrieve and process content from web pages, converting HTML to markdown for easier consumption.85.8k
  3. 3Gitmodelcontextprotocol/serversA Model Context Protocol server for Git repository interaction and automation. This server provides tools to read, search, and manipulate Git repositories via Large Language Models.85.8k
  4. 4Memorymodelcontextprotocol/serversA basic implementation of persistent memory using a local knowledge graph. This lets Claude remember information about the user across chats.85.8k
  5. 5Sequential Thinkingmodelcontextprotocol/serversAn MCP server implementation that provides a tool for dynamic and reflective problem-solving through a structured thinking process.85.8k
  6. 6Timemodelcontextprotocol/serversA Model Context Protocol server that provides time and timezone conversion capabilities. This server enables LLMs to get current time information and perform timezone conversions using IANA timezone names, with automatic system timezone detection.85.8k