Token compression · V2

Fewer tokens.
Same context.

Compression V2 upgrades Edgee's Compress pillar for coding agents. Three techniques across two layers: sharper tool result trimming, new task-aware tool surface reduction, and Layer 2 output brevity, semantically lossless on code tasks.

On real fleets that is a 15-20% cut in token spend, which is the number our customers actually see. On SWE-bench Lite, under controlled conditions, the same three techniques measure 50%. We publish both, and we tell you which is which.

Drop-in CLI · works with your existing API keys and plans · no code changes

compression.toml
SWE-bench Lite

tokens the model sees

9,210

of 18,420 uncompressed

one benchmark session

−50%

reduction

kept · sent to model9,210 removed

Toggle a technique to see the benchmark session move.

Install
  • 15-20%

    lower token bills in production

    what active customers see from compression alone

  • −50%

    cost on SWE-bench Lite

    controlled benchmark, all three techniques on

  • <12ms

    P50 gateway overhead

    compression time at the edge

  • 0

    code changes

    drop-in CLI wrapper

Production figure: rolling 30-day aggregate across active Edgee customers, compression alone, no routing. Benchmark figure: SWE-bench Lite, all three techniques on, full methodology on the blog. They are different measurements on different baselines. Never add them together.

What you actually save

Compression vendors quote benchmark numbers and let you assume they are invoice numbers. We would rather you knew the gap before you sign anything, so here are both, side by side.

In production

15-20%

lower token bills, from compression alone

This is the range our active customers actually see on their bill, measured on a rolling 30 days across live fleets. Plan your budget around it. Anything above it is upside, not a promise.

Real sessions are not benchmark tasks. Workloads are mixed, sessions are shorter, some agents are already tuned for terseness, and most teams enable a subset of the three techniques per API key. Every one of those pulls the number below what a controlled run produces.

Rolling 30-day aggregate, active Edgee customers. Compression alone, no routing.

In the lab · SWE-bench Lite

−50%

cost on SWE-bench Lite, all three techniques on

Khaled Maâmra, Research Engineer at Edgee, evaluated each technique on SWE-bench Lite: 300 real GitHub issues from popular Python repositories, run in agent mode. It is a ceiling measured under controlled conditions, not a forecast for your fleet.

  • Vanilla Claude Code against Edgee, paired per task, replicate order shuffled within each task.
  • A random nonce per replicate so every run starts on a cold prefix cache.
  • Token counts read from Claude Code session logs, cost computed from the published price table.
  • Paired sign test for direction, 10,000-sample bootstrap for magnitude.
Read the full methodology

Compression is one of three pillars. The 15-20% above is compression working on its own, with no routing involved. Teams that also run budget-driven Strategies on top of it are where the up to 70% combined reduction we quote elsewhere comes from. Different layer, different baseline, quoted separately.

Two layers. Three techniques.

Token compression splits in two. Layer 1 (Input) handles what enters the context window — tool results, tool definitions, codebase context — roughly 99% of token volume in a coding session. Layer 2 (Output) trims the model's response: small in volume, high in ROI. V2 sharpens both and adds a new Layer 1 technique that compresses the tool surface itself.

What's new in Compression V2

Each technique is a named config flag you toggle independently. The percentages below are each technique's share of one SWE-bench Lite session, measured in tokens and never summed across techniques. Each card also carries the measured result for that technique on its own, including the cases where the signal is directional rather than statistically significant.

Tool result trimming

Improved in V2

tool_result_trimming · Layer 1 (Input)

−10%1,842 tokens

Filters CLI and tool results before they reach the model — boilerplate, pagination markers, ANSI escape sequences, repeated headers, and verbose framing. Inspired by RTK. V2 trims harder while keeping the output semantically intact for code tasks: a 980-token directory listing becomes a dense 340-token one the model reads just as well.

Measured on SWE-bench Lite

10.4% median per-task cost reduction over 6 SWE-bench Lite tasks, 4 of 6 favor Edgee. Directional rather than significant at this sample size, and the gains compound on longer sessions.

Tool surface reduction

New in V2

tool_surface_reduction · Layer 1 (Input)

−10%1,842 tokens

Agents send the model the union of every MCP server, skill, and tool definition on every request — even when 95% are irrelevant to the task. V2 runs a small, fast classifier that scores each tool against the classified task, then strips or down-scopes the rest before the request hits the model. Your IDE still exposes everything; the model only sees a curated, task-relevant subset. No more toggling MCP servers on and off by hand like a mixing desk.

Measured on SWE-bench Lite

33.0% fewer total tokens over 8 tool-heavy MCP tasks, 8 of 8 favor Edgee, sign test p = 0.008. Cost falls around 10%: the tokens removed are cache reads, the cheapest class.

Output brevity

New in V2

output_brevity · Layer 2 (Output)

−30%5,526 tokens

Reduces the verbosity of model responses without losing technical content — same answer, fewer tokens. Pick the level (light, medium, hard) to trade aggressiveness against tone. Small in token volume, high in ROI: output is the ~1% of traffic you pay the most for.

Measured on SWE-bench Lite

27.5% median per-task cost reduction over 6 SWE-bench Lite tasks, 6 of 6 favor Edgee, sign test p = 0.031. Bootstrap 95% CI on the token ratio: 0.41x to 0.84x, entirely below 1.

Compression is designed to be semantically lossless for code-oriented tasks. Khaled Maâmra, our Research Engineer, validated this on SWE-bench Lite, where the compressed prompt produced outputs statistically indistinguishable from the original. Extremely short prompts compress less, tool-use schemas are passed through untouched, and when in doubt Edgee skips compression.

See it on a live session

Watch the three V2 techniques applied to a real Claude Code session — what gets trimmed, which tools get cut, and how the token bill moves.

A walkthrough of the three V2 techniques applied to a live Claude Code session: tool result trimming, tool surface reduction, and output brevity.

Drop-in install

Install the CLI once. Launch any supported coding agent through it. Compression V2 runs per session — your CLAUDE.md and MCP servers stay put.

# Install the Edgee CLI
curl -fsSL https://edgee.ai/install.sh | bash

# Launch Claude Code through the compression proxy
edgee launch claude

Full CLI guide in the Edgee documentation.

Measure every saved token

Compression without observability is flying blind. Every session reports its compression ratio, tokens saved, and estimated cost avoided — at developer and team level.

Technical FAQ

Stop sending verbose prompts. Ship Compression V2.

15-20% off the token bill in production, semantically lossless on code tasks. On SWE-bench Lite the same three techniques take a session from 18,420 to 9,210 tokens, a 50% reduction.

Works with your existing API keys and plans. No lock-in.