Skip to main content
Kodelyth ECC
guide

How to Cut Claude Code Token Costs — Measured, Not Guessed

Where tokens actually go in an AI coding session, and the three levers that move the bill: filtering tool output before it reaches the model, compressing replies, and matching model tier to task. With a real 1,285-command ledger showing 64.1% input reduction.

Most advice about AI coding cost is some version of "write shorter prompts." That is close to irrelevant. Your prompts are not where the tokens go.

In a real coding session the bill is dominated by three things, in roughly this order:

  1. Tool output fed back to the model — the full text of every file read, every git status, every test run, every npm install log.
  2. The model's own replies — explanation, restatement, preamble.
  3. Context carried turn to turn — everything above, re-sent on each request.

Prompt length is a rounding error against those. So the levers that matter are the ones that act on tool output, reply length, and model tier.

Lever 1: filter tool output before the model sees it

This is the largest and least-discussed lever. When an assistant runs git status in a busy repo, the raw output can be hundreds of lines, nearly all of it noise the model does not need to answer the question.

Filtering that output before it reaches the context window is pure savings — the model's answer does not change, because the information it used was never in the discarded part.

Kodelyth ECC bundles RTK (Rust Token Killer) to do this, transparently, via a hook that rewrites commands. Here is an actual ledger from a working machine — 1,285 commands, not a benchmark:

total_commands:    1,285
total_input:       7,964,612    (raw)
total_output:      2,858,501    (after RTK filter)
total_saved:       5,107,394    ← 64.1% average reduction

Per-command savings range from 60-90% depending on how noisy the command is. git status, ls, docker ps, cargo test and 100+ others are covered.

You can check your own number at any time:

rtk gain

The thing to internalise: this lever is free. It costs no quality, because the filtered bytes were never load-bearing.

Lever 2: shorten the replies

Output tokens are typically billed several times higher than input tokens, so reply length punches above its weight.

Most assistant replies carry a lot of connective tissue — restating the question, narrating what is about to happen, summarising what just happened. Stripping that while preserving every technical detail is worth 40-70% of output tokens on a typical reply.

ECC ships this as Terse Mode, a four-level dial:

/terse        # on
/terse 3      # more aggressive
/terse off

The constraint that makes it safe: code blocks and shell commands are preserved byte-exact. Compression applies to prose only. A terse reply contains the same facts as a verbose one — it just stops explaining them twice.

Lever 3: stop using a frontier model for trivial work

The expensive habit is running every task on the heaviest available model. Renaming a variable, fixing a typo, adding a JSDoc comment, and formatting a file do not need frontier reasoning — and a large share of turns in a real session are exactly that.

A workable split:

TaskTier
Rename, format, typo, single comment, one-line fixsmallest
Build a component, write a test, debug one bug, refactor a filemid
Architecture, migration planning, security audit, cross-cutting buglargest

Most sessions should sit in the middle tier, drop to the smallest for mechanical edits, and reserve the largest for decisions that would take a senior engineer an hour of thought.

ECC encodes this as an always-on routing rule that classifies the task by scope, ambiguity and risk, and says so when the active model is badly mismatched — rather than silently switching, which would be worse.

Context hygiene, briefly

Two habits that cost nothing:

  • Clear context between unrelated tasks. Stale context is re-sent on every subsequent request, so an irrelevant 20k-token file read keeps being paid for.
  • Compact before long sessions get long, not after. Compaction near the end of the window is both more expensive and more lossy.

What this adds up to

The levers compound, because they act on different parts of the bill:

  • Input filtering cuts the largest component by roughly 60-90%.
  • Reply compression cuts output by 40-70%.
  • Tier routing cuts the per-token rate on the share of turns that are trivial.

None of the three costs output quality, which is what separates them from the usual advice. Filtering removes bytes the model never used. Compression removes restatement, not facts. Tier routing only moves work that did not need the reasoning it was getting.

Try it

RTK and Terse both install and wire themselves up:

npm i -g kodelyth-ecc && kodelyth-ecc

Then check your own ledger rather than trusting the numbers above:

rtk gain
kodelyth-ecc dashboard

The dashboard runs on localhost, reads only local files, and sends nothing anywhere.

Last updated: 2026-10-05T00:00:00.000Z · v2.24.5