Most advice about AI coding cost is some version of "write shorter prompts." That is close to irrelevant. Your prompts are not where the tokens go.
In a real coding session the bill is dominated by three things, in roughly this order:
- Tool output fed back to the model — the full text of every file read, every
git status, every test run, everynpm installlog. - The model's own replies — explanation, restatement, preamble.
- Context carried turn to turn — everything above, re-sent on each request.
Prompt length is a rounding error against those. So the levers that matter are the ones that act on tool output, reply length, and model tier.
Lever 1: filter tool output before the model sees it
This is the largest and least-discussed lever. When an assistant runs git status
in a busy repo, the raw output can be hundreds of lines, nearly all of it noise the
model does not need to answer the question.
Filtering that output before it reaches the context window is pure savings — the model's answer does not change, because the information it used was never in the discarded part.
Kodelyth ECC bundles RTK (Rust Token Killer) to do this, transparently, via a hook that rewrites commands. Here is an actual ledger from a working machine — 1,285 commands, not a benchmark:
total_commands: 1,285
total_input: 7,964,612 (raw)
total_output: 2,858,501 (after RTK filter)
total_saved: 5,107,394 ← 64.1% average reductionPer-command savings range from 60-90% depending on how noisy the command is.
git status, ls, docker ps, cargo test and 100+ others are covered.
You can check your own number at any time:
rtk gainThe thing to internalise: this lever is free. It costs no quality, because the filtered bytes were never load-bearing.
Lever 2: shorten the replies
Output tokens are typically billed several times higher than input tokens, so reply length punches above its weight.
Most assistant replies carry a lot of connective tissue — restating the question, narrating what is about to happen, summarising what just happened. Stripping that while preserving every technical detail is worth 40-70% of output tokens on a typical reply.
ECC ships this as Terse Mode, a four-level dial:
/terse # on
/terse 3 # more aggressive
/terse offThe constraint that makes it safe: code blocks and shell commands are preserved byte-exact. Compression applies to prose only. A terse reply contains the same facts as a verbose one — it just stops explaining them twice.
Lever 3: stop using a frontier model for trivial work
The expensive habit is running every task on the heaviest available model. Renaming a variable, fixing a typo, adding a JSDoc comment, and formatting a file do not need frontier reasoning — and a large share of turns in a real session are exactly that.
A workable split:
| Task | Tier |
|---|---|
| Rename, format, typo, single comment, one-line fix | smallest |
| Build a component, write a test, debug one bug, refactor a file | mid |
| Architecture, migration planning, security audit, cross-cutting bug | largest |
Most sessions should sit in the middle tier, drop to the smallest for mechanical edits, and reserve the largest for decisions that would take a senior engineer an hour of thought.
ECC encodes this as an always-on routing rule that classifies the task by scope, ambiguity and risk, and says so when the active model is badly mismatched — rather than silently switching, which would be worse.
Context hygiene, briefly
Two habits that cost nothing:
- Clear context between unrelated tasks. Stale context is re-sent on every subsequent request, so an irrelevant 20k-token file read keeps being paid for.
- Compact before long sessions get long, not after. Compaction near the end of the window is both more expensive and more lossy.
What this adds up to
The levers compound, because they act on different parts of the bill:
- Input filtering cuts the largest component by roughly 60-90%.
- Reply compression cuts output by 40-70%.
- Tier routing cuts the per-token rate on the share of turns that are trivial.
None of the three costs output quality, which is what separates them from the usual advice. Filtering removes bytes the model never used. Compression removes restatement, not facts. Tier routing only moves work that did not need the reasoning it was getting.
Try it
RTK and Terse both install and wire themselves up:
npm i -g kodelyth-ecc && kodelyth-eccThen check your own ledger rather than trusting the numbers above:
rtk gain
kodelyth-ecc dashboardThe dashboard runs on localhost, reads only local files, and sends nothing anywhere.