Rust crate published. npm and PyPI next.
Every model has its own cheapest spelling.
Same payload, same information, same lossless round trip. Change only how the symbols are spelled and the token bill moves by a quarter. Tokendoo measures which spelling costs least for each model, and keeps up to date with every new release of the models people actually use.
median token ratio, cl100k, derived alphabet, shipping fixtures
Wire ratio: compressed payload alone. The format legend adds 97 to 174 tokens and can erase the gain on small payloads.
published synthetic corpus · exact cl100k_base counts · 2026-08-06
01 / The problem
JSON pays for structure, over and over.
A 1,000-row JSON array repeats its keys 1,000 times. Braces, quotes, field names, repeated strings: the API bills every one of them, on every request, at per-token prices.
Meaning lives in the values. Most of the bulk is repetition you can strip and put back exactly. That is the only part Tokendoo touches.
02 / How it works
Four transforms, in order.
Each step is structural and reversible. We keep it only when the payload gets smaller. If nothing helps, we emit the input unchanged and say so in the header.
Canonicalisation
Drop whitespace, sort keys, normalise numbers. Gives every later step a deterministic input.
{ "sku": "a-114", "qty": 2 }{"qty":2,"sku":"a-114"}Key dictionary, off by default
Repeated keys become base-36 ordinals plus a table. On the shipping corpus it missed the ≥10% token-gain bar, so it stays off until it earns a place.
{"quantity":2,"status":"active"}{"0":2,"1":"active"} k:["quantity","status"]Tabular collapse
An array of same-shape objects becomes one schema row plus value rows. Keys are written once, not once per record.
[{"qty":2,"sku":"a-114"},
{"qty":1,"sku":"a-115"}]["__T",["qty","sku"], [[2,"a-114"],[1,"a-115"]]]
Value interning
Strings that repeat across the payload move into a table and get referenced by ordinal. The table ships with the output, and every figure on this page includes it.
"eu-west-1" × 40
"0" × 40 v:["eu-west-1"]
The result is a T3 wire string, T388:…, plus a JSON dictionary. Both count in every ratio. Quoting savings without the dictionary would overstate them.
03 / Why subscribe
The codec is free. Knowing how to spell it is not.
The transforms are four hundred lines of deterministic code, published under Apache 2.0 with the corpus and the bench. Take them. What is hard is not compressing the payload, it is choosing the symbols it compresses into.
Symbols cost whatever the tokenizer charges for them. Control characters looked obviously right, one byte each, and turned out to be the worst possible choice: they participate in no merge, so each costs a full token while encoding nothing. Deriving symbols from the target vocabulary instead, so each one is exactly one token, does better again.
| alphabet | median token ratio |
|---|---|
| control characters | 0.6575 |
| printable prefix | 0.5241 |
| bare base-36 | 0.4207 |
| derived from vocabulary | 0.3868 |
And the answer does not travel. The table derived for cl100k and the table derived for a Llama tokenizer share 438 symbols out of 1,377: two thirds of the optimal spelling for one family is wrong for the other.
That is the whole service. A vendored library freezes the answer on the day it was copied. Every model family that ships is a column that has to be measured again, and a client running last year's table against this year's model pays for it without ever finding out.
04 / Measurements
The numbers we ship against.
Exact cl100k_base counts on a synthetic, published corpus, benched 2026-08-07 with the shipped defaults: derived_cl100k_v1 alphabet, tabular collapse and value interning on, key dictionary off. Every after-count includes the dictionary.
| fixture | what it is | tokens in | tokens out | token ratio | byte ratio |
|---|---|---|---|---|---|
homogeneous_100 | 100 records, one schema | 2 903 | 1 229 | 0.4234 | 0.327 |
homogeneous_1000 | 1,000 records, one schema | 29 003 | 8 387 | 0.2892 | 0.246 |
realistic_api | mixed API response | 3 200 | 2 415 | 0.7547 | 0.689 |
size_10k | ≈10 KB mixed payload | 3 860 | 1 493 | 0.3868 | 0.304 |
size_100k | ≈100 KB mixed payload | 39 983 | 11 663 | 0.2917 | 0.246 |
size_250k | ≈250 KB mixed payload | 98 603 | 29 156 | 0.2957 | 0.247 |
prose_heavy | long text values, thin structure | 2 965 | 2 980 | 1.0051 | 1.002 |
shipping fixtures, median token ratio 0.3868. Corpus and bench code are in the repository; re-running the bench reproduces every count.
The last row is where the tool does nothing. Structural compression has nothing to grip on prose, and prose_heavy comes out 0.5% larger. We keep the row: it is the case where this does not help, and you should see it.
Round-trip is enforced in the open repository: property-based tests require decompress(compress(x)) to equal x on JSON values.
05 / Get the library
06 / How it plugs in
The path we refuse.
What we never do
# We never sit in front of the model provider. # caller ──(your OpenAI/Anthropic key)──▶ tokendoo-proxy ──▶ model ✗ # caller ──▶ compress (local) ──▶ you ──(your key)──▶ model ✓ # We never receive your provider key. # We never see the model's response. # We never receive your payload.
Language not covered yet? The codec is published under Apache 2.0: about four hundred lines of deterministic code. Reimplement it. There will be no compression API to wait for.
07 / Zero retention
Nothing to leak, by design.
The payload never leaves the machine. We do not keep it because we never receive it.
The only outbound call is the planned daily table refresh. It carries an account key and nothing else: no payload, no counters, no identifiers.
This section is the privacy notice.
There is no paste box on this page, on purpose. A product built on not collecting payloads does not ask you to upload JSON to prove that it does not keep JSON.
08 / Status & roadmap
What ships, what we are measuring, what is planned.
as of 2026-08-08
Shipped
- tokendoo 0.1.0 on crates.io; API docs on docs.rs/tokendoo
- Lossless round-trip over JSON values, checked by 1,000+ property-test cases
- Local Rust library; compression never requires a network hop
- JSON mode only; inputs up to 256 KiB
Being measured now
- Whether a model answers as well on compressed input as on the original. We do not claim it does. The harness exists, and the results will be published whatever they say.
- Cost of the legend that teaches a model the format: 97 to 174 tokens, enough to erase the gain on small payloads. Counted, not hidden.
Planned
- Code mode: a useful one is necessarily lossy, which breaks the round-trip guarantee. It waits until that trade-off can be stated clearly.
- Text mode: prose sits at ratio 1.005. Structural compression has nothing to do there.
- Library on npm and PyPI from the same Rust core.
- Table distribution endpoint: versioned alphabet tables only, never payloads. The wire format must carry a versioned remote table identifier; today's T3 header encodes an alphabet id in the flags but not that.
- Subscription: current alphabet tables by model family covered.
09 / Price
You pay for current tables, not for compression.
Planned. Tier structure only; prices come when the tiers exist.
A local library cannot bill usage honestly: a self-reported counter is manipulable, and catching unreported use would need the telemetry we refuse. So the subscription sells access to alphabet tables, the way GeoIP vendors sell the database. The engine does not break when tables age; it just ages. That is the whole argument.
Free
Tables frozen at the library's version. No updates. Compresses, round-trip holds, works offline indefinitely. Not crippled: it simply ages.
One family
Current tables for OpenAI, Anthropic, Google or Llama.
All families
Current tables for every family we cover.
Custom
Tables derived from your own payloads. The answer to "your fixtures are synthetic", not a premium upsell.
No invoice depends on a token count, so the exact-tokenizer-in-production question closes again.
We do not withhold a better table from a paying customer to manufacture retention. Coverage is the axis, not update cadence.
10 / Questions
The objections, before you raise them.
Seven questions a sceptical reader asks in the first two minutes. The answers are the ones we would give in a code review.
Why not just gzip the payload?
Doesn't prompt caching already solve this?
1,024 tokens, and writes cost 1.25x base. So it pays on stable prefixes with sustained traffic, and does nothing for payloads that change per call, sporadic traffic, or multi-tenant contexts where no prefix is shared. The two also compose: cache writes and reads are billed per token, so compressing first lowers both. We do not replace the cache, we shrink what it is billed on.Does the model actually understand the compressed form?
Why a subscription? This is four hundred lines of code.
Apache 2.0, along with the corpus and the bench. Take it. What the subscription sells is not the transformation but knowing which one to apply: the optimal symbol alphabet is a property of the tokenizer, not of the codec, and every new model family invalidates the answer. A vendored library freezes that decision on the day it was copied. Models, meanwhile, keep shipping.Your fixtures are synthetic and you wrote them.
Zero retention. How would I know?
You are adding a network hop to save a few cents.
11 / npm and PyPI
Leave an email. We send one when the library reaches npm and PyPI.
The Rust crate is published. One email when npm and PyPI ship, nothing before, nothing after.