Skip to content

Rust crate published. npm and PyPI next.

Every model has its own cheapest spelling.

Same payload, same information, same lossless round trip. Change only how the symbols are spelled and the token bill moves by a quarter. Tokendoo measures which spelling costs least for each model, and keeps up to date with every new release of the models people actually use.

0.3868

median token ratio, cl100k, derived alphabet, shipping fixtures

Wire ratio: compressed payload alone. The format legend adds 97 to 174 tokens and can erase the gain on small payloads.

published synthetic corpus · exact cl100k_base counts · 2026-08-06

JSON pays for structure, over and over.

A 1,000-row JSON array repeats its keys 1,000 times. Braces, quotes, field names, repeated strings: the API bills every one of them, on every request, at per-token prices.

Meaning lives in the values. Most of the bulk is repetition you can strip and put back exactly. That is the only part Tokendoo touches.

Four transforms, in order.

Each step is structural and reversible. We keep it only when the payload gets smaller. If nothing helps, we emit the input unchanged and say so in the header.

Canonicalisation

Drop whitespace, sort keys, normalise numbers. Gives every later step a deterministic input.

{ "sku": "a-114",  "qty": 2 }
{"qty":2,"sku":"a-114"}

Key dictionary, off by default

Repeated keys become base-36 ordinals plus a table. On the shipping corpus it missed the ≥10% token-gain bar, so it stays off until it earns a place.

{"quantity":2,"status":"active"}
{"0":2,"1":"active"}   k:["quantity","status"]

Tabular collapse

An array of same-shape objects becomes one schema row plus value rows. Keys are written once, not once per record.

[{"qty":2,"sku":"a-114"},
 {"qty":1,"sku":"a-115"}]
["__T",["qty","sku"],
 [[2,"a-114"],[1,"a-115"]]]

Value interning

Strings that repeat across the payload move into a table and get referenced by ordinal. The table ships with the output, and every figure on this page includes it.

"eu-west-1" × 40
"0" × 40   v:["eu-west-1"]

The result is a T3 wire string, T388:…, plus a JSON dictionary. Both count in every ratio. Quoting savings without the dictionary would overstate them.

The codec is free. Knowing how to spell it is not.

The transforms are four hundred lines of deterministic code, published under Apache 2.0 with the corpus and the bench. Take them. What is hard is not compressing the payload, it is choosing the symbols it compresses into.

Symbols cost whatever the tokenizer charges for them. Control characters looked obviously right, one byte each, and turned out to be the worst possible choice: they participate in no merge, so each costs a full token while encoding nothing. Deriving symbols from the target vocabulary instead, so each one is exactly one token, does better again.

Same codec, same corpus, same fixtures. Only the symbol spelling changes.
alphabetmedian token ratio
control characters0.6575
printable prefix0.5241
bare base-360.4207
derived from vocabulary0.3868

And the answer does not travel. The table derived for cl100k and the table derived for a Llama tokenizer share 438 symbols out of 1,377: two thirds of the optimal spelling for one family is wrong for the other.

That is the whole service. A vendored library freezes the answer on the day it was copied. Every model family that ships is a column that has to be measured again, and a client running last year's table against this year's model pays for it without ever finding out.

The numbers we ship against.

Exact cl100k_base counts on a synthetic, published corpus, benched 2026-08-07 with the shipped defaults: derived_cl100k_v1 alphabet, tabular collapse and value interning on, key dictionary off. Every after-count includes the dictionary.

Token and byte ratios per fixture, exact cl100k_base counts, derived_cl100k_v1, 2026-08-07
fixturewhat it istokens intokens outtoken ratiobyte ratio
homogeneous_100100 records, one schema2 9031 2290.42340.327
homogeneous_10001,000 records, one schema29 0038 3870.28920.246
realistic_apimixed API response3 2002 4150.75470.689
size_10k≈10 KB mixed payload3 8601 4930.38680.304
size_100k≈100 KB mixed payload39 98311 6630.29170.246
size_250k≈250 KB mixed payload98 60329 1560.29570.247
prose_heavylong text values, thin structure2 9652 9801.00511.002

shipping fixtures, median token ratio 0.3868. Corpus and bench code are in the repository; re-running the bench reproduces every count.

The last row is where the tool does nothing. Structural compression has nothing to grip on prose, and prose_heavy comes out 0.5% larger. We keep the row: it is the case where this does not help, and you should see it.

Round-trip is enforced in the open repository: property-based tests require decompress(compress(x)) to equal x on JSON values.

Install the Rust crate.

One line in your manifest. Source on GitHub, versions on crates.io.

cargo add tokendoo

The path we refuse.

What we never do

# We never sit in front of the model provider.
#   caller ──(your OpenAI/Anthropic key)──▶ tokendoo-proxy ──▶ model   ✗
#   caller ──▶ compress (local) ──▶ you ──(your key)──▶ model          ✓
# We never receive your provider key.
# We never see the model's response.
# We never receive your payload.

Language not covered yet? The codec is published under Apache 2.0: about four hundred lines of deterministic code. Reimplement it. There will be no compression API to wait for.

Nothing to leak, by design.

The payload never leaves the machine. We do not keep it because we never receive it.

The only outbound call is the planned daily table refresh. It carries an account key and nothing else: no payload, no counters, no identifiers.

This section is the privacy notice.

There is no paste box on this page, on purpose. A product built on not collecting payloads does not ask you to upload JSON to prove that it does not keep JSON.

What ships, what we are measuring, what is planned.

as of 2026-08-08

Shipped

  • tokendoo 0.1.0 on crates.io; API docs on docs.rs/tokendoo
  • Lossless round-trip over JSON values, checked by 1,000+ property-test cases
  • Local Rust library; compression never requires a network hop
  • JSON mode only; inputs up to 256 KiB

Being measured now

  • Whether a model answers as well on compressed input as on the original. We do not claim it does. The harness exists, and the results will be published whatever they say.
  • Cost of the legend that teaches a model the format: 97 to 174 tokens, enough to erase the gain on small payloads. Counted, not hidden.

Planned

  • Code mode: a useful one is necessarily lossy, which breaks the round-trip guarantee. It waits until that trade-off can be stated clearly.
  • Text mode: prose sits at ratio 1.005. Structural compression has nothing to do there.
  • Library on npm and PyPI from the same Rust core.
  • Table distribution endpoint: versioned alphabet tables only, never payloads. The wire format must carry a versioned remote table identifier; today's T3 header encodes an alphabet id in the flags but not that.
  • Subscription: current alphabet tables by model family covered.

You pay for current tables, not for compression.

Planned. Tier structure only; prices come when the tiers exist.

A local library cannot bill usage honestly: a self-reported counter is manipulable, and catching unreported use would need the telemetry we refuse. So the subscription sells access to alphabet tables, the way GeoIP vendors sell the database. The engine does not break when tables age; it just ages. That is the whole argument.

Free

Tables frozen at the library's version. No updates. Compresses, round-trip holds, works offline indefinitely. Not crippled: it simply ages.

One family

Current tables for OpenAI, Anthropic, Google or Llama.

All families

Current tables for every family we cover.

Custom

Tables derived from your own payloads. The answer to "your fixtures are synthetic", not a premium upsell.

No invoice depends on a token count, so the exact-tokenizer-in-production question closes again.

We do not withhold a better table from a paying customer to manufacture retention. Coverage is the axis, not update cadence.

The objections, before you raise them.

Seven questions a sceptical reader asks in the first two minutes. The answers are the ones we would give in a code review.

Why not just gzip the payload?
Because the model does not decompress anything. It reads text. A gzipped blob in base64 is unreadable to it and costs more tokens than the original. The compression has to stay legible to the model, which is the entire constraint and the reason this is not trivial.
Doesn't prompt caching already solve this?
On a segment we do not compete for, yes. Caching applies to a shared prefix across differing requests, at a tenth of the input price on read, and no local cache replicates that. But the default TTL is five minutes, the minimum cacheable prefix is 1,024 tokens, and writes cost 1.25x base. So it pays on stable prefixes with sustained traffic, and does nothing for payloads that change per call, sporadic traffic, or multi-tenant contexts where no prefix is shared. The two also compose: cache writes and reads are billed per token, so compressing first lowers both. We do not replace the cache, we shrink what it is billed on.
Does the model actually understand the compressed form?
That question has no single answer, because it is a property of each model, not of the codec. A model has a tokenizer, which decides what the compressed form costs, and a capacity to read it, which decides what it costs you in accuracy. Both are measured per model and published together, alongside the optimal alphabet for that model. Where a model is not yet covered, the coverage page says so rather than implying otherwise. And if a transform turns out to cost accuracy on a given model, it ships disabled for that model.
Why a subscription? This is four hundred lines of code.
It is, and it is open source under Apache 2.0, along with the corpus and the bench. Take it. What the subscription sells is not the transformation but knowing which one to apply: the optimal symbol alphabet is a property of the tokenizer, not of the codec, and every new model family invalidates the answer. A vendored library freezes that decision on the day it was copied. Models, meanwhile, keep shipping.
Your fixtures are synthetic and you wrote them.
True, which is why they are published and reproducible. Two of them are on this page precisely because they are unflattering: a realistic mixed API response barely gains, and prose-heavy JSON comes out slightly larger. Add your own payloads to the corpus and re-run the bench.
Zero retention. How would I know?
Three things, in the order of what actually protects you. First, there is nothing to keep: the request body lives for the duration of the computation, in memory, and is written nowhere. No log, no cache, no store. There is no retention policy to read because there is no data to retain. Second, we are not the sole judge of that. Tokendoo is a European entity, bound by the GDPR for all of its processing, wherever its clients are. Data minimisation is not our marketing line, it is an obligation whose breach is sanctioned, and our regulator can establish that without asking our opinion. Third, the codec and its tests are public. We will not pretend that proves what runs in production, since publishing one codebase and deploying another is always possible. It is one element of verification, not a proof, and we would rather say so.
You are adding a network hop to save a few cents.
We do not. Compression runs on your machine; there is no compression hop to pay for. The free tier is that library with tables frozen at its version. Subscribe only if you want those tables kept current as models ship.

Leave an email. We send one when the library reaches npm and PyPI.

The Rust crate is published. One email when npm and PyPI ship, nothing before, nothing after.