The Same Sentence Costs More in Shona
The number everyone quotes is true.

The number everyone quotes is true. AI inference cost fell more than 90% in 18 months, from $0.50 to $0.03 per query (TAUR canonical figure). It may be the most important price chart in the industry. It also has a footnote nobody prints. The chart counts tokens, and a token is not a unit of meaning. It is whatever scrap of text the tokenizer happened to learn.
Your user does not pay per token. They pay per sentence. And the same sentence costs a different amount depending on the language it arrives in.
A vocabulary is a census
Before a model reads anything, a tokenizer cuts the text into subword pieces drawn from a fixed vocabulary. That vocabulary is learned from training data: find the pieces that occur most often, merge them, repeat. Byte Pair Encoding, one of the two dominant methods, is essentially that loop 2. Common English words stay whole. Everything else goes under the knife.
Now feed it isiZulu. Ngiyakuthanda, "I love you," is one word assembled from parts: ngi- (I), -ya- (present tense), -ku- (you), -thand- (love), -a. Shona works the same way. Ndichakudzidzisa, "I will teach you," packs subject, tense, object, root and a causative ending into a single string. Bantu grammar is a masterclass in compression. A tokenizer that never saw enough of it undoes the work, splitting the word into whatever fragments it happens to know. Those are rarely the real parts.
Modern tokenizers don't fail outright. Their fallback mechanisms guarantee every word gets some encoding 2. It just gets a longer one. The vocabulary is a census of whose text was in the room.
The tax, itemised
The measure is fertility: tokens per word. English on GPT-4o comes in at about 1.3 7. The premium is a language's fertility divided by English's on the same tokenizer. A premium of 2.38× means that language costs 2.38× more in tokens, and therefore in API spend and latency 7.
A study published in June 2026 measured this across 19 African languages: 16 in Latin script, 2 in Ethiopic and 1 in N'Ko 17. On cl100k_base, the tokenizer behind a generation of OpenAI models, the mean African premium is 3.31× 1. Amharic reaches 9.68× 1. Read those figures as a price list. The same clinic reminder or crop-price query is billed at 3.31× the English rate on average, and far more in Ge'ez script.
The bill is only the first line. A fixed context window holds fewer words, leaving less room for conversation history or retrieved documents. Every extra output token burns more compute, so replies arrive later, and on a USSD session with a hard timeout, later can mean never. Mobile data already costs 10–20% of monthly income (TAUR canonical figure). A deployment that passes this cost down lands it on the people least able to pay.
One honest correction to the usual story: the same study finds that a language's premium and the model's accuracy in it are unrelated. Swahili, for instance, has a low premium and high accuracy 1. The tax is a cost problem, not a quality signal. That turns out to be the good news.
Why the market hasn't fixed it
Nobody designed this out of malice. Tokenizers are trained on text from the markets that pay, and changing a tokenizer means retraining the model built on top of it. Labs do that rarely, and only for their biggest markets. Earlier research already argued that charging by the token "is flawed and extremely unfair towards speakers of the over-fragmented languages" 3. Pricing pages still don't publish token counts per language. The incentive model does exactly what it was built to do. That is the whole problem.
Who is closing the gap
Part of the fix is already on the shelf. A leaderboard built from the study ranks current tokenizers by mean African premium: Gemma 4 at 2.38×, Llama 4 at 2.46×, BLOOM at 2.59×, Qwen3 at 2.63× and o200k_harmony at 2.7× 7. Switching from cl100k_base to Gemma 4 cuts tokenization overhead by 28% across African-language deployments. For Amharic, the premium drops from 9.68× to 2.47×, a 74% cut in inference cost 1. Because a lower premium doesn't cost accuracy, the authors call tokenizer choice "a free optimization variable" 1. Free is a rare word in this industry. Take it.
The other routes carry real costs:
- Byte-level and tokenizer-free models cover every language, but they produce much longer sequences, so cost and latency climb 6.
- Translating to English and back saves tokens but adds a second model call and more latency. It can also strip out the nuance that made a local-language service worth building.
- Caching common intents works for the crop-price query. It does not work for a real conversation.
What to do on Monday
- Measure before you price. Run real messages, in your users' actual language, through every tokenizer you're considering. The premium is measured, not guessed.
- Budget per conversation, not per token. A per-token price is quoted in a currency your users never spend.
- Treat the tokenizer as a product decision. Choosing a model means choosing whose sentences are cheap.
The cost collapse is real. It reaches your users only if you check what their language costs to process. Until then, the receipt arrives itemised, in fragments.
Ready to build at the edge of where AI ends and people begin?
Get Early Access