The Answer Is Cheaper Than the Envelope
The token costs less than the text message it travels in.

The token costs less than the text message it travels in. That is the whole story, and most of the industry has not noticed yet.
Put two current numbers side by side. A full language-model round-trip — a short question in, a useful answer out — now runs to roughly three cents of inference, on the trajectory that took the marginal cost of a query from about fifty cents to about three over eighteen months. Sending a single international SMS to carry that answer costs more: the average price for a brand to send one text internationally rose from $0.033 to $0.0646 between 2021 and 2023 9. The intelligence is cheaper than the envelope. We spent a decade assuming the reverse.
Two curves, one crossing
The point is not that AI got cheap. It is that two cost curves moving in opposite directions finally crossed.
On one side, inference collapsed. For constant quality, LLM pricing has dropped roughly 10x per year — a factor of about 1,000 in three years 17, faster than PC compute during the microprocessor era or bandwidth during the dotcom boom 6. What cost $30 per million input tokens and $60 per million output tokens in early 2023 now buys GPT-4-class performance at around $0.40 per million tokens, and open-weight models like Llama 3.2 3B run at $0.06 per million 3. Reasoning quality became a commodity input.
On the other side, carrier delivery barely moved — and where it moved, it moved up. Mobile carriers have raised A2P SMS termination rates by 15% to 75% since 2021 7, enough that regulators noticed: Ofcom ran a multi-year market review into wholesale A2P pricing and closed it only on 31 October 2025 1. The pipe is not a falling cost. It is a defended one.
The cost center flipped
For years the implicit model was simple: intelligence is the scarce, expensive thing; distribution is solved plumbing. That assumption is now backwards.
The model is the commodity. The channel is the constraint. When a round-trip of reasoning costs less than the message that delivers it, the unit economics of serving a low-ARPU user are no longer set by who has access to a good model — everyone does — but by who controls the last-mile delivery economics. The meter, not the model, decides whether you can profitably answer a farmer's question. This is the inversion, and it reorders the whole problem.
Where it holds, and where it doesn't
The honest part: SMS is the expensive floor, not the cheap one, and the inversion is channel-dependent. A skeptic should push here, so let me push first.
Per-message A2P fees do not fall with scale the way tokens do — carriers are now layering on per-message pass-through charges, with T-Mobile charging $0.0025 per inbound message for 10DLC traffic in the US 18. USSD, the channel that actually reaches feature phones, carries its own toll: through an aggregator like Africa's Talking, a dedicated code runs KES 17,400 in monthly maintenance per network, on top of an end-user telco charge and a per-session infrastructure fee 14. Voice minutes and the aggregator-plus-regulatory tax stack the same way. None of these are inference costs. All of them are delivery costs, and they are sticky.
So the claim is not "delivery is free." It is the opposite: delivery is now the part that isn't getting cheaper, which is precisely why it has become the thing that matters.
Optimize the round-trip, not the model
If intelligence is nearly free and delivery is the cost, the engineering discipline inverts with it.
You stop optimizing the model and start optimizing the round-trip. Fewer, denser messages — every segment you save is worth more than every token you shave. Channel arbitrage becomes a first-class decision: route by cost-per-resolved-query, not by habit, choosing WhatsApp, SMS, or USSD per interaction. Cache the predictable, because re-asking the model costs less than re-sending the answer. And follow the channel users already trust with money — USSD carried 63.5% of all mobile money transaction volume in Africa in 2024 16, across 1.1 billion registered mobile money accounts in Sub-Saharan Africa 16. The rail that moves the cash is the rail that should move the intelligence.
The frontier moved. It used to sit inside the model, where the smartest people and the largest budgets were already pointed. It now sits at the meter — the per-message, per-session, per-minute economics of reaching someone over the channel they already hold. Distribution was always the real frontier. The cost curves just made it undeniable.
Ready to build at the edge of where AI ends and people begin?
Get Early Access