The Model Was Never the Expensive Part

**The token got cheaper than the text.** Start with the number, because the number is the whole argument.

taur.ai / insights
The Model Was Never the Expensive Part
June 20264 minBy Virgil

The token got cheaper than the text.

Start with the number, because the number is the whole argument. Large-language-model inference has fallen more than 90% in roughly 18 months — from about $0.50 to $0.03 per query (TAUR's own figure, and consistent with the direction everyone else reports). Gartner projects another 90%-plus drop in inference cost by 2030 against a 2025 baseline 6. The longer arc is steeper still: a GPT-4-class workload that cost around $20 per million tokens in late 2022 runs near $0.40 per million in early 2026 — a thousandfold collapse in three years 17, tracking the roughly 10× annual decline measured across the open market for constant quality 19.

Now price the other side of the transaction — the side most builders never put on the ledger. A marketing SMS in the US runs $0.01 to $0.05 per message in 2026 3. WhatsApp business-initiated marketing templates cost $0.025 to $0.1365 per message, depending on the recipient's country, under Meta's per-message regime 8. A USSD session in Nigeria carries a flat charge of about ₦6.98 per transaction 14. International A2P termination — the wholesale rail the message actually rides — can reach roughly 15 pence per message or more 15.

Read the two columns together and the inversion is plain. At $0.03, the model is the cheap part of the interaction. The channel is the expensive part. In last-mile AI, the cost center has moved from the thing that thinks to the thing that speaks.

Where it holds, and where it doesn't. Be honest about the seams, because a skeptic will find them first. The flip is sharpest on premium delivery — WhatsApp marketing templates, USSD per-transaction billing, international SMS termination — where the per-message price sits at or above a query's cost. It is not universal: the cheapest raw-API bulk SMS still floors near $0.0083 per message 4, which undercuts a $0.03 query. WhatsApp service messages inside a customer-initiated 24-hour window are free 8. So the claim isn't "every text costs more than every thought." It's narrower and harder to dodge: the two cost curves have crossed, and they cross more often every quarter.

This is structural, not a blip. Inference rides a compute-and-competition curve that bends down by design — cheaper silicon, fiercer model competition, the 10×-a-year cadence 19. Channel cost sits on regulated rails: termination rates, spectrum, carrier rent. Those move slowly, by committee, and sometimes the wrong way. When Kenya's regulator adjusted mobile termination, it moved from 41 to 37 cents per minute 20 — a negotiated nudge, not a collapse. One line falls because engineers compete to make it fall. The other holds because a regulator and a carrier agree to let it. Two cost lines diverging on purpose. That is a direction, not a moment.

What flips when the cost center flips. The optimization that went into smaller, cheaper models now belongs on the wire. For a decade the hard question was can we afford to run the model. That question is closing. The open one is can we afford to deliver the answer. The discipline moves downstream: batching messages, compressing payloads, designing conversations to minimize round-trips, and — the decision most teams still make by habit instead of arithmetic — choosing the channel by its economics. A verbose agent that fires four SMS segments where one would do has not overspent on intelligence; it has overspent on postage. When utility-category WhatsApp messages run $0.004 to $0.0456 against marketing's $0.025 to $0.1365 8, the category you send in is a margin decision, not a formatting one.

Margin in last-mile AI is now a distribution problem. The unit economics of serving an underconnected user no longer turn on model choice. They turn on how few, how short, and over which rail you let the answer travel.

The takeaway. The model was never the moat, and it was never the bottleneck either. We spent years assuming the frontier was the next decimal of inference cost. The frontier was always the channel — and the channel just became the line item that decides whether a deployment clears. Whoever engineers the delivery layer, not whoever ships the smartest model, wins the user who can only text.

Which leaves the question worth carrying into the next build: when thinking costs almost nothing and speaking still costs real money, what is the cheapest honest way to say the answer back?

Ready to build at the edge of where AI ends and people begin?

Get Early Access