Last month, the market's average price of a million tokens crossed below one dollar for the first time — Silicon Data's LLM Token Expenditure Index closed August at $0.97, down more than half from its $2.05 May peak. This week Anthropic cut Claude Haiku 5.5 prices 90% to match GPT-6 Luna at $0.10 in, $0.50 out. This isn't a discount war. It's the end of the model as a durable moat.
The unit economics are breaking
Epoch AI found the cost of hitting the same benchmark fell 47% a quarter since 2023; GPT-5.6 Luna now matches o3's GPQA Diamond score for $0.0004 a problem — a 725x drop in 18 months. On the commercial side, JPMorgan tracked a 47% jump in August token volume on OpenRouter against just 7% growth in dollar spend. Goldman Sachs warned that two years of AI valuations were built on 'more tokens equals more revenue' — August was the first month to break that assumption.
The harder truth: we keep paying 6x
MIT research from Frank Nagle found closed models still capture 80% of usage and 96% of revenue while costing six times more than open models that reach 90%+ of their performance — and close the gap within months. Decagon's numbers show open-source share of enterprise LLM spend actually fell to 11% from 19% last year. We are paying a premium for a shrinking lead, mostly out of habit.
Where the moat moves
If raw inference is commoditizing, the durable assets sit around it: proprietary context, workflow integration, memory, distribution, and the controls that keep agents off systems they shouldn't touch. That is where defensible value accumulates — not in the API call.
The takeaway
Stop modeling budgets on token costs that just fell more than half in a single quarter. Redirect the premium you pay for frontier models into the data, context, and governance you own. In 18 months the model will be the cheapest layer in your stack — and the only thing you'll still control is everything around it.