
AI Model Churn Is Rising as Open-Weight Token Use Grows
Published by AINave Editorial • Reviewed by Ramit
Vercel’s AI Gateway Production Index points to a striking mismatch in AI model use: open-weight models accounted for 56% of token volume in its September index, but just 14% of token spend. The figures, reported by TechRadar, also suggest that usage is turning over quickly, with most tokens going to models available for less than four months. Vercel’s index figures and model-turnover finding
Usage is moving toward newer models
The index describes a market where newer releases displace older ones. TechRadar cites Anthropic’s Fable 5 losing out to the company’s Opus 5 and OpenAI’s Astra. That is evidence of change in the index, not proof that every company has stopped using older models or that those models no longer work for specific tasks. The examples of model turnover
The practical issue is how teams manage commitments as model options change. If a purchase or workflow is tied to a particular older model, newer alternatives may make that commitment less useful. But the article does not quantify stranded spend, or show that churn raises every company’s overall AI costs.
Falling token prices do not equal lower total bills
The September article reports that the index’s average price per token fell 23.2% for a third consecutive month. Vercel’s data is presented as reflecting several influences, including falling prices across the industry. The reported change in average price per token
That metric matters, but it is not a company-level savings figure. A lower average token price does not establish that any particular business’s total AI bill fell, especially as its usage and model mix may also change.
Open-weight models separate volume from spend
The open-weight figures help explain why the market shift matters. TechRadar reports that these models’ share of token use rose from 10% in December 2025 to 13% in April 2026, before reaching 56% of token volume in the September index. Their share of token spend in that September index was 14%. The reported open-weight usage and spend shares
The gap indicates that token volume and spending are telling different stories. Open-weight models can be self-hosted or run on cloud servers, and the article points to that flexibility as part of their appeal to enterprises. But the index figures alone do not establish what a particular deployment costs or whether it matches a closed model’s capabilities.
Meta Llama 4 is one example cited in the report. It also highlights TypeSafe AI’s Jev, which takes data rather than conversation as input and which Vercel describes as the fastest-adopted model since the index launched. That different input type is a reminder that “model adoption” can cover products serving distinct workflows, not just interchangeable chat assistants. The examples and Jev’s described input type
The signal is not that every team should switch models. It is that model choice, usage share and token pricing are moving together, while the index does not reveal how those shifts translate into a particular organization’s workload or total bill.



















