AI Token Costs Plunge Reshapes Enterprise AI Economics and Margins
forbes.com

AI Token Costs Plunge Reshapes Enterprise AI Economics and Margins

Tech News
4 min read

Published by AINave Editorial • Reviewed by Ramit

TL;DREnterprise AI spending cuts are driving token prices down, squeezing model provider margins and reshaping profitability across the generative AI stack. Hybrid model usage and tokenomics are key strategies.

AI token costs are plunging as enterprises aggressively cut spending by switching to cheaper models, squeezing margins for premium providers like OpenAI and Anthropic while benefiting chipmakers and Chinese LLM suppliers. This shift has given rise to "tokenomics" as a discipline for optimizing AI budgets, and hybrid model usage is becoming the norm for cost-conscious teams.

What happened

Companies are clamping down on AI spending by opting for lower-priced models, even for complex tasks. The potential cost savings are considerable: building a web browser from scratch costs more than $10,000 on OpenAI's GPT-5.5, whereas a combination of Cursor's Composer and Anthropic's Opus 4.8 gets the job done for $1,339, according to Cursor research featured by the Wall Street Journal. That is an 87% savings.

This price competition has brought "tokenomics" to life as a new field that helps businesses track rapidly changing token prices and optimize AI budgets. By mid-2026, Chinese suppliers of less expensive large language models including DeepSeek, Qwen, GLM, Kimi and MiniMax had won 46% market share.

Profitability is concentrating in hardware and cloud services. Nvidia's net margin is 63%, TSMC's is 46.5%, Micron's is 56%, and SanDisk's is 34.2%. Among cloud providers, AWS operates at a 37.7% operating margin and Google Cloud at 35.6%. Meanwhile, OpenAI at roughly $25 billion annualized revenue runs about a 33% gross margin, dragged down by consumer ChatGPT. Anthropic, which gets about 80% of its $30 billion run-rate from enterprise use, saw inference margins climb from 38% to over 70% during 2026 due to the popularity of Claude Code, according to SemiAnalysis.

Why AI builders should care

The economics of AI are shifting from model quality alone to cost per token and inference margins. Investors and builders should monitor token prices and inference margins to identify winners and losers in the AI value network. With IPOs looming for OpenAI and Anthropic, the diminishing profit potential of the AI chatbot market does not bode well for retail investors. The chip and cloud services providers enjoy the highest profit potential while the model providers are being squeezed the most.

Heavy agentic use is forcing pricing changes. Cursor replaced its $20 flat plan with credits-based repricing. Uber used up its entire 2026 budget in four months. Salesforce changed its Agentforce pricing models three times in eighteen months. These examples show that flat-rate pricing is unsustainable when token consumption varies wildly.

Practical implications

Enterprises can reduce AI spend by pairing cheaper models for routine tasks with premium models for critical tasks. Telnyx, whose AI agent costs were $100,000 a day using Anthropic, switched to open models. The company's 1,400 agents now use Chinese startup Z.AI's product which costs around $100 per agent per day. Anthropic's most powerful model, Fable, acts as a conductor that plans out work while open-weight models do the implementation. OpenAI's Sol handles a review of what the open-weight models produce.

Budgeting should include token-cost dynamics and potential variance in token prices as a core input. The discipline of tokenomics helps decision-makers track the rapidly changing price of tokens and the cost of providing those tokens to users. Asset-light suppliers and cloud/hardware providers may outperform pure-play model providers as costs deflate.

Caveats

Evidence across sources varies in depth; some statements are syntheses from headlines or related reporting rather than a single detailed source. Numbers like exact token-cost reductions and margin percentages are subject to revision as the market evolves. The parent article relies on company-provided research and analyst estimates, which may not reflect actual enterprise costs in all cases.

Sources

Latest Tech News