AI Agent Token Usage Is 5x Human Volume on OpenRouter
tomshardware.com

AI Agent Token Usage Is 5x Human Volume on OpenRouter

Tech News
3 min read

Published by AINave Editorial

TL;DROpenRouter’s August figures show about five times as many agent tokens as human tokens on the platform. More than 85% of agent tokens were cached, so the volume is not a direct measure of spending, though cached context still occupies memory.

OpenRouter’s agent token lead is a platform-specific measure

AI agent token usage on OpenRouter reached 7.3 trillion tokens in an August snapshot, compared with 1.4 trillion for human users. The figures come from an a16z chart of OpenRouter data cited by Tom’s Hardware, and measure token volume on that platform, not AI use or spending across the industry. Agents had passed human usage about six months earlier, around February. The August comparison and crossover make the scale of agent traffic clear, but not how many people or dollars it represents.

OpenRouter classifies API keys as agentic, mixed or human using a seven-signal weighted score that includes tool-call rate, turn count and timing between turns. Since the February crossover, agent usage rose 14-fold, while human usage rose 2.8-fold. Mixed traffic grew 4.7-fold by Tom’s Hardware’s calculation, so how that category is assigned matters to the comparison. The reported trend also had dips in April and July; it was not a steady climb. The classification method and growth figures describe activity on one gateway, not a universal division between people and agents.

Cached context changes what token totals mean

More than 85% of agent tokens in the cited data came from cached prompts, and a16z said cached tokens accounted for nearly all the relative growth. In practical terms, agents often send context they have already used again as a task continues. A call-center consultancy’s September Claude Code logs provide one example: it reported that 96% of its input was rereading old conversation. That is a single organization’s result, not a general benchmark. The cached-token share and consultancy example help explain why raw token totals can look much larger than the amount of newly processed prompt content.

Cached tokens cost less to process than prompts handled from scratch, according to the reporting, so a fivefold token count does not mean a fivefold bill. But a lower processing cost does not make repeated context disappear from the hardware picture. Models keep stored context in a KV cache, which must reside in memory; a16z links the growth in cached context to demand for high-bandwidth memory, or HBM. The cost distinction and memory connection point to two different measures of scale: what inference costs to process and how much context infrastructure must hold.

That distinction is more useful than treating token volume as a proxy for agent economics. OpenRouter’s figures show substantial growth in one platform’s agent traffic; they do not establish the cost of a typical agent workload. The infrastructure question is whether memory capacity can keep pace with the context agents repeatedly carry forward.

FAQs

In the August OpenRouter figures, agents used 7.3 trillion tokens versus 1.4 trillion for human users, roughly five times as many. These are token volumes on one platform, not a measure of all AI usage or spending. The figures come from an a16z chart cited by Tom’s Hardware.

Sources

Latest Tech News