AI is using more and more AI. According to data from OpenRouter analyst Peter Walker, February 6, 2026 may have been the last day humans consumed more tokens than AI agents.

Agentic token usage has grown 14x since then. Over the same period, human usage is up just 2.8x. Growth on the machine side is running roughly five times faster than on the human side.

The two rates

  • Agent usage — up 14x since February 6, 2026.
  • Human usage — up 2.8x over the same period.
  • Share of agent tokens served from cache — roughly 70 percent.

Why it grows so fast

The reason lies in how agents work. Agents increasingly operate on their own over longer stretches, spinning up additional AI processes along the way. An agent handed a task makes further model calls for subtasks, and each of those carries its own context.

That produces non-linear growth: even with a flat user count, token consumption compounds as agents' autonomous stretches get longer.

But the bill is not growing at the same rate

The figure should not be read at face value. Nearly 70 percent of agent token usage comes from cached prompts, which are billed at much lower rates.

The reason is structural: an agent resends the same system instructions, the same tool definitions and the same project context at every step. That repeating portion can be cached. So while raw token counts rise 14x, actual costs do not rise proportionally.

The distinction matters, because "token usage grew 14x" gives the impression that industry costs grew 14x. They did not.

The limits of the data

Where the measurement comes from is worth noting. OpenRouter skews toward open-weight models, which tend to be less token-efficient than models from the major labs. The absolute figures here may not represent the whole industry.

The trend, however, likely looks similar at the major labs. Token inflation had already started with reasoning models, which think longer before they respond — even when they should not.

One more caveat: this is one platform's data read by one analyst. There is no independent industry-wide confirmation. The direction of the trend is credible, but the exact ratios deserve caution.

Why it matters

This number changes how the industry's revenue figures should be read. When a model provider's token volume grows, it used to imply "more people are using it." It no longer does: much of the volume consists of calls initiated by other software.

A second consequence concerns pricing. Agent workloads are structurally different from human ones: more repetitive, more cacheable, more predictable. It is no accident that providers have separated cache pricing so aggressively — they know where the volume comes from.

Third, and perhaps most importantly: it means the largest customer in this economy is another machine. Anyone forecasting demand needs to watch agent autonomy duration rather than human user counts.