According to OpenRouter's own usage data, agentic AI token consumption overtook human token consumption for the first time on February 6, 2026. Since that crossover, agent usage has climbed 14x, from a 7-day average of 0.51 trillion tokens to 7.3 trillion. Human usage hasn't shrunk in that window, it's up 2.8x, from 0.5 trillion to 1.4 trillion. Both lines are climbing. One is just climbing at five times the rate of the other.
That gap is the actual story. This isn't humans getting replaced by agents in some zero-sum sense. It's a second, much larger form of demand appearing alongside human demand, one that didn't meaningfully exist eighteen months ago.
What changed
OpenRouter classifies traffic as agentic, human, or mixed using a weighted score built from seven signals per API key, tool-call frequency, response gaps, turn count, and similar behavioral markers, not self-reported labels. What the data is really tracking is a behavior shift: letting AI systems run autonomously over long stretches, calling tools, checking their own output, and iterating without a person in the loop for every step. That shift didn't arrive all at once. It's been building for over a year, and February 6 is simply the date the crossover became visible in the numbers, not the date the underlying change happened.

One structural detail explains part of the curve's shape: nearly 70% of agentic token volume comes from cached prompts, billed at a much lower rate than fresh tokens, because multi-turn agent work repeatedly reuses context. That means the raw token count understates how much actual compute and cost sits behind agentic workloads, but it also means the economics of running agents at scale are more favorable than a naive per-token read would suggest, which is likely part of why usage climbed as fast as it did.
Why this points straight at the open-weight model race
Anytime a cost driver explodes 14x in six months, someone starts optimizing against it, and that's exactly what's playing out in the shift toward open-weight and lower-cost models. Not every open-weight model has been capable of genuinely useful agentic work, and the ones that are capable now often weren't six months ago, which is why it can feel like a new model quietly "tips" into real usefulness almost every week. That's not marketing hype about model releases, it's a direct consequence of a market where the dominant cost driver is agent volume, not chat volume, and cheaper-per-token models suddenly matter far more than they used to.
What this means for how businesses should be planning
If your organization is still budgeting for AI the way you would for a chat tool, one person, one conversation, one cost estimate, this data says that's already the wrong mental model. The workloads driving actual token growth are autonomous processes running in the background, not people typing prompts, and the cost structure, model selection, and reliability requirements for that kind of usage are genuinely different. Planning your AI spend and tooling strategy around human-scale usage while agent-scale usage is what's actually compounding is a good way to be surprised by your bill, or by a competitor moving faster than you.
If your team is trying to figure out where agentic workflows fit into your own operations before the cost curve catches you off guard, that's exactly the kind of forward-looking work our growth strategy practice does, and our AI marketing services team can help you build agent-ready workflows instead of retrofitting them later.


Writing Team