AI Costs Fell 1,000x. Spending Rose 320%.
The cost of running an AI query has plummeted. Per-token prices have fallen a thousandfold in three years. And yet, total enterprise spending on AI surged 320% in 2025 alone. If that contradiction doesn’t stop you in your tracks, it should. Because buried inside that paradox is the single biggest shift in business economics since cloud computing: tokens, the fundamental unit of AI work, are becoming the most contested resource in your organisation. Not hardware. Not headcount. Tokens.
This isn’t a technology story. It’s an economics story. And if you run a business, it’s about to become your story too.
The Unit of Work Nobody Budgeted For
If you’ve used any AI tool in the past two years, you’ve consumed tokens without thinking about it. Every question you ask, every document you summarise, every image you generate is measured in tokens. They’re the atomic unit of AI. And unlike a software subscription where you pay a flat fee regardless of usage, tokens make your AI costs inherently variable. The more you use, the more you spend.
That variability is what makes this different from any technology cost that came before. Deloitte’s research on AI token economics puts it plainly: unlike previous technology waves where costs were tied to subscriptions or virtual machines, AI economics now revolve around tokens, making costs inherently variable and often unpredictable. Some firms already report that AI consumes up to half of their IT spend, and cloud computing bills rose 19% in 2025 for many enterprises.
For most businesses, this crept up. You started with a ChatGPT subscription. Then a few team members got access. Then someone integrated an API. Then an AI agent started handling customer queries. Each step felt small. But the token meter was running the whole time, and nobody was watching it the way they watch the P&L.
The Paradox That Should Worry You
Here’s where it gets uncomfortable. The cost per token has been falling dramatically. Stanford’s 2025 AI Index found that inference costs for GPT-3.5-level performance dropped over 280-fold between 2022 and 2024. The price of achieving GPT-4-level performance fell from $20 per million tokens to roughly $0.40 in the same period. That sounds like brilliant news for your budget.
It isn’t. Total enterprise spending on AI surged 320% in 2025 despite those falling per-token costs. Analysis of this inference cost paradox reveals that inference costs fell 1,000-fold, but demand rose 10,000-fold. Cheaper tokens didn’t reduce spending. They unleashed it.
Economists have a name for this: the Jevons Paradox. In the 19th century, William Stanley Jevons observed that as coal use became more efficient, total coal consumption actually increased, not decreased. Efficiency made it cheaper, which made it more attractive, which created new use cases that hadn’t been viable before. The same pattern played out with email, spreadsheets, and mobile data. Now it’s playing out with AI tokens.
Microsoft CEO Satya Nadella acknowledged as much when DeepSeek launched its low-cost model, writing simply: “Jevons Paradox strikes again.” I’ve written before about why AI makes most people busier, not more productive. The Jevons Paradox is the economic engine behind that phenomenon. Every efficiency gain creates new demand that outstrips the saving.
This means your AI bill is going up, not down. The question is whether you’re getting proportional value in return.
Nations Are Already Fighting Over This
If you think the token allocation problem is just a business issue, zoom out. Governments are treating AI compute as strategic infrastructure, the same way they treat oil reserves and energy grids.
Foreign Affairs published a piece in December 2025 titled, bluntly, “Compute Is the New Oil.” The article examines how computing power has become a pillar of the US-Gulf relationship, with advanced chip sales and AI mega-campuses becoming the currency of modern diplomacy.
The numbers back up the framing. Nearly $100 billion is expected to be invested in sovereign AI compute by the end of 2026, as nations prioritise strategic independence in AI capabilities. Roughly 90 countries have now established national AI strategies or formal governance frameworks. The Gulf states, once defined by oil wealth, are repositioning as compute powers. Saudi Arabia launched “HUMAIN,” a full-stack AI ecosystem backed by its sovereign wealth fund. India launched its own sovereign large language model. The EU created OpenEuroLLM across its 24 official languages.
Sam Altman, CEO of OpenAI, has gone further. He’s proposed what he calls “Universal Basic Compute,” a system where AI-compute tokens would be distributed to all citizens as a form of universal wealth. His reasoning is that in an AI-driven future, raw computational power will be more valuable than cash. You could use your allocation, sell it, rent it out, or donate it. It’s a radical idea, but the underlying logic is sound: tokens are becoming a finite, tradeable resource with real economic value.
The geopolitical dimension matters because it sets the context for what’s coming at the business level. When nations are competing over compute allocation, deciding whether to spend tokens on defence, healthcare, or industrial policy, the same logic will cascade down to every organisation with an AI budget.
Inside Your Business, the Same Fight Is Brewing
This is where it gets personal. Inside your business, departments are about to start fighting over tokens the way they currently fight over headcount and budget.
Think about it. Your marketing team wants to use AI agents to personalise campaigns across channels. Your operations team wants AI to optimise logistics. Your customer service team wants AI to handle first-line support. Your finance team wants AI for forecasting and anomaly detection. Every one of those use cases consumes tokens. And the more sophisticated the task (complex reasoning, multi-step agent workflows, large document analysis), the more tokens it burns.
Gartner predicts that 40% of enterprise applications will feature task-specific AI agents by the end of 2026, up from less than 5% in 2025. That’s an eightfold increase in one year. Every single one of those agents will consume tokens. And the businesses deploying them will need to decide which agents get priority, which get throttled, and which get shut down entirely when the budget is tight.
This is the new cash flow question. Not “can we afford another hire?” but “where do we allocate our tokens to get the most return?” The businesses that treat this as a vague IT problem will burn through their compute budget the same way companies burn through marketing budgets without measurement. The ones that treat it as a strategic allocation problem will pull ahead.
The Allocation Problem Is the Strategy Problem
Here’s what makes token allocation particularly tricky: not all tokens are created equal. A token spent on a simple chatbot interaction that saves a customer service agent five minutes has a very different return profile from a token spent on a complex reasoning task that surfaces an entirely new revenue opportunity. But both show up on the same line item.
McKinsey’s research reveals that only 19% of enterprises report revenue increases greater than 5% from AI, while 36% see no measurable change at all. Over 80% reported no meaningful impact on enterprise-wide EBIT. The problem isn’t that AI doesn’t work. It’s that most businesses aren’t allocating their AI resources to the places where the return justifies the spend.
This mirrors a pattern I’ve seen repeatedly in marketing. The issue is almost never the total budget. It’s where the budget goes. Incentives drive behaviour, and if your token allocation is driven by whichever department shouts loudest rather than where the data shows the greatest return, you’ll get exactly the results you deserve.
The businesses getting this right are approaching token allocation the way sophisticated organisations approach capital allocation. They’re measuring return by use case. They’re right-sizing models for specific tasks rather than throwing the most expensive model at everything. They’re building what Deloitte calls “FinOps discipline,” meaning real-time monitoring, forecasting, and spend management applied to AI the same way it’s applied to every other significant cost centre.
What Smart Businesses Are Doing Differently
The organisations pulling ahead aren’t necessarily spending more on AI. They’re spending differently.
First, they’re matching model capability to task complexity. Not every query needs the most powerful (and expensive) model. A customer FAQ can run on a lightweight model at a fraction of the token cost of a frontier model, with virtually identical results. The businesses that understand this routing layer, directing different tasks to appropriately sized models, are seeing dramatically better unit economics.
Second, they’re measuring obsessively. Industry research indicates that enterprises tracking all five key operational metrics achieve 34% efficiency gains within 18 months, compared to 12% for those tracking fewer than three. You can’t optimise what you don’t measure, and most businesses aren’t measuring token consumption at a granular enough level to make intelligent allocation decisions.
Third, they’re treating AI investment as augmentation, not replacement. I’ve argued before that the future of work is humans with machines, not humans replaced by machines. The token allocation version of this principle is straightforward: the highest-return token spend is usually the one that makes a capable person significantly more effective, not the one that tries to automate away a function entirely. The fully autonomous agent that runs unsupervised burns through tokens at a rate that rarely justifies the output quality.
The Cash Flow Problem of the Future
Token consumption is about to become as fundamental to business operations as energy costs, labour costs, and capital expenditure. Gartner forecasts worldwide AI spending of $2.52 trillion in 2026, up 44% year on year. The hyperscalers are collectively investing over $600 billion in AI infrastructure. This is not a trend that’s slowing down.
For business owners, the parallel to cash flow management is exact. Cash flow problems don’t kill businesses because there’s no money. They kill businesses because money is in the wrong place at the wrong time. Token allocation will work the same way. The total compute available to your business is finite. How you distribute it across teams, tasks, and time periods will determine whether AI becomes your competitive advantage or your most expensive line item with the least to show for it.
The businesses that will thrive are the ones that start treating tokens as what they are: a scarce, valuable resource that demands the same rigour and strategic thinking as any other critical input. Not an IT cost to be absorbed. Not a technology trend to be followed. A fundamental economic resource that needs to be allocated, measured, and optimised with the same discipline you’d apply to your most important investment.
The token economy is here. The only question is whether you’re going to manage it, or let it manage you.




