What is Token-Based Pricing?
Imagine if your electricity bill was based not just on how long you kept the lights on, but on the complexity of thoughts you had in each room. That’s the new reality of AI costs. Instead of a flat monthly subscription, many AI services from providers
like OpenAI, Google, and Anthropic charge based on “tokens”. A token is the basic unit of information an AI model processes—think of it as a piece of a word, with one English word being roughly 1.3 tokens. Every part of an AI interaction, from the user's query (input tokens) to the AI's response (output tokens), consumes this new currency. This usage-based model directly ties the cost to the computational effort required, a fundamental shift from the predictable, seat-based licenses of traditional software.
The Budgeting Black Box
The primary challenge for corporations is the unpredictability. A finance department can easily budget for a fixed software license, but forecasting token consumption is like predicting the weather a year out. Costs can fluctuate wildly depending on usage. An employee running a simple summarisation task might use a few hundred tokens, while a developer accidentally creating an infinite loop in an AI-powered application could rack up a bill of thousands in minutes. This potential for “bill shock” is a major concern, as costs are no longer capped. Furthermore, not all tokens are priced equally; output tokens, which require the AI to generate new information, can be three to five times more expensive than input tokens. This complexity makes traditional ROI calculations difficult, with one study finding only about half of organisations feel confident they can even evaluate the return on their AI investment.
The Rise of AI Governance
In response, a new corporate discipline is emerging: AI governance, or “tokenomics”. This isn't just about cutting costs, but about spending smarter. Companies are implementing strategies to manage and optimise their AI spend. A key tactic is model routing. This involves using powerful, expensive models like GPT-5 or Claude Opus only for complex reasoning tasks, while routing simpler jobs like data extraction or basic classification to smaller, cheaper models. For example, a high-volume task might be sent to a model that costs ten times less than a flagship one, reserving the premium models for work that truly needs their power.
New Strategies for Cost Control
Beyond model routing, companies are adopting several other cost-control measures. Prompt engineering—teaching employees to write shorter, more efficient queries—is crucial, as every word adds to the token count. Caching common queries to avoid repeatedly paying for the same answer is another quick win. On a more technical level, API gateways are being used to set rate limits and quotas, providing a safety net against runaway usage. This creates guardrails that prevent a single user or faulty application from draining the budget. The goal is to move from reactive bill-paying to a proactive strategy where AI consumption is tracked, managed, and aligned with specific business outcomes.
Linking Cost to Value
Ultimately, experts argue that tracking token consumption alone is a vanity metric; it shows activity, not value. The most advanced companies are building systems that connect AI costs directly to business outcomes. Instead of AI spending being a nebulous item in the IT budget, it’s being allocated to the departments that benefit, such as marketing, engineering, or customer service. This requires a new level of collaboration between finance, technology, and business leaders to create a shared view of AI cost, usage, and return. By treating AI spend as a direct cost of delivering a service or product, companies can make more informed decisions about where and how to deploy this transformative technology for maximum economic impact.











