What is Token-Based Pricing? Understanding the New AI Economy

September 7, 2026

Get Started with Pricing Strategy Consulting

Join companies like Zoom, DocuSign, and Twilio using our systematic pricing approach to increase revenue by 12-40% year-over-year.

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
What is Token-Based Pricing? Understanding the New AI Economy

What Is Token Based Pricing Understanding the New AI Economy

A procurement leader opens an AI vendor’s invoice and sees a familiar problem in an unfamiliar unit: millions of tokens. A product leader sees the same unit from the other side, as a fast way to protect gross margin while models, prompts, and usage patterns keep changing. Both are asking the same question: should tokens be the price customers see, or only the cost the company manages?

The distinction matters because AI has introduced a real variable cost into products that were once sold mostly through predictable seats and annual subscriptions. Yet a token is a unit of model processing, not a unit of business value. A customer does not wake up wanting 20 million output tokens. They want a support case resolved, a sales lead qualified, a contract reviewed, or a developer made faster.

Monetizely's position is clear: token-based pricing should be the primary customer price only when the buyer is deliberately purchasing model capacity, usually through an API. For most B2B AI applications, tokens should sit underneath the offer as a cost and risk meter, while the customer-facing price should track the human user or the work the agent completes.

Tokens reveal the cost of intelligence, not the value of work

A token is a chunk of text or other model input that an AI system processes. Models consume tokens when they read a prompt, retrieve a document, produce an answer, reason through a task, or call a tool. Providers commonly price input, cached input, and output at different rates, while some tools create separate charges. OpenAI also notes that visible response length can understate total usage because cached and reasoning tokens may appear in usage data.

That structure makes tokens an excellent meter for the model provider. It ties revenue closely to compute demand. For an application vendor, however, the same structure creates a commercial problem: the customer’s bill can move because of system prompts, retrieval design, tool calls, model routing, or a longer conversation history - choices the customer may neither see nor control.

The anatomy of a token bill explains why a simple price per million tokens rarely produces a simple customer experience.

Exhibit 1: A token bill contains more than the prompt and answer

Bill component What commonly creates it Why it changes cost What an operator should monitor
Input tokens User prompt, system instructions, retrieved documents, tool schemas Longer context raises spend before the model produces an answer Tokens per request and retrieved-content size
Cached input tokens Reused prompts, policies, instructions, and context Caching can materially reduce repeat-context cost Cache-hit rate and cache-write cost
Output tokens Written response, code, structured data, tool arguments Output is often priced well above input Output length and completion limits
Reasoning tokens Internal model work on complex tasks Usage may exceed what the user sees on screen Cost per completed task, not cost per visible word
Tool and workflow charges Search, grounding, computer use, databases, or third-party actions Some services bill beyond token volume Cost per tool call and cost per workflow run

OpenAI’s documentation identifies input, output, cached-input, and reasoning-token usage as separate reporting categories, while Google Cloud states that repeated context in a live session can be processed and billed on every turn. The practical lesson is straightforward: token volume is a precise measure of model activity, but a noisy measure of the customer’s intent.

Token pricing is not one universal rate. Each model has its own input, output, cache, context-window, and service-tier economics. Output typically costs more because generating text or code consumes more inference resources than reading a prompt. Caching changes the picture again.

The following public list prices show why an AI product team cannot treat “one million tokens” as a stable cost unit.

Exhibit 2: Token rates vary sharply by model and token type

Public list prices checked September 7, 2026, in U.S. dollars per one million text tokens.

Provider and model Input Cached input Output What the rate card signals
OpenAI GPT-5.6 Sol $4.00 $0.40 $20.00 Premium reasoning and professional work create a large output-cost exposure. (developers.openai.com)
Anthropic Claude Sonnet 5 $2.00 $0.20 $10.00 Cached context is far cheaper than fresh input, rewarding disciplined prompt design. (docs.anthropic.com)
Google Gemini 2.5 Flash $0.30 $0.03 $2.50 High-volume workloads can support low unit costs, but output still costs far more than input. (cloud.google.com)

A workload with 10 million input tokens and 2 million output tokens would produce a model charge of $80 on GPT-5.6 Sol, $40 on Claude Sonnet 5, and $8 on Gemini 2.5 Flash at those published rates. Those figures do not establish that the models deliver the same quality, reliability, or task-completion rate. They show why a lower token rate alone is not a pricing strategy.

A product team therefore needs to manage cost per completed task. A more capable model may use fewer turns, need less human correction, and generate less output. OpenAI makes that point directly in its model guidance: a model with a higher per-token price can still have a lower estimated cost per task if it uses materially fewer output tokens.

The market already shows the difference between a cost meter and a selling meter. Vendors that sell raw model access commonly charge on tokens. Vendors that sell a complete application commonly use seats, messages, actions, conversations, or outcomes.

That pattern is not cosmetic. Each meter tells the buyer what they are purchasing and which party bears usage risk.

Exhibit 3: B2B AI vendors price the customer’s job, not merely model consumption

Public list prices checked September 7, 2026.

B2B SaaS product Primary customer-facing meter Published price What the meter says
ChatGPT Business Named seat $20 per user per month on annual billing, or $25 monthly The buyer purchases a secure AI workspace for employees. (openai.com)
Cursor Teams Standard Named seat $40 per user per month The product is sold as a developer tool, even though usage and model choice affect Cursor’s underlying cost. (cursor.com)
Microsoft Copilot Studio Messages $200 per 25,000 messages per month The buyer purchases agent interactions across a tenant, not raw model tokens. (microsoft.com)
Salesforce Agentforce Conversations or actions $2 per conversation, or $500 per 100,000 Flex Credits The buyer pays for agent activity in customer and employee workflows. (help.salesforce.com)
Intercom Fin AI Agent Outcome $0.99 per outcome The buyer pays when Fin resolves, qualifies, disqualifies, or completes a defined handoff. (intercom.com)

The pattern is decisive. Tokens are the external meter when the customer is building with the model; seats, interactions, and outcomes are stronger meters when the customer is buying a product that uses the model on their behalf.

A customer-support leader can forecast resolutions. A sales leader can forecast qualified leads. A CIO can forecast employee seats. Few can forecast how many tokens a retrieval system will add after a vendor changes its system prompt or expands an agent’s tool set.

Monetizely's 5-Step Pricing Framework puts the pricing metric in its proper place. It begins with goals and segmentation: what the company must achieve and which buyers have meaningfully different needs. It then moves to packaging, where offers, features, services, and terms are shaped for those segments. Only then does the company choose the pricing metric, set price points, and operationalize the model through metering, entitlement rules, billing, reporting, and sales execution. The sequence matters because a token rate chosen before those decisions often solves a finance problem while creating a customer problem. The full treatment appears in Monetizing Agentic AI.

Applied to token pricing, the framework forces a harder and better set of decisions.

Exhibit 4: The five decisions determine whether tokens belong on the invoice

Pricing decision Question leaders must answer Implication for token pricing
Goals and segmentation Is the priority market adoption, ARR growth, margin protection, or enterprise expansion? Which buyers have different willingness to pay? A startup developer and a regulated enterprise should not inherit the same allowance and overage logic.
Packaging Are buyers purchasing a tool, a team workflow, or a delegated job? Include tokens inside a seat package for human-led tools; attach an action or outcome meter to delegated work.
Pricing metric What unit can the buyer understand, forecast, and connect to value? Use tokens as the primary meter for model capacity, not as a default for every AI feature.
Price points What price captures value while protecting gross margin? Set included usage from observed cost per successful task, then price overages deliberately.
Operationalization Can usage be measured, attributed, explained, invoiced, and governed? Record input, output, cache, reasoning, model, tool, customer, and workflow data before launch.

Monetizely’s position is that Step 3 - selecting the metric - carries the greatest strategic weight, but it cannot be done well without the work in Steps 1 and 2.

The Agentic Monetization Spectrum, or AMS, helps answer a narrower question: when should an AI product move away from seats and toward a measure of completed work? It scores an agent along three dimensions: zero-human ability, meaning how little human work remains; operational domain, meaning whether the agent handles a task, a business-function workflow, or work across functions; and output/cost ratio, meaning whether the value created rises only with compute or greatly outpaces it. The first dimension matters most because a seat remains a credible anchor when a person still does most of the job.

For the table below, 1 means small or linear, 2 means medium or inflecting, and 3 means large or exponential. The scores are Monetizely’s commercial assessment of the product archetype, not vendor performance rankings.

Exhibit 5: Higher autonomy pushes the selling meter toward work completed

Product archetype and example Zero-human ability Operational domain Output/cost ratio AMS read Recommended primary meter
Employee AI workspace - ChatGPT Business 1 3 2 Human remains the work anchor Seat
Coding assistant - Cursor Teams 2 2 2 Human delegates and reviews meaningful work Seat, with included use and controlled overage
Customer-service agent - Salesforce Agentforce 3 2 2 Agent can conduct a bounded customer workflow Conversation or action
Resolution agent - Intercom Fin 3 2 2 Agent produces a measurable service result Outcome

The scores support a clear commercial architecture: the more fully an agent performs the work, the less defensible a seat becomes, and the less useful a token rate becomes as the buyer-facing price.

Intercom’s outcome model is particularly instructive. A customer does not pay more because the question needed a longer knowledge-base search or a more expensive reasoning path. The customer pays when an outcome occurs, once per conversation. That puts model-selection and workflow-efficiency risk with Intercom, where it can be managed, while giving the buyer a bill tied to a business event.

A company can use more than one meter without confusing the customer, provided each one has a distinct role. Monetizely does not recommend a vague blend of seats, tokens, credits, and outcomes. We recommend a named primary meter, supported by internal token controls.

Exhibit 6: The correct architecture starts with the buyer’s purchase

What the buyer is purchasing Primary price Role of tokens Commercial rationale
Model access for developers Input and output tokens External revenue meter The buyer controls prompts, model calls, and application design.
Employee productivity tool Seat Internal COGS meter and fair-use guardrail The buyer is equipping people, not buying inference capacity.
Agent handling a defined workflow Action, message, or conversation Internal margin and capacity meter The buyer can forecast workflow demand more easily than model use.
Agent delivering a verified business result Outcome Internal cost and quality meter The vendor earns more only when the customer receives a defined result.

This architecture also clarifies what to put in contracts and billing systems before the first enterprise renewal:

Operational discipline is not back-office cleanup. Monetizely notes that agentic pricing must connect token metering, tier entitlements, usage rating, credit balances, and understandable invoices.

Token pricing has earned a permanent place in the AI economy. It makes variable inference cost measurable. It supports API businesses. It gives product and finance teams a direct view of margin by model, workflow, and customer.

But it should not become the default answer to every AI pricing question. A token invoice transfers technical uncertainty to the buyer. A well-designed commercial model keeps that uncertainty with the party best able to manage it: the vendor that chooses the model, designs the prompts, controls retrieval, and improves the agent.

Our committed position is therefore simple: sell tokens when customers are buying compute. Sell seats when people are using a tool. Sell actions, conversations, or outcomes when an agent is doing the work.

  1. Separate the product price card from the model-routing policy. Product leaders should be free to improve models, prompts, and caching without reopening every customer contract.

  2. Build a margin view around successful work. Report cost per resolved case, qualified lead, completed document review, or accepted code change alongside cost per token.

  3. Create a portfolio rule for AI features. Decide which offers are API capacity, human copilots, and delegated agents, then assign one primary meter to each category.

  4. Make token telemetry a board-level operating metric. Track concentration risk by customer, model, and workflow before a small number of heavy users reshape gross margin.

  5. Reward sales teams for the right commitment. Compensation should favor durable seat adoption, workflow volume, or outcome commitments - not token consumption for its own sake.

Footnotes

  1. Monetizing Agentic AI. https://www.amazon.com/Monetizing-Agentic-AI-Handbook-Transformation/dp/B0H7Z13VKJ/
  2. OpenAI, GPT-5.6 Sol model pricing and token-usage documentation, accessed September 7, 2026. (developers.openai.com)
  3. Anthropic, Claude API model-pricing documentation, accessed September 7, 2026. (docs.anthropic.com)
  4. Google Cloud, Gemini and Vertex AI agent-platform pricing, accessed September 7, 2026. (cloud.google.com)
  5. Intercom, Fin AI Agent pricing and outcome definitions, accessed September 7, 2026. (intercom.com)

Get Started with Pricing Strategy Consulting

Join companies like Zoom, DocuSign, and Twilio using our systematic pricing approach to increase revenue by 12-40% year-over-year.

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.