CFO's Guide: How to Effectively Evaluate AI Usage-Based Pricing Models?

September 7, 2026

Get Started with Pricing Strategy Consulting

Join companies like Zoom, DocuSign, and Twilio using our systematic pricing approach to increase revenue by 12-40% year-over-year.

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
CFO's Guide: How to Effectively Evaluate AI Usage-Based Pricing Models?

CFOs Guide How to Effectively Evaluate AI Usage Based Pricing Models

An AI vendor’s promise to charge only for usage can sound like disciplined purchasing. Finance pays when the business receives value, rather than buying idle capacity. Yet the unit on the invoice often decides whether the model creates operating leverage or an unpleasant surprise at quarter-end.

The CFO’s question is not, “Is $0.99 or $2 a reasonable rate?” It is, “What event creates the charge, who controls that event, how will we verify it, and what will we actually pay over three years?” Salesforce currently offers Agentforce at $2 per conversation or through Flex Credits priced at $500 per 100,000 credits. Intercom charges $0.99 for several Fin AI Agent outcomes, while a qualified lead costs $9.99. Cursor combines seat fees with included model usage and on-demand spend. Cognition’s Devin uses included quotas and credits for self-service plans, and Agent Compute Units for enterprise agreements. These are not small variations in packaging. They assign cost risk to the buyer in fundamentally different ways.[^3][^4][^5]^6

Monetizely’s position is clear: approve AI usage-based pricing only when the meter records a verified unit of business work that the customer can influence and finance can forecast. For autonomous agents, the primary meter should be an annual committed volume of outcomes or conversations, with tightly defined overages. For tools that still rely on human judgment, the primary meter should remain the named user or team, with usage limits protecting the vendor’s economics.

The headline rate often hides the financial risk being transferred

Usage pricing is not one model. A charge per token, action, conversation, resolution, qualified lead, or compute unit can each be called “pay for what you use.” Only some of those units map cleanly to a business result.

Consider the contrast in the current market. Salesforce gives buyers a choice between $2 per conversation and action-based Flex Credits, where one action consumes 20 credits and costs $0.10 at list price. Intercom’s Fin bills for a defined outcome, rather than every message sent. Cursor charges a team seat fee, then allows on-demand usage after included model capacity is exhausted. Devin’s enterprise unit is an ACU, a measure of agent compute rather than a completed engineering task.[^3][^4][^5]^6

Exhibit 1: Four current AI pricing designs and the CFO question each creates

Vendor and pricing design, as accessed September 7, 2026 Published commercial unit Main financial question for the buyer
Salesforce Agentforce[^3] $2 per conversation, or $500 per 100,000 Flex Credits. One action uses 20 credits, or $0.10. Does a conversation or action represent meaningful work, and can demand spike without a matching business benefit?
Intercom Fin AI Agent[^4] $0.99 per resolution, procedure handoff, or disqualification; $9.99 per qualification. Is the vendor’s definition of “successful” aligned with the company’s definition of value?
Cursor Teams[^5] $40 per user per month for Standard, with included usage and on-demand model charges after the included amount is used. How much spend is fixed by headcount, and how much can rise with model choice and heavy use?
Cognition Devin[^6] Self-service plans use included quota and on-demand credits; enterprise plans use ACUs at the order-form rate. Can the company forecast compute use from the engineering work it expects to assign?

The pattern matters: the closer the meter is to a verified business output, the stronger the case for variable pricing; the closer it is to model activity, the more carefully the buyer must cap and govern spend.

A conversation is not necessarily a resolved customer problem. An action is not necessarily a productive employee task. An ACU is not necessarily a merged pull request. Finance should resist treating technical activity as a proxy for value simply because it is easy for a vendor to count.

Monetizely’s 5-Step Pricing Framework begins with goals and segmentation, then moves to packaging, choosing the pricing metric, finding price points, and operationalizing the model. The sequence matters because a price should be the result of commercial choices, not the starting point. Goals and segmentation establish which buyers and business problems matter. Packaging determines which features, service levels, and terms those buyers receive. The pricing metric identifies what will be counted. Price points set the rate and commitment. Operationalization makes the model work in billing, reporting, and renewal processes. The same logic gives a CFO a disciplined way to evaluate a vendor proposal, as explored in Monetizing Agentic AI.[^1]^2

The buyer should reverse the usual pricing conversation. Rather than asking the vendor to explain a meter first, finance should establish the company’s business case, operating constraints, and required controls before accepting any unit rate.

Exhibit 2: The CFO’s reverse application of Monetizely’s 5-Step Pricing Framework

Framework step CFO question Evidence required before approval
Goals and segmentation Which business group will use the AI, and what economic problem must it solve? A baseline for ticket volume, engineering capacity, sales coverage, or another relevant workload.
Packaging Which capabilities are required for the target group, and which are unnecessary? A clear list of included channels, integrations, security features, support levels, and implementation services.
Choosing the pricing metric Does the billed unit track value, effort, or vendor cost? A written definition of the unit, exclusions, attribution rules, and a sample event log.
Finding price points What will the company pay at low, expected, and high demand? A three-year spend model including committed volume, overages, annual increases, and unused credits.
Operationalizing Can finance measure, approve, reconcile, and forecast the charges? A dashboard, invoice detail, spend alerts, and a documented monthly reconciliation process.

A vendor may have a credible product and still present a weak pricing model for a particular buyer. The failure often begins before the rate card: the package does not fit the deployment, or the unit cannot be governed by the operating team.

Cursor offers a useful contrast. Its pricing separates individual and team needs through seats, centralized administration, pooled usage for enterprise customers, and on-demand capacity. That structure recognizes that a developer tool still depends on human work and human review. The company is not buying autonomous output in the same way that a customer-service operation buys completed case resolution.^5

The Agentic Monetization Spectrum, or AMS, addresses the question that many AI rate cards obscure: how far has the product moved from assisting a person to doing work on its own? It scores an agent across three dimensions. Zero-human ability asks how much human effort remains: more than half, 20% to 50%, or less than 20%. Operational domain asks whether the agent handles a single task, an end-to-end workflow in one function, or work across functions. Output-to-cost ratio asks whether value rises roughly with compute cost, rises faster than cost, or far exceeds it. As autonomy, domain breadth, and output value increase, the pricing logic should move away from the seat and toward a measurable output or outcome.^2

For CFO review, a simple 1-to-3 score makes the implications visible. A score of 1 denotes the smaller end of each AMS dimension, 2 the middle, and 3 the larger end. The numbers are not a substitute for diligence. They force the operating team to state what the AI is actually doing.

Exhibit 3: AMS scoring points to the primary meter finance should prefer

Product archetype Zero-human ability Operational domain Output-to-cost ratio Total Primary meter that fits the work
Cursor-style coding copilot 1 1 2 4 Named user or team seat, with included usage and controlled overages
Devin-style autonomous coding agent 3 2 2 7 Committed compute capacity while reliability is proven; move toward accepted work units as results become consistent
Intercom Fin-style support agent 3 2 2 7 Committed resolutions or defined support outcomes
Salesforce Agentforce service deployment 3 2 2 7 Committed conversations or resolutions, provided the eligible event is auditable

The score explains why a pure seat model is sound for a copilot but weak for an autonomous service agent. A company buying 100 support-agent seats does not necessarily receive 100 times the customer work. A company paying for 100,000 completed and verified support outcomes can connect expenditure to a workload that management already tracks.

The same reasoning explains why outcome pricing should not become a reflex. An autonomous coding agent may perform work, but a merged pull request can be blocked by code review, test failures, or a changed product decision. Until the work unit is stable and independently verifiable, compute capacity may be the more workable commercial unit. Devin’s use of ACUs reflects that practical constraint, even though it leaves finance with a harder forecasting task.^6

A CFO should not accept a usage price until the team can model it under three demand cases. The exercise is simple, but it changes the conversation. A rate that looks attractive at forecast volume may become expensive when a product launch, seasonal demand spike, or workflow redesign doubles consumption.

The following model uses Intercom’s published $0.99 outcome price as the starting rate. It compares pure pay-as-you-go billing with a contract whose primary meter is an annual commitment of 120,000 outcomes. The commitment is priced at the same $0.99 base rate; overage is modeled at a 20% premium to show how a contract can exchange some low-volume flexibility for budget control.^4

Exhibit 4: Three-year spend under three service-volume cases

Annual outcomes Three-year outcomes Pay-as-you-go at $0.99 per outcome 120,000-outcome annual commitment plus 20% overage Finance interpretation
80,000 240,000 $237,600 $356,400 A commitment creates avoidable spend if adoption falls short.
120,000 360,000 $356,400 $356,400 Forecast usage makes the economic choice neutral before other contract benefits.
180,000 540,000 $534,600 $570,240 Overage costs matter, but the committed volume establishes a known operating baseline.

The point is not that commitment always lowers spend. It establishes a budgetable base, makes the overage rule explicit, and prevents finance from discovering the true rate only after demand accelerates.

For a customer-service deployment, the model should also test what happens when containment improves. If the AI resolves more cases but the company does not reduce outsourced volume, redeploy employees, or avoid new hiring, the system may create operational value without producing the financial return assumed in the approval case. Finance needs both measures.

A rate card does not provide control. Contract language, event data, and operating ownership provide control. Intercom’s own outcome definitions show why the detail matters: a resolution is counted when no further help is requested after Fin’s last answer, while a procedure handoff is chargeable when Fin completes a configured procedure that ends in a human or workflow handoff. Both may be useful units, but each requires the buyer to understand the trigger and challenge process before signing.^4

Salesforce similarly gives buyers a Digital Wallet to track Agentforce use across its buying models. Visibility is valuable, but a dashboard does not remove the need for spending authority, invoice reconciliation, and a shared interpretation of the billed event.^3

Exhibit 5: Five approval tests for an AI usage-based contract

Approval test What finance should demand Failure signal Required response
Unit definition A plain-language definition with examples and exclusions “Usage” is described only as model activity, credits, or requests Reprice around a business event, or impose a hard usage ceiling
Customer control The operating team can influence demand through routing, policies, and user access The vendor’s model behavior can create charges without a buyer decision Require approval controls and a no-charge rule for vendor-caused retries
Shared evidence Buyer and vendor can access the same event-level records Invoice totals cannot be traced to individual conversations, actions, or outcomes Require exportable logs and a formal dispute window
Budget protection Alerts at set thresholds and a stated rule for overage Spend can continue without a cap or executive approval Set a contractual ceiling or suspend service at the agreed limit
Renewal integrity The renewal baseline separates real demand growth from one-time implementation spikes The vendor uses peak-period volume as the default renewal commitment Rebase the commitment on trailing, normalized usage

A model that fails any of these tests may still be usable for a time-boxed pilot. It should not be scaled across the enterprise. A pilot exists to learn whether the business unit is real, not to normalize unbounded spending.

Vendor cost matters, especially for inference-heavy products. Cursor’s documentation makes that reality visible: third-party model requests are charged at underlying model API rates, and certain business usage also carries a Cursor Token Rate. Those mechanics may be commercially reasonable. They are still vendor-cost measures, not the buyer’s measure of business value.^5

Our view is that buyers should not object to a vendor protecting margins. They should insist that margin protection is not disguised as a value metric. When a vendor needs to recover compute costs, a committed pool of credits or capacity can be appropriate. When the agent has crossed into independent business work, the agreement should shift to a business unit such as a verified resolution, completed qualification, or completed workflow.

That distinction also protects the vendor. A well-defined outcome meter rewards the supplier when adoption grows because the product is producing more useful work. A poorly defined compute meter invites buyer scrutiny every time quality dips, the model retries a task, or the vendor changes its architecture.

Monetizely’s position is therefore not to seek the lowest variable rate. It is to seek the most defensible primary meter: seats for human-led tools, committed outcomes or conversations for autonomous agents, and committed compute only while output cannot yet be measured reliably.

  1. Create an AI investment portfolio, not a collection of software subscriptions. Group every AI purchase by the business capacity it affects: engineering output, service volume, sales coverage, legal throughput, or internal productivity. A portfolio view exposes duplicate spend and makes trade-offs visible.

    Make realized operating change a condition of scale. For service AI, define whether increased resolution will reduce outsourcing, avoid hiring, improve retention, or release employees to higher-value work. For coding AI, define whether faster delivery changes roadmap capacity, defect rates, or contractor spend.

    Set a separate hurdle rate for autonomous agents. A tool that drafts, suggests, or summarizes should be assessed as productivity software. An agent that independently engages customers, updates records, or completes workflow steps should meet a stronger standard for attribution, controls, and financial return.

    Assign one executive owner for demand and one for economics. The functional leader should own adoption and work quality. Finance should own forecasting, variance review, and the decision to expand the annual commitment. Shared accountability prevents either side from treating the invoice as someone else’s problem.

    Use renewal as a redesign point, not a procurement event. After two or three quarters, reassess the AMS score. An agent that began as assisted work may have become autonomous enough to justify an outcome meter. A product that has not delivered reliable output should not receive a larger usage commitment merely because the vendor requests one.

    [^1]: Monetizing Agentic AI: A Handbook for AI Transformation. https://www.amazon.com/Monetizing-Agentic-AI-Handbook-Transformation/dp/B0H7Z13VKJ/

    [^2]: Monetizely, “Step 1: Goals and Segmentation,” “Step 2: Packaging - Designing Offers That Fit,” “Step 3: Choosing the Right Pricing Metric,” “The Agentic Monetization Spectrum,” “Step 4: Finding the Right Price Points,” and “Step 5: Operationalizing Agentic AI Pricing,” accessed September 7, 2026.

    [^3]: Salesforce, “Agentforce Pricing,” accessed September 7, 2026.

    [^4]: Intercom, “Fin AI Agent Outcomes,” July 30, 2026.

    [^5]: Cursor, “Models & Pricing” and “Pricing Policy,” accessed September 7, 2026.

    [^6]: Cognition, “Devin Billing,” accessed September 7, 2026.

Get Started with Pricing Strategy Consulting

Join companies like Zoom, DocuSign, and Twilio using our systematic pricing approach to increase revenue by 12-40% year-over-year.

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.