
Frameworks, core principles and top case studies for SaaS pricing, learnt and refined over 28+ years of SaaS-monetization experience.
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
Join companies like Zoom, DocuSign, and Twilio using our systematic pricing approach to increase revenue by 12-40% year-over-year.
The pricing question for agentic AI is no longer whether customers will pay for intelligence. They already do. The harder question is what customers should pay for when software can make decisions, call tools, update systems, and complete work that once belonged to an employee.
That choice will shape more than revenue. It determines whether a product feels like a useful assistant, a digital worker, or an unpredictable infrastructure bill. It also determines whether falling inference costs expand margin or force repeated price cuts. A support agent that resolves a customer’s delivery problem should not be priced like a text box. Nor should a coding assistant that still needs a developer to review every change be priced like a fully autonomous engineer.
Monetizely’s position is clear: autonomous B2B agents should be priced with a platform commitment and a primary meter tied to a completed, auditable business outcome. Seats remain right for human-led copilots; compute credits remain useful while reliability is immature. Neither should be the long-term commercial center of an agent that independently makes and executes business decisions.
The temptation in AI is to begin with the meter. A company sees high model costs and reaches for credits. Another sees a customer-service outcome and jumps to pay-per-resolution. Both moves can be premature.
Monetizely’s 5-Step Pricing Framework puts the decisions in the order customers actually experience them. As Monetizing Agentic AI argues, pricing works only when each choice supports the one before it. The five steps are:
This order matters more in agentic AI because value is uneven. A solo developer may use an agent to accelerate code review. A 5,000-person engineering organization may use the same product to enforce security policies, route work, and reduce backlog. The underlying model may be similar, but the offer should not be.
Cursor provides a useful current example. Its public plans separate individual and team buyers through administration, security, and pooled usage, while preserving a familiar per-user buying motion. Cursor lists Pro at $20 per month and Teams Standard at $40 per user per month, with usage beyond included allowances billed separately as needed. The lesson is simple: do not package an agent around how impressive its model is. Package it around what each buyer must control, deploy, and defend.
Public pricing pages now show a market in active experimentation. Several vendors combine a recurring base with usage, but they use different units because their products sit at different points on the path from assistance to autonomy.
The table below establishes the reference set. It also shows why “AI pricing” is too broad a category to guide a commercial decision.
The pattern is not a random collection of pricing experiments. Products closer to a human’s daily workflow retain seats; products that can complete a defined unit of work move toward output, resolution, action, or consumption.
That distinction separates a commercial model from a billing mechanism. Credits can tell a vendor how much computation ran. They rarely tell a buyer whether the agent created value. A completed support resolution, qualified prospect, reviewed pull request, or processed claim comes much closer.
The Agentic Monetization Spectrum, or AMS, makes the decision more disciplined. It evaluates an agent on three dimensions.
First, zero-human ability asks how much human work remains. Small means the human still performs at least half the work. Medium means the agent executes while a person delegates and reviews. Large means the agent does most of the work, with less than 20% human involvement. Second, operational domain measures scope: one task, an end-to-end workflow within a function, or a broader cross-functional role. Third, the output/cost ratio asks whether the value created rises roughly with compute cost, outpaces it materially, or dwarfs it.
Those dimensions matter because a human remains the natural commercial anchor only while a human remains the real worker. Once the agent acts with limited supervision, the buyer begins comparing its output with labor, cycle time, revenue, risk reduction, or avoided service volume.
The AMS scores below are Monetizely’s commercial assessment of the public product positions and pricing structures shown in Exhibit 1.
| Product | Zero-human ability | Operational domain | Output/cost ratio | Current commercial read | Recommended primary meter |
|---|---|---|---|---|---|
| Cursor | Medium | Medium | Inflecting | Human developers remain responsible for the shipped code | Seat-led subscription |
| GitHub Copilot | Medium | Medium | Inflecting | Agent work assists software delivery, but developers still review and own the result | Seat-led subscription with usage guardrails |
| Devin | Large | Medium | Inflecting | The agent can execute tasks, but work complexity and reliability still vary materially | Work credits during the transition to output pricing |
| Replit Agent | Large | Medium | Inflecting | The product can build and deploy, yet task size drives cost and quality variance | Subscription plus effort meter |
| 11x Alice | Large | Medium | Inflecting | Prospect creation is observable; downstream revenue attribution is not immediate | Per-prospect output package |
| Intercom Fin | Large | Medium | Inflecting | Resolution can be observed and linked to a support workflow | Verified resolution |
| Salesforce Agentforce | Large | Large | Inflecting | A broad platform supports many workflows, some measurable and some not | Outcome meter by workflow, not one global credit model |
The scoring produces a non-obvious conclusion. A broad agent platform does not earn the right to charge for outcomes simply because it is autonomous. It earns that right only when each workflow has a clear finish line that the customer accepts.
Intercom Fin meets that test more often than most. Its published rules define a $0.99 outcome as a resolution, procedure handoff, disqualification, or qualification, and it charges at most once per conversation. That is not merely a clever price point. It is a measurable agreement about what counts as delivered work.
Salesforce’s menu makes the opposite point. Agentforce offers actions, conversations, user licenses, and a fixed-access option because its platform spans customer service, employee work, and broader enterprise processes. A single universal meter would hide important differences in autonomy, value, and proof.
The practical question is not whether a company should use fixed or variable pricing. The question is which unit should lead the buyer’s mental model.
For the category of autonomous decision-making agents, Monetizely recommends a two-part architecture:
The platform commitment is not a compromise designed to avoid choosing. It pays for a real set of fixed costs and customer needs. The outcome fee remains the primary meter because it is the part of the contract that expands as autonomous value expands.
The table points to a hard commercial rule: a mature autonomous agent should not grow revenue only when more people log in. If the agent can resolve 100,000 customer issues while the customer’s human support team stays flat, the value has expanded even if no additional employee receives a seat.
A credit model has one obvious advantage: it protects gross margin when inference costs are uncertain. That protection is valuable early in a category. Yet a price anchored too closely to compute has a predictable weakness. As model efficiency improves, the vendor’s own meter becomes an argument for lower prices.
OpenAI’s April 14, 2025 release of GPT-4.1 made the direction visible. The company stated that GPT-4.1 was 26% less expensive than GPT-4o for median queries, while GPT-4.1 mini reduced cost by 83% on its stated comparison. More recently, OpenAI has continued to position new models around cost-performance trade-offs, rather than one stable cost curve.
Compute will remain a vital internal control. It should influence model routing, rate limits, premium tiers, and when an agent asks for human help. It should not become the main customer promise once the product has proven it can complete valuable work.
The modeled example below shows the commercial effect. It uses the same 100,000 verified outcomes under two pricing logics.
| Annual volume | Cost-indexed meter at launch | Cost-indexed meter after 50% inference-cost decline | Outcome-led meter | Revenue preserved after lower inference cost |
|---|---|---|---|---|
| 100,000 verified outcomes | $0.60 per unit = $60,000 | $0.30 per unit = $30,000 | $0.99 per outcome = $99,000 | $99,000 |
The implication is straightforward: when the customer still receives the same resolved outcome, a price linked to the outcome holds value more effectively than a price linked to the computation used to generate it.
A company should not interpret that as permission to charge for vague ROI. Pricing a sales agent on closed revenue is usually too distant from the agent’s actions, too dependent on human follow-up, and too easy to dispute. 11x’s current choice to price Alice by new prospects rather than by email sends is more defensible because the output is observable and the unit does not depend on how many touches a sequence requires.
The most common failure in outcome pricing is not the rate. It is the definition.
A support agent may answer a customer’s question, but the customer may return tomorrow. A sales agent may book a meeting that never occurs. A finance agent may process an invoice that later needs correction. If the vendor cannot define the outcome before invoicing begins, finance teams will treat the meter as discretionary and sales teams will discount around it.
The contract should make five elements visible to the buyer before the first bill arrives.
| Contract element | Minimum definition | Example for an autonomous support agent |
|---|---|---|
| Eligible event | What work enters the meter | A new inbound support conversation |
| Agent action | What the agent must actually do | Authenticate, retrieve order data, and provide a policy-compliant answer |
| Completion condition | What makes the event billable | Customer accepts the answer and does not reopen the issue within the agreed window |
| Exclusions | What does not count | Human escalation, failed tool call, duplicate ticket, abuse, or refund-related reversal |
| Audit record | Which data settles a dispute | Conversation ID, system actions, timestamps, outcome code, and reopen status |
Intercom’s published policy offers a practical benchmark: unsuccessful attempts are not billed, and an explicit request for a human prevents the interaction from counting as an outcome. That approach places performance risk where it belongs - with the vendor claiming autonomy.
Operational readiness is equally important. Product telemetry must create the billable event. Entitlements must control access. Finance must reconcile usage records to invoices. Customer success must be able to explain why one interaction counted and another did not. Monetizely’s fifth step exists because pricing that cannot be measured, billed, and defended is not a pricing model. It is a slide in a launch deck.
Many AI companies are currently adding an “AI fee” to an existing SaaS plan. That can fund early experimentation, but it rarely creates a durable category position. Buyers soon ask whether the fee pays for access, usage, labor replacement, or a bundle of features they cannot separate.
A better path begins with the role the agent is meant to occupy in three years. If the company wants to sell a better tool for employees, it should design seats, permissions, and premium usage tiers. If it wants to sell a digital worker, it should build the data and contractual machinery to sell completed work.
The shift need not happen in one release. Devin’s current structure shows why. Its self-serve plans combine recurring access with quotas and on-demand credits, which gives customers control while agent tasks vary in duration and complexity. That structure makes sense while buyers are still evaluating what autonomous coding can reliably deliver. The destination, however, should be a better measure of accepted code, completed issue remediation, or another verifiable engineering output - not permanent dependence on compute units.
Choose the one workflow where the agent can credibly own the final result. Do not begin with the broadest product claim. Begin with a workflow that has a clear start, finish, and system of record.
Set a three-year meter destination before launching the next AI feature. Teams should know whether the product is becoming a seat-led copilot, a work-unit agent, or an outcome-led digital worker.
Make verified outcomes a core company metric alongside ARR and gross margin. Track outcome volume, success rate, reopen rate, human-escalation rate, and customer dispute rate from the first deployment.
Use platform commitments to fund the fixed work that outcomes do not cover. Integrations, governance, deployment, analytics, and support deserve recurring revenue. They should not be hidden inside an inflated outcome rate.
Treat credits as a control system, not the customer value proposition. Use them to manage premium models, exceptional workloads, and uncertain tasks while the product earns the right to charge for business results.

Join companies like Zoom, DocuSign, and Twilio using our systematic pricing approach to increase revenue by 12-40% year-over-year.