
Frameworks, core principles and top case studies for SaaS pricing, learnt and refined over 28+ years of SaaS-monetization experience.
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
Join companies like Zoom, DocuSign, and Twilio using our systematic pricing approach to increase revenue by 12-40% year-over-year.
AI pricing tests are moving from a product-management curiosity to a commercial necessity. A customer-support leader can now buy an agent that charges only when it resolves a case. An engineering leader can buy an autonomous coding product through a mix of seats, included capacity, and usage credits. A law-firm buyer may still expect a price per lawyer, even when AI does much of the underlying work.
The risk is not waiting too long to test. The larger risk is testing the wrong thing. Teams often put a new AI price in market before they have defined the buyer, the job, the proof of value, or the billing record needed to defend the invoice. They then call the result “price sensitivity” when the real problem was an unready commercial model.
Monetizely’s position is clear: every industry should begin AI pricing tests now, but only customer service should move quickly to outcome pricing. Software engineering should lead with a platform-plus-consumption model; legal, healthcare, and regulated financial work should price access and documented work before they price business or clinical decisions. Cost-based meters should protect margin during the transition, not become the long-term basis of value.
A useful AI pricing test begins when three things meet: the agent performs a bounded job, the customer can observe whether that job happened, and finance can produce a bill the customer will recognize. Product readiness alone is not enough. An agent may work well in a demo while still lacking a fair unit of value.
Consider the contrast between two customer-service agents. Intercom’s Fin charges $0.99 for a resolution, procedure handoff, or lead disqualification, while a successful lead qualification costs $9.99. Its published definition is specific: a resolution occurs when no further help is requested after Fin’s final answer. That is a customer-visible event with a clear record.
By comparison, a coding agent may open a pull request, pass tests, and still create work for a senior engineer who must review architecture, security, and product fit. Charging for an “accepted pull request” may sound outcome-based, yet the accepted result reflects both AI work and human judgment. The right early test is therefore not a paid outcome. It is a pricing structure that gives buyers access, contains compute risk, and produces data on what kinds of work the agent can complete.
Monetizely’s 5-Step Pricing Framework puts this sequence in the proper order. As developed in Monetizing Agentic AI, the framework treats price as the result of five linked commercial decisions, rather than the first decision a team makes. Goals and segments establish whom the company is trying to win. Packages turn those customer differences into offers. The pricing metric identifies what the customer will pay for. Price points set the actual rate. Operationalization makes the model work in product, billing, sales, and renewal processes.
The five steps are:
The practical implication is simple: a pricing test should start at Step 3 only after Steps 1 and 2 are settled. Teams that skip those steps may learn that a customer dislikes an invoice, but not whether the customer dislikes the price, the package, or the meter.
The market has not converged on a single AI pricing model. That is a sign of commercial maturity, not confusion. Each model reflects a different mix of autonomy, buyer expectations, cost exposure, and ability to measure value.
The comparison below grounds the discussion in published B2B SaaS offers as of September 3, 2026.
The table reveals an important point: many offers described as “flat” are actually fixed commitments with a capacity boundary. That distinction matters. A true flat fee leaves the vendor exposed as usage rises; a capacity plan preserves budget certainty while limiting that exposure.
The Agentic Monetization Spectrum, or AMS, answers a more useful question than “Should we use outcome pricing?” It asks how far the agent has moved away from the human worker who once anchored the software purchase. The AMS scores an agent on three dimensions: zero-human ability, meaning how little human effort remains in the job; operational domain, meaning whether the agent handles a task, a workflow, or work across functions; and output/cost ratio, meaning whether the customer value rises faster than the cost to produce the work. As autonomy, domain breadth, and the output/cost ratio rise, pricing should move away from seats and toward measurable output or outcomes.
The scoring below does not claim that every customer deployment behaves identically. It identifies the commercial position of the workflow each vendor is selling.
| Product or commercialized workflow | Zero-human ability | Operational domain | Output/cost ratio | Pricing implication |
|---|---|---|---|---|
| Cursor coding assistant and agent | Medium | Medium | Inflecting | Keep the developer seat as the primary anchor; use usage charges to cover heavy AI work |
| Devin autonomous software engineering | Large | Medium | Inflecting | Use a platform-plus-consumption model while task quality remains uneven |
| Intercom Fin routine support resolution | Large | Medium | Inflecting | Charge per defined resolution where the close condition is observable |
| Salesforce Agentforce action-based deployment | Medium | Large | Inflecting | Charge for actions or conversations until a specific business outcome can be attributed cleanly |
| Sierra enterprise customer-service agent | Large | Large | Exponential | Charge per business outcome when the outcome is immediate, measurable, and attributable |
| 11x Alice outbound prospecting | Large | Medium | Inflecting | Use a capacity-based lead or prospect meter; do not charge per booked meeting before attribution is proven |
The AMS explains why two apparently similar agents deserve different commercial tests. Fin and Sierra both automate customer interactions, yet Fin’s published resolution definition supports a discrete, low-friction meter for routine service. Sierra’s broader enterprise model can credibly price a larger business outcome because it is designed around multi-system workflows and outcome measurement.
Coding presents a different case. Cursor remains tied to the productivity of an individual developer, which supports a seat anchor. Devin does more work independently, but its output still requires engineering judgment. Its current mix of recurring access, seats, included quota, and on-demand credits is therefore closer to the AMS position than a premature “per merged feature” promise would be.
Before an industry tests a new meter, leadership should assess four conditions. The exercise should be run on a specific workflow, not on the entire AI product.
Four passes justify a paid outcome test. Three passes justify a platform-plus-consumption test. Two or fewer passes mean the company should collect product data without changing its commercial architecture.
That standard is demanding by design. A disputed invoice can erase the trust created by a successful product pilot. It also aligns with NIST’s AI Risk Management Framework, which calls for organizations to specify intended use, document human oversight, measure performance, and monitor deployed systems rather than treating testing as a one-time event.
Customer service is the clearest industry for immediate outcome-based tests. The work is frequent, the event trail is rich, and many routine contacts have a natural stopping point: an issue is resolved, an order is changed, an appointment is booked, or a return is initiated. Begin with one narrow workflow, such as order-status questions or password resets, and bill only when the system records the stated completion condition.
The primary meter should be a verified resolution. A modest platform fee may fund implementation, security, and reporting, but the strategic meter should remain the resolution. Charging per token, message, or minute would make the buyer absorb the vendor’s operating choices rather than pay for the reduced service workload.
Software engineering should also test immediately, but its primary meter should be platform access plus AI work consumed, not an outcome fee. The best early contract gives a team predictable access through seats or a monthly minimum, includes a defined amount of work, and charges for additional use through shared credits or other consumption. Cursor and Devin both show variants of this design.
A completed ticket should be monitored as a shadow metric. It should not yet be the invoiced unit. Engineering leaders need evidence that the code is accepted, secure, maintainable, and worth operating six months later. Until the system can separate AI-created value from the judgment of the reviewer, an accepted pull request is not a clean economic outcome.
Legal, healthcare, and regulated financial workflows should not wait to test pricing. They should refuse the wrong test.
Legal AI can begin with per-lawyer, per-practice-group, or firm-access pricing, plus a premium usage band for high-volume research, document review, or due-diligence work. The human lawyer still owns the judgment, the client relationship, and the final work product. A seat anchor reflects how law firms budget and how they allocate responsibility, even where the AI creates a large share of the draft work. Harvey’s enterprise focus illustrates the commercial logic of serving that buyer through a controlled, professional-user model.
Healthcare administration should price a completed, documented administrative transaction - for example, a finished intake packet, appointment change, or prior-authorization submission prepared for clinician review. It should not price a clinical decision, a treatment result, or an authorization approval. CMS has explicitly described AI and machine learning in prior authorization alongside clinician review, while encouraging secure innovation, transparency, evaluation, and ongoing monitoring.
Regulated financial operations should follow the same discipline. An agent that prepares a claims file, gathers know-your-customer documentation, or routes an exception can be tested on completed case preparation. Do not pay the vendor for approving a claim, rejecting a customer, or making a credit decision. Paying for those decisions creates an incentive problem precisely where the buyer needs independent judgment.
Revenue teams belong between customer service and regulated work. Outbound AI can be sold today through a capacity-based prospect meter. 11x Alice’s current published plan prices new prospects rather than the number of touches sent, which is a better fit than billing by email volume. A booked meeting should remain a monitored result until the vendor and customer agree on qualification criteria, attendance rules, routing, and sales follow-up. Otherwise, the agent is being paid for an event whose value depends heavily on people outside the product.
The following view translates AMS position into a commercial starting point.
The table means that timing is not a choice between “move now” and “wait.” The practical choice is whether an industry is ready for an outcome invoice, a work invoice, or an access invoice.
Compute matters. It can determine whether an AI product has a viable gross margin during its first years in market. Yet compute should rarely define the customer’s long-term price.
The reason is straightforward. Imagine an agent priced at $2 per unit of model work. If model routing, caching, or a better underlying model cuts the cost of completing the same job in half, a competitor can offer the same customer result at $1 and preserve the same margin. The original $2 meter has no durable link to what the buyer values.
That pressure is strongest when the AI produces work that the customer can already measure in business terms. A customer does not value a support agent because it used 80,000 tokens. The customer values fewer tickets, shorter waits, and retained revenue. Nor does an engineering leader value the number of inference calls. The leader values shipped, reliable software.
Monetizely’s view is that cost-based pricing has a proper but limited role:
Salesforce’s Flex Credits provide a useful bridge. The company can charge $0.10 for a standard Agentforce action while customers experiment across workflows and channels. As a deployment proves that it resolves a specific problem, such as a defined customer-service issue, the vendor has the option to move toward a more valuable meter. The bridge should not become the destination.
A pricing test should run long enough to produce invoice-level evidence, not merely survey-level reactions. Ninety days is usually sufficient for a recurring workflow because it captures onboarding, first use, repeated behavior, and at least one billing conversation.
The test should place comparable customers into no more than two commercial offers. One offer should retain the existing meter. The second should introduce the new meter with a spend ceiling and a clear credit or dispute rule. Sales compensation must be neutral across the two offers, or sellers will steer deals toward the easier contract.
Track five measures together:
A test that improves usage but creates billing disputes has not succeeded. Nor has a test that protects gross margin by making the agent too hard to buy. The winning structure is the one that customers can understand, finance can operate, and the business can carry from pilot to renewal.
Choose one primary meter for each AI category in the portfolio. Customer-service agents should have a resolution path; coding agents should have a platform-plus-consumption path; regulated workflows should have an access or documented-work path.
Make AMS reassessment a product-release decision, not an annual pricing exercise. Re-score an agent when autonomy expands, human review falls, or its domain grows from one task to an end-to-end workflow.
Set a migration rule before signing early customers. State what product evidence will justify moving from credits to output pricing, or from seats to a premium usage layer, at renewal.
Fund telemetry as commercial infrastructure. Event records, audit trails, entitlement systems, and readable invoices are not back-office details. They determine whether a new meter can survive procurement and renewal.
Decide where the company will not monetize a decision. In clinical, legal, and regulated financial settings, preserving independent human judgment can be more valuable than capturing a short-term outcome fee.

Join companies like Zoom, DocuSign, and Twilio using our systematic pricing approach to increase revenue by 12-40% year-over-year.