
Frameworks, core principles and top case studies for SaaS pricing, learnt and refined over 28+ years of SaaS-monetization experience.
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.

You’ve probably read that AI companies operate at 50-60% gross margins. That’s true for many of them today. ICONIQ Capital’s January 2026 State of AI report found that inference costs average 23% of total revenue at scaling-stage AI B2B companies. Bessemer Venture Partners documents AI gross margins at 50-60% against 70-90% for mature SaaS.
Those numbers describe companies paying retail API prices for frontier model inference on every request. That’s how most agentic companies start (it’s the fastest path to market), but it is not where a well-run agentic company stays. The 72% gross margin in the agentic P&L above reflects what happens when a company gets serious about its inference stack.
The assumption that you need a $15/million-token proprietary frontier model for every inference call is no longer supported by the data.
According to Epoch AI, open-weight models now trail state-of-the-art proprietary models by only about three months on average. DeepSeek-V3.2-Speciale has surpassed GPT-5 and reached Gemini-3.0-Pro-level reasoning on benchmarks. Microsoft’s Phi-4, at 14 billion parameters, outperforms models 10x its size through curated training. A 3B parameter model fine-tuned on medical literature can outperform GPT-5 on clinical documentation tasks. Open-source models match or beat closed alternatives on coding (SWE-Bench), multilingual tasks, and domain-specific work after fine-tuning.
The proprietary frontier still holds an edge on the absolute ceiling of complex, open-ended reasoning. That ceiling is relevant to perhaps 5-10% of the inference calls in a typical agentic workflow. The other 90-95% (classification, extraction, tool-call formatting, summarization, retrieval, validation) can be handled at parity or better by models that cost a fraction of the price.
Serving a 7B parameter SLM is 10-30x cheaper than running a 70-175B parameter model, cutting cloud and energy expenses by up to 75%. The cost gap between inference on a 7B model versus a 700B model is not incremental. It is orders of magnitude.
For high-volume agentic workloads, that gap decides profitability.
The companies reaching 70%+ gross margins are doing two things well: choosing the right models and intelligently routing between them.
Gartner has published guidance codifying this approach: routine, high-frequency tasks must be routed to efficient small and domain-specific language models, while expensive frontier-level inference must be heavily gated and reserved exclusively for high-margin, complex reasoning tasks.
The architecture is converging on a standard pattern. SLMs handle 90-95% of queries. Cloud-hosted frontier LLMs handle the 5-10% requiring complex multi-step reasoning.
Claude Code itself implements this pattern internally. It uses Haiku (Anthropic’s smallest, cheapest model) for sub-agent exploration tasks like codebase search and quick analysis, reserving Opus (the most expensive, most capable model) for planning, orchestration, and complex reasoning. Cheaper model for the bulk of the work, expensive model for the hard parts.
The math: inference represents 23% of revenue at retail API prices. Hybrid routing reduces the blended inference cost by 70-80%, so that line item drops to roughly 5-8% of revenue. Add the remaining COGS (hosting, infrastructure, support) at roughly 15-20%, and you land at the 72% gross margin in the agentic P&L. (see Figure 2) Figure 2. How agentic AI reaches 72% gross margin - hybrid routing collapses inference from roughly 23% of revenue at retail API prices to 5-8%.
Companies are already executing this transition and reporting results.
Sully.ai builds AI employees that handle routine tasks for physicians. Their proprietary closed-source models created three bottlenecks: unpredictable latency, inference costs scaling faster than revenue, and insufficient control over model updates. They switched to open-source models deployed on NVIDIA Blackwell GPUs. According to the company, inference costs dropped by 90%. Response times improved by 65%. They returned over 30 million minutes to physicians. That last number deserves a pause. Thirty million minutes of a doctor’s time, returned, by swapping one model architecture for another.
Latitude (AI Dungeon) faces a business model where every player action triggers an inference request. They moved to open-source MoE models on Blackwell. Cost per million tokens dropped from 20 cents to 5 cents, a 4x cost reduction while maintaining quality.
Replit’s gross margins have fluctuated dramatically during 2025, ranging from 36% to negative 14% as the company iterated on pricing models and navigated the high cost of LLM inference for its coding agents. The margin trajectory for agentic companies is not fixed. It is a function of how seriously you manage the model layer.
The genuine, structural caveats remain. Switching from closed to open models has real costs. MIT Sloan notes ecosystems built around closed models are not trivially ported. Self-hosting requires GPU procurement (lead times exceeding 30 weeks), monitoring, and operational investment. Total task costs can rise even as per-token costs fall, because agentic workflows trigger 5-20 model calls per user action. In regulated industries, routing a high-stakes task to the wrong model tier can have severe consequences. Vendor risk shifts rather than disappears.
The directional conclusion is clear: 40-60% gross margins describe where the median agentic company is today. 70%+ gross margins describe where the well-run agentic company lands once it optimizes its inference stack. The COGS line in the agentic P&L is a moving target, and it’s moving in the right direction.
The operating expense lines reveal the second structural advantage that makes the agentic P&L work. It’s the one that makes the EBITDA story sing even if you’re skeptical about the gross margin trajectory.
A traditional SaaS company at $10M ARR typically employs 60-100 people: 30-50 engineers building and maintaining the product, 15-25 in sales and marketing, 8-12 in customer success and support, and 5-8 in G&A. Average fully loaded cost per employee runs ~$150K. At the midpoint of those ranges, total people cost runs roughly $12M against $10M of revenue, which is why so many SaaS companies at this stage burn cash. Revenue per employee across the band: $100-170K.
The SaaS company in Figure 1 is the disciplined version of this org: roughly 55 people and about $8.2M in payroll, running leaner than the benchmark band on engineering. Even that company only manages a small loss.
An agentic AI company at the same revenue typically employs 25-40 people: 10-20 in engineering and ML combined (smaller because AI writes significant portions of the code; Anthropic’s head of Claude Code stated that 100% of his contributions to Claude Code were written by Claude Code, and because model tuning, evaluation, and harness engineering sit inside the same team), 6-10 in sales (often founder-led longer, with outcomes-based selling that shortens cycles), 4-6 in customer success (sitting in COGS, since the product handles much of its own support), and 3-5 in G&A.
Average fully loaded cost per employee runs ~$190K (higher per head due to technical talent concentration; harness engineers, eval specialists, and ML ops roles command premium compensation). The agentic company in Figure 1 runs 27 people and roughly $5.2M in payroll. Revenue per employee: $250-400K.
The payroll gap is the single biggest driver of the EBITDA swing between the two P&Ls: roughly $3M between the two lean companies in Figure 1, and $4M or more at the midpoints of the benchmark bands. Even if you’re conservative on the gross margin line and assume the agentic company only reaches 60%, the operating margin is still positive and competitive with SaaS. The denominator of the payroll equation is half as large.
In SaaS, you need a 25-person sales and CS team because the product is a tool that customers need help using. In agentic AI, the product does the work. The customer doesn’t need as much hand-holding, and you don’t need as many people to sell, onboard, and support them. Salesforce’s CEO Marc Benioff noted that the company isn’t hiring more engineers, citing a 30%+ increase in engineering productivity from AI internally. Agent-native startups are born into this reality. They never build the 80-person org in the first place.
Join companies like Zoom, DocuSign, and Twilio using our systematic pricing approach to increase revenue by 12-40% year-over-year.