Monetization
Monetization & Pricing in AI
A pricing model has four parts (Scale, What, Amount, When), must pass three tests (Customer View, Growth Loops, Cost of Revenue), and converts via one equation: Perceived Value > Perceived Price + Friction.
Core thesis
A pricing model has four parts: Scale (per seat, per usage, per outcome), What (the unit you charge for), Amount (the price), and When (billing frequency). It must pass three tests: Customer View (does the customer perceive more value than price plus friction?), Growth Loops (does the pricing model accelerate or decelerate acquisition and retention?), and Cost of Revenue (are margins sustainable as usage scales?). AI breaks traditional SaaS pricing in three specific ways: cost-to-serve variance is extreme (one power user can cost 50x a casual user), frontier model prices stay constant because customers always switch to the best available model, and outcome attribution enables 25 to 50 percent value capture that seat-based pricing cannot access.
Why AI breaks traditional SaaS pricing
Traditional SaaS pricing works on a simple assumption: the cost to serve each additional user approaches zero. A seat in Salesforce costs roughly the same to deliver whether the user logs in daily or monthly. AI products break this assumption. Every inference call has a real compute cost. A power user who runs 1,000 queries per day costs substantially more than a casual user who runs five. The LangChain survey found that 57.3 percent of organizations have AI agents in production, and token consumption per task doubles approximately every six months as models get more capable and users ask more of them. Flat subscription pricing for AI products is mathematically broken because the cost curve slopes upward while the revenue curve stays flat. The companies that survive will have pricing models that track costs.
The Frontier Model Trap explained
Flat AI subscriptions are mathematically broken
Frontier models will always cost roughly $20 per million input tokens and $60 per million output tokens at current market rates, because customers always switch to the best available model. Token-per-task doubles every six to nine months as models improve and users ask more.
The Frontier Model Trap works like this. You launch an AI product with a $30 per month subscription. Your users love it. They use it more each month. Their token consumption grows 15 percent month over month. Your inference costs grow with it. By month 12, your power users are consuming 5x the tokens they consumed at signup. Your revenue is still $30 per month. Your gross margin on those users has gone from 70 percent to negative 20 percent. Meanwhile, a new frontier model launches. Your users expect you to upgrade. The new model costs the same per token as the old one (frontier pricing is remarkably stable at roughly $20 per million input and $60 per million output for top-tier models). But it handles more complex tasks, so users give it harder problems. Token consumption jumps again. The trap tightens. The only way out is a pricing model where revenue scales with usage.
Real pricing data from AI companies
The AI pricing landscape has bifurcated. Consumer AI products default to flat subscriptions: ChatGPT Plus at $20 per month, Claude Pro at $20 per month, Perplexity Pro at $20 per month. These companies are in a land grab phase. Their margins on power users are negative or near zero. They are subsidizing usage to build data moats and brand loyalty. Enterprise AI products are moving toward usage-based pricing. OpenAI's API charges per token. Anthropic's API charges per token. Cohere charges per token. The pattern is clear: where switching costs are low (consumer), flat pricing wins the land grab. Where switching costs are higher (API, enterprise workflow integration), usage-based pricing protects margins. The Writer 2026 survey found that 54 percent of C-suite leaders cite integration difficulty as the primary AI adoption blocker. Integration is hard. Pricing should capture the value of solved integration.
Flat subscriptions vs usage-based pricing
- Flat subscription advantages: predictable revenue, simple sales process, easy customer comparison, faster initial adoption. The customer knows exactly what they will pay. No surprise bills. This matters most in consumer and SMB markets where budget predictability drives purchase decisions.
- Flat subscription risks: margin erosion as usage grows, free-rider problem (casual users subsidized by heavy users leaving for cheaper alternatives), no upside from power users who would pay more, competitive vulnerability when a usage-based competitor undercuts on casual-user pricing.
- Usage-based advantages: revenue scales with costs, power users pay their fair share, casual users get a lower entry price, the product gets cheaper as efficiency improves, value capture aligns with value delivered. This model works best when usage variance is high and power users derive disproportionate value.
- Usage-based risks: revenue unpredictability makes forecasting hard, customers fear surprise bills, usage can decline during economic downturns, sales cycles are longer because procurement needs to understand the unit economics. Gartner data suggests that by 2027, over 60 percent of B2B software contracts will include a usage-based component, up from roughly 35 percent in 2024.
- Hybrid models: a base subscription plus usage overages. This is the most common enterprise model emerging in AI. The base covers fixed costs and a baseline usage tier. The overage captures value from heavy users. OpenAI's ChatGPT Team and Enterprise plans follow this pattern. The base subscription buys seats and baseline GPT-4 access. Additional API usage is metered.
The four-part pricing model in detail
Every pricing model has four variables. Scale is the unit of measurement: per seat, per API call, per token, per outcome, per gigabyte processed, per successful resolution. What is the specific thing you charge for. Do not charge for inputs. Charge for the thing the customer values. No one values a token. They value a resolved support ticket, a generated report, a translated document. Amount is the price per unit. This should reflect the value delivered, not the cost to produce. A report that saves a customer $10,000 in analyst time should cost more than $5. When is the billing frequency: monthly, quarterly, annual, prepaid credits, postpaid invoice. Annual contracts reduce churn but slow adoption. Monthly billing accelerates adoption but increases churn risk. AI products with high switching costs can push for annual. AI products with low switching costs need monthly so customers feel they can leave.
The three tests framework overview
Every pricing model must pass all three
Customer View, Growth Loops, and Cost of Revenue. Fail one and the model breaks. Pass all three and the model scales.
I use three tests with every team that is designing AI pricing. The Customer View test asks: does the customer perceive the value as greater than the price plus the friction of buying? Perceived value is not the same as actual value. A product that saves $100,000 per year but the customer perceives as worth $500 will not sell at $50,000. The Growth Loops test asks: does the pricing model accelerate or decelerate your acquisition and retention loops? A pricing model that creates negative surprises (usage overages that shock the customer) breaks the retention loop. A pricing model that makes the product free for the first user in an organization accelerates the acquisition loop. The Cost of Revenue test asks: at projected usage volumes, do margins hold? If your gross margin drops below 50 percent at scale, the pricing model is broken. AI products should target 60 to 80 percent gross margins at steady state, accounting for inference costs, hosting, and support.
The Customer View test
The Customer View test is the hardest because it requires understanding not what your product does, but what your customer believes it does. I have watched teams price based on cost-plus logic: inference costs $0.30 per query, so we charge $0.50. The customer sees a query that saves them 30 seconds. Thirty seconds is worth maybe $0.10 to them. They do not buy. The Customer View test has three sub-questions. First, what is the before-state cost? What does the customer spend today to achieve the outcome your product delivers? Second, what is the perceived improvement? Not your measured improvement. What improvement does the customer believe they will experience? Third, what is the risk discount? Customers discount value by the probability that the AI gets it wrong. If your AI is 85 percent accurate, the customer mentally discounts the value by 15 percent or more. Price against the perceived value, not the cost.
The Growth Loops test
Pricing is a growth lever. The right pricing model pulls customers through the funnel. The wrong pricing model pushes them away. The Growth Loops test evaluates three dynamics. First, does the pricing model create free or low-cost acquisition? A generous free tier or a usage-based model where the first X units are free turns the product into its own acquisition channel. Users try it, get value, and upgrade when their usage exceeds the free tier. Second, does pricing encourage expansion revenue? Seat-based pricing grows when organizations add seats. Usage-based pricing grows when usage increases. Outcome-based pricing grows when the customer succeeds. The best pricing models align expansion with customer success. Third, does pricing discourage churn? Annual contracts, prepaid commitments, and integrated workflows that break on cancellation all reduce churn. But they only work if the product delivers ongoing value. A contract traps an unhappy customer. It does not retain one.
The Cost of Revenue test
The Cost of Revenue test is where most AI pricing models fail. The math is simple but unforgiving. Calculate your cost to serve one average user at projected scale. Include inference costs, hosting, data storage, support allocation, and payment processing. Divide that by your projected revenue per user. That is your unit margin. Now stress-test it. What happens when inference costs drop 50 percent? Good. Your margin improves. What happens when your most active 10 percent of users increase their usage by 3x? That is the scenario that kills flat-subscription AI products. I recommend modeling three scenarios: baseline (current usage growth), power-user surge (top decile usage doubles), and model-upgrade shock (a new frontier model launches, token-per-task increases 40 percent, and you must upgrade to stay competitive). If your pricing model survives all three scenarios with gross margins above 50 percent, you have a viable structure.
Outcome-based pricing and value capture
Outcome-based pricing is the most defensible model in AI. You do not charge for tokens or seats. You charge for the outcome the AI delivers: per resolved support ticket, per generated sales-qualified lead, per completed financial audit, per translated and localized document. The advantages are significant. First, the customer perceives the price as a cost of goods sold, not an overhead expense. That changes the budget conversation entirely. Second, the pricing scales with the customer's success. When the customer grows, your revenue grows. Third, outcome-based pricing is harder for competitors to undercut because it bundles the AI quality, the workflow integration, and the outcome guarantee into a single price. The disadvantage is measurement complexity. You need to agree with the customer on what counts as a completed outcome. That agreement takes time. But once established, it creates switching costs that seat-based pricing cannot match.
Margin math for AI products
Three margin realities define AI product economics. First, inference is not the only variable cost. Hosting, data pipelines, evaluation runs, monitoring, and support all scale with usage. A product that looks profitable on inference costs alone may lose money when you include the full cost stack. Second, margin on the first dollar of revenue is not the same as margin on the millionth dollar. AI infrastructure has volume discounts. Reserved capacity reduces per-token costs by 30 to 50 percent. Unit economics improve at scale, but only if the pricing model captures that improvement as margin instead of passing it all to the customer. Third, the biggest margin killer in AI is not inference cost. It is evaluation cost. The LangChain survey identified quality as the number one barrier for 32 percent of teams. Ensuring output quality requires evaluation runs that can cost as much as the inference itself. A product that evaluates every output spends 2x on compute. A product that evaluates 10 percent of outputs risks quality degradation that loses customers. The margin sweet spot is somewhere in between, and it changes as models improve.
Vertical integration as a pricing strategy
Vertical integration is a defensive pricing strategy for AI products. The logic: you may lose money on inference, but you capture margin on the layers around it. Hosting the model (instead of calling an API) gives you margin on infrastructure. Managing the database gives you margin on storage. Handling deployment gives you margin on DevOps. Monitoring and observability give you margin on operations. The enterprise customer pays one bill. You make your margin on the bundle, not on any individual layer. Deloitte's 2026 State of AI report found that worker access to AI rose by 50 percent in 2025, but fewer than 20 percent of companies have more than 40 percent of AI projects in production. The gap between access and production is an integration gap. Companies that bridge that gap with a vertically integrated offering capture margin that pure-model companies cannot. The model is a commodity. The integration is the product.
Common pricing mistakes in AI products
- Pricing too low at launch and training customers to expect free value. Raising prices later triggers churn that could have been avoided with a higher initial price and a generous launch discount.
- Pricing based on cost-plus instead of value. The customer does not care what your inference costs. They care what the outcome is worth. Price the outcome, not the input.
- Ignoring usage variance. A small percentage of power users will consume a large percentage of your compute. If your pricing model does not account for this, those power users will destroy your margins.
- Making pricing too complex. Customers need to understand what they will pay within 30 seconds of looking at your pricing page. If they need a calculator, you have already lost a portion of your funnel.
- Anchoring to competitor pricing without understanding competitor economics. OpenAI can subsidize consumer usage because they have enterprise API revenue and Microsoft's balance sheet. Your startup cannot. Price for your economics, not theirs.
The playbook for getting started
Start with a cost-to-serve analysis. For your ten most active users, calculate exactly what they cost you in inference, hosting, support, and evaluation. Compare that to what they pay you. If the ratio is above 70 percent for any user, your pricing model is broken. Next, run the three tests. Customer View: ask five customers what they would pay for your product and why. Do not lead them. Listen. Growth Loops: map how a free user becomes a paying user in your current pricing model. Where is the friction? Cost of Revenue: model the power-user-surge scenario. If your margins die, redesign the pricing before you acquire more power users. Finally, if you are still on flat subscriptions, add a usage-based component within the next quarter. Start with a generous included usage tier so casual users are not affected. The business survives on the power users who pay for what they consume.
Explore other frameworks
The AI Growth Imperative
Strategy
AI Growth Defensibility
Strategy
Acquisition Strategy in AI
Acquisition
Retention & Engagement in AI
Retention
AI Prototyping
Product
AI Product Teams
Product
Enjoyed this framework? Get more research and practical notes in your inbox.
Subscribe to Build Notes →