Key takeaways
- Seat-based SaaS pricing assumed headcount scales with revenue — AI broke that chain.
- A new line item, token spend ("variable cognitive labor"), now sits between headcount and infrastructure — a metered variable cost where salaried fixed cost used to be.
- The denominator is collapsing: new customers hit the same output with far fewer seats, so they never buy them in the first place — a structural pricing collapse, not churn.
- SaaS economics is now a three-body problem: falling headcount cost (but pricier talent), rising token spend (but cheaper per unit), and revenue squeezed from both sides.
- Winners price on outcomes, treat token cost as a core competency (lowest cost per resolution), and build proprietary feedback loops that compound.
The SaaS business model was designed for a world where headcount scaled linearly with revenue. That world ended about eighteen months ago. Most CEOs and CFOs just haven't updated the spreadsheet.
In several AI-native companies and client situations we've studied, substantially smaller teams are performing work that previously required much larger functional groups — engineering organizations shipping comparable output with a fraction of the headcount, plus an AI compute budget that would have seemed absurd three years ago. That is an operating observation from our own engagements, not yet an economy-wide benchmark.
The cost structure did not shrink. It shifted. And that shift is rewriting the fundamental economics of how software companies price, grow, and survive.
01The vanishing denominator
Seat-based pricing was never really about seats, which were a proxy for organizational complexity. The more people a company employed, the more software licenses it needed, and the more a SaaS vendor could charge. Salesforce built a $30 billion annual revenue business on this assumption. So did ServiceNow, Atlassian, and virtually every mid-sized SaaS company founded between 2005 and 2020. When I was the CPO of social media and e-commerce platforms, we charged enterprises a flat platform fee and then a seat-based licensing fee for different teams.
The assumption was elegant and, for two decades, correct. Revenue growth at the customer correlated with headcount growth. Headcount growth correlated with seat expansion. Seat expansion correlated with net revenue retention above 120 percent. The entire SaaS valuation framework — from public market multiples to venture capital term sheets — was constructed on this chain of correlations.
AI broke the chain.
Klarna said in February 2024 that in its first month its AI assistant handled 2.3 million conversations — two-thirds of its customer service chats — doing work equivalent to 700 full-time agents. The implication for seat-based vendors was immediate: every SaaS tool selling per-seat licenses into that support organization lost much of its addressable market inside the account, not because Klarna stopped needing the capability, but because it stopped needing the humans who required the seats.
The sequel matters just as much, and most citations of Klarna omit it. In May 2025 Klarna reversed course and began rehiring human agents, with CEO Sebastian Siemiatkowski conceding that treating cost as the dominant evaluation factor produced "lower quality." By late 2025 the company reported the assistant doing the work of 853 full-time agents — alongside rehired humans, not instead of them. The honest reading is that AI absorbed the high-volume tier and humans came back for the complex and premium tier. Seat compression is real; total labor replacement is not what actually happened.
In our own engagements we have seen directionally similar compression — support, marketing and engineering functions delivering prior output with materially smaller teams. Those are operating observations across our client base, not a measured industry average.
The denominator in the per-seat equation is collapsing. And it is not coming back.
02The token line item
Something strange has appeared in the financial models of well-run startups. Between headcount costs and infrastructure costs — the two line items that historically defined a SaaS company's burn rate — a new category has emerged: token spend.
At OpenAI's current pricing, a company making heavy use of GPT-5.5 pays $5.00 per million input tokens and $30.00 per million output tokens. Anthropic's frontier models sit in a comparable range. What that converts to as a monthly line item depends enormously on architecture, model choice, caching and workload — but in the AI-native companies we work with, it has stopped being a rounding error and started being a budget line that leadership reviews next to headcount.
The useful comparison is not a precise dollar figure, which varies by an order of magnitude between companies. It is the category shift: a meaningful share of cognitive work now arrives as a metered variable cost rather than a salaried fixed cost.
This is not a technology expense in the traditional sense. Token costs represent a fundamentally new category: variable cognitive labor. The company is purchasing thinking by the unit, and the units are getting cheaper while the volume increases.
The trajectory matters enormously. Across the major providers, the capability available at a given price point has risen sharply while per-token prices for equivalent capability have fallen — the reason a frontier model costs $5 per million input tokens while an earlier lightweight model cost a fraction of that for far less capability. The models are getting dramatically more capable while the cost per unit of useful output falls. This creates a compounding advantage for companies that architect their operations around token consumption rather than human headcount.
But it also creates a problem that no one in the SaaS pricing conversation seems to be addressing honestly.
03The pricing model paradox
Consider a hypothetical SaaS company — call it WorkflowCo — that sells project management software at $25 per seat per month. In 2023, a typical mid-market customer with 300 employees had 40 WorkflowCo licenses, generating $12,000 in annual recurring revenue. By 2026, a similarly new customer achieves similar operational output with 100 employees needing only 10 licenses. WorkflowCo's revenue from the new account drops to $3,000 — a 75 percent decline in new-account revenue from a successful, growing market.
The evidence from seed investors is not that existing SaaS customers are immediately cutting seats at scale. The stronger signal is that the next generation of startups is being advised to build around agents, services replacement, and outcome-based economics from the beginning.
Y Combinator has explicitly solicited companies that perform work traditionally bought as services, rather than merely making existing services faster. a16z has argued the same thesis from the investment side — that software is beginning to eat labor, and that per-seat is no longer the natural atomic unit when an agent, not a person, completes the work. The result is subtle but dangerous for seat-based SaaS: the new customer does not churn seats later. They simply never buy them in the first place.
This is not a churn problem. This is a structural collapse of the pricing model.
The reflex response from the SaaS industry has been to pivot toward usage-based or outcome-based pricing. Charge for what gets consumed, not who consumes it. On paper, this sounds reasonable. In practice, it introduces a set of problems that most companies are deeply unprepared for.
Usage-based pricing requires the vendor to absorb variable costs — including token costs — that scale unpredictably with customer behavior. A customer that discovers a particularly effective AI workflow might increase token consumption by 400 percent in a single quarter. The SaaS vendor's cost of goods sold spikes, but the customer expects the price to reflect the marginal cost of computation, not the full cost of the vendor's R&D, go-to-market, and operational overhead.
Variable inference costs can put real pressure on traditional SaaS gross-margin assumptions, particularly when usage rises with the value delivered. The actual effect depends heavily on model choice, architecture, caching, utilization, pricing design and workload — this is an engineering and pricing problem, not a law of nature. But the tension is structural: companies that pass token costs through to customers start to look less like software businesses and more like services businesses, while companies that absorb them to preserve headline margin end up subsidizing their heaviest users.
Neither option produces the economics that venture capital has spent two decades optimizing for.
04The compression effect
What makes this moment genuinely different from previous SaaS transitions — the shift from on-premise to cloud, or from perpetual licenses to subscriptions — is the speed at which organizational compression is occurring.
We have spoken with over forty founders in the past several months who describe the same pattern. A company raises a Series A with twelve to fifteen people. Within six months, that team is shipping products at a velocity that would have required thirty-five to forty people in 2022. The AI tooling — Cursor for engineering, Claude for strategy and analysis, various agents for customer operations — does not just make individuals faster. It eliminates entire roles.
The product manager who spent three days writing specifications now works with an AI that generates comprehensive specs in three hours. The QA team of four becomes a single engineer running AI-powered test generation. The data analyst who built dashboards is replaced by a natural language interface that any operator can query directly. The customer success manager handling fifty accounts can now manage two hundred because AI drafts every email, summarizes every call, and flags every risk signal automatically.
Each of these compressions is individually unremarkable. Collectively, they represent a structural transformation of what a company needs to look like at each stage of growth.
Cursor's parent company, Anysphere, was reported by TechCrunch to have reached roughly $300 million in ARR by mid-April 2025, having tripled from $100 million in January of that year, and to have passed $500 million by that June. The striking number is the slope, not a revenue-per-employee ratio — we have not seen a reliable headcount figure for that period, and deriving one would be guesswork.
Harvey, the legal AI company, describes its product as taking lawyers from a blank page to a reviewable first draft, with the lawyer retaining final judgment over every word that reaches the client or counterparty. That is not the elimination of legal work. It is a shift in the unit of work — from blank-page creation toward human review and judgment — and it is a more defensible reading of what AI is doing to professional services than the replacement story.
The companies being built today are not leaner versions of their predecessors. They are architecturally different organisms.
05The three-body problem of SaaS economics
The old SaaS equation had two variables: headcount cost and infrastructure cost. The margin between revenue and those two cost categories determined whether a company was investable. The new equation has three variables that interact in ways that break conventional financial modeling.
Headcount cost is falling — but talent is pricier
Fewer people, more skilled, paid more each.
Headcount cost is declining as a percentage of total spend, but the humans who remain are more expensive because they need to be significantly more skilled. A team of four engineers who can effectively orchestrate AI-augmented development workflows commands higher individual salaries than a team of twelve engineers performing more narrowly defined tasks. Total headcount cost drops, but cost per head increases.
Token cost is rising — but cheaper per unit
Bigger absolute budget, falling cost per unit of output.
Token cost is rising as a percentage of total spend, but on a per-capability basis it is falling rapidly. A company might spend $80,000 per month on tokens today to achieve what would have cost $300,000 per month in human labor eighteen months ago. Next year, the same capability might cost $40,000 in tokens. The absolute token budget grows because companies keep finding new applications, but the cost per unit of output keeps declining.
Revenue is under structural pressure
Customers pay less while demanding more.
The revenue model is under structural pressure because customers are simultaneously shrinking their teams (reducing seat-based revenue) and demanding more AI-native capabilities (increasing the vendor's cost to serve). The customer is getting more value while paying less, and the vendor is delivering more capability while spending more on compute.
The companies that figure out how to navigate this three-body problem will define the next era of software economics.
The companies that do not will find their unit economics deteriorating quarter by quarter, even as their products get better and their customers get happier.
06What actually works
The emerging answer is not a single pricing model. It is a fundamental rethinking of what a SaaS company sells. The companies we see navigating this transition most effectively share three characteristics. They price on outcomes rather than inputs. They treat token costs as a core competency rather than a pass-through expense. And they build proprietary feedback loops that make their AI capabilities compound over time, creating a moat that raw token access cannot replicate.
We recently worked with a healthcare company to develop a pricing model based on per-patient retention. The clinics pay when a problem gets solved, regardless of whether a human or an AI solved it. This aligns the vendor's revenue with the customer's value creation rather than the customer's organizational size. It also means the vendor is intensely motivated to reduce its own token costs per resolution — a competitive advantage that compounds with every improvement in model efficiency.
This is the structural insight that most SaaS founders are missing. In a world where the primary input is variable AI compute, the competitive advantage belongs to whoever can deliver the most value per token consumed. Not per seat licensed. Not per API call billed. Per unit of customer outcome achieved.
That requires a depth of domain-specific optimization that generic AI tools cannot provide. It requires proprietary data flywheels, fine-tuned models, and workflow architecture that turns raw intelligence into reliable business outcomes. It requires, in other words, exactly the kind of deep vertical expertise that the best SaaS companies were always supposed to build — but that seat-based pricing never forced them to develop.
References
- Klarna. (2024, February 27). Klarna AI assistant handles two-thirds of customer service chats in its first month. Reports 2.3 million conversations handled and work equivalent to 700 full-time agents in the first month. Klarna subsequently began rehiring human agents in May 2025, citing quality.
- OpenAI. API Pricing. GPT-5.5 at $5.00 per million input tokens and $30.00 per million output tokens.
- Andreessen Horowitz. Software is Eating Labor. See also AI Is Driving A Shift Towards Outcome-Based Pricing (December 2024).
- TechCrunch. (2025, June 5). Cursor's Anysphere nabs $9.9B valuation, soars past $500M ARR.
- Harvey. How Lawyers Use AI for Contract Drafting. Describes the lawyer retaining final judgment over every word that reaches the client or counterparty.
Note on evidence: third-party findings are cited above. Statements describing what "we've seen" or "in our engagements" are StatsLateral operating observations from client work, not industry benchmarks. The WorkflowCo scenario is an illustrative model, not a real company.