On the Dwarkesh podcast, Dario Amodei laid out a case for how AI labs eventually make money: individual models are already profitable, and the cash burn is an artifact of each new training run costing far more than the last. I find the model convincing, but it raises a sharper question that the interview only partly answers. When does the company, not just the model, turn profitable?
Current gross margin trajectory bears out Dario's model (negative 94% in 2024 to ~50% in 2025, projected 77% by 2028): individual models are already profitable, and companies bleed cash only because each successive training run costs an order of magnitude more than the last. But a scaling plateau alone doesn't guarantee sustained profitability. It requires two conditions: training cost growth slows, and revenue per unit of intelligence doesn't collapse under competitive pressure. Foundation model companies start making real money when they convert temporary model superiority into enterprise workflow lock-in that outlasts the commoditization cycle.
1. Training cost growth slows. The exponential climb in cost per generation has to decelerate toward incremental, or the burn never stops.
2. Revenue per unit of intelligence holds. Pricing has to resist collapse even as raw model quality commoditizes. This is what workflow lock-in protects.
Intelligence, compute, and capital are necessary but insufficient moats
Google has the best compute infrastructure and deepest enterprise relationships on earth. Meta has distribution to 3 billion users and competitive open-source models. If the moat were model quality plus infrastructure, they should be winning.
Anthropic and OpenAI offer a useful natural experiment. Anthropic went enterprise-first and hit $30B ARR by April 2026, surpassing OpenAI's $25B while spending a quarter of what OpenAI spends on training. OpenAI went consumer-first with 900M weekly users but higher burn. Enterprise-first converts to defensible revenue faster because enterprise contracts expand over time, carry higher retention, and crucially, enterprise customers buy workflow outcomes rather than model access, making them far less sensitive to benchmark leapfrogging. Consumer subscriptions, cancelled in a click when the next model drops, do not have this property.
The revenue divergence is not explained by benchmark superiority. On SWE-bench Verified, the top six frontier models score within 0.8 points of each other, yet Anthropic's 500+ enterprises spending $1M+ annually chose Claude despite this parity. Something beyond model intelligence is producing durable differentiation.
Workflow lock-in is the mechanism that defends margins
Enterprise switching costs are structurally underestimated. For mid-market enterprises and up, procurement cycles run 6 to 9 months, compliance certifications take 12 to 24 months, and signed contracts create binding commitments. Singapore's government recently deployed agentic AI across public sector agencies with provider-specific risk frameworks and governance structures requiring deep institutional integration that cannot be swapped on a quarterly cycle.
Client-specific post-training creates model variants that are economically and organizationally costly to replicate in time. Once embedded in an enterprise workflow, a lab fine-tunes on proprietary data, internal feedback signals, and domain-specific RLHF that the client would never expose to a third party. Could a competitor approximate this via synthetic data or similar domain training? In principle, yes. In practice, replicating 18 months of accumulated fine-tuning on a Fortune 500's internal processes requires access, context, and organizational trust that take years to build. The barrier is not technical impossibility but the time cost of replication, and in a market moving this fast, time is the scarcest resource.
These layers compound. Orchestration frameworks, memory architecture, evaluation stacks, enterprise integrations: each individually replicable, but collectively forming an ecosystem far harder to replicate as a whole.
Even if the full stack were replicated, switching friction persists. As 2025 Nobel laureate Joel Mokyr has argued, technology adoption is shaped not just by suppliers and customers but by non-market actors: regulators, compliance committees, procurement boards. These actors don't benchmark models. They ask: "If this goes wrong, can we defend the decision to use this vendor?" A lab's safety track record, compliance history, and institutional relationships form trust networks that function as a moat independent of model quality.
Claude Code's journey has been instructive. It reached $2.5B in annualized revenue within nine months and now authors roughly 4% of all public GitHub commits, despite GPT-5.3-Codex leading Terminal-Bench 2.0 at 77.3% versus Claude's 65.4%. The product won on deployment depth, not model intelligence alone.
Profit becomes sustainable when both conditions hold
Return to the two conditions from the opening. Both need to hold.
Condition 1: training cost growth slows. I assume it decelerates from exponential (10x per generation) to incremental (~1.2x) as log-linear returns to scale weaken the incentive for ever-larger runs and physical constraints (power, chip supply) impose ceilings. This is an assumption, but consistent with Anthropic's own projection of positive free cash flow by 2028. The illustrative table below traces the unit economics across four generations, from exponential burn to a post-plateau surplus:
| Gen 1 | Gen 2 | Gen 3 | Gen 4 (post-plateau) | |
|---|---|---|---|---|
| Training cost | $1B | $10B | $100B | ~$120B |
| Inference cost | $1B | $10B | $100B | $130B |
| Revenue | $4B | $40B | $300B | $450B |
| Gross profit (revenue minus inference) | $3B | $30B | $200B | $320B |
| Net profit (gross profit minus training) | $2B | $20B | $100B | $200B |
| Next gen's training cost | $10B | $100B | ~$120B | ~$140B |
| Company cash after funding next gen | -$8B | -$80B | ~-$20B | +$60B |
Numbers are illustrative, not predictive. They encode assumptions about pricing, demand, and cost curves, but show the structural dynamic: company-level profitability emerges when next-generation training costs grow incrementally rather than exponentially.
Condition 2: revenue per unit of intelligence doesn't collapse. If labs were selling tokens at market rates, competition would compress margins to near zero. But enterprises embedded in provider-specific workflows buy outcomes, not tokens, and outcomes can be priced on value delivered rather than compute consumed. The lock-in described above prevents revenue from collapsing even as the underlying intelligence becomes commoditized.
Both conditions are reinforced by Jevons paradox: as inference costs fall, consumption rises even faster because entirely new use cases emerge.[1] Agentic workflows, always-on AI assistants, and real-time enterprise data processing all become viable at lower price points, expanding the pie faster than competition erodes margins.
What to watch for
Open-source models will close the raw intelligence gap; distillation makes this nearly inevitable. The real question is whether the ecosystem and trust layers also become portable or standardized. If they do, margins collapse. If algorithmic breakthroughs or new paradigms (agents, multimodality) restart exponential capex cycles, the cost plateau never arrives.
The labs that survive won't be the ones with the highest benchmark scores but the ones that treated model superiority as a distribution wedge, converting a three-month intelligence lead into a multi-year infrastructure position.