Foundation Model Usage-Based Subscription (Model Layer Monetization)
1. API Token Usage-Based Billing: Charging developers or application providers based on input/output token counts or que
Key Fields
FIELD STAMPS📌 Background
Capital is accelerating into the foundation model layer: by strict metrics, AI-native financing grew by approximately 146% year-over-year in 2025, reaching a scale of about $136 billion (based on institutional statistics). Meanwhile, enterprise adoption of large models has spread from experimentation to production. The MaaS (Model-as-a-Service) model is maturing, and inference usage has become a measurable consumption-based revenue stream. Anthropic's ARR rose from $1 billion at the beginning of 2025 to $14 billion in February 2026 (as disclosed by the company), validating the scalable monetization capability of the model layer itself.
👤 Target Customers
Enterprises and SMB developers who need to rapidly integrate AI capabilities (text, image, video, code generation, or Agent invocation) pay via API usage or enterprise subscriptions. Payers include both developers calling models directly and organizations purchasing enterprise-grade services, driven by the desire to reduce operational costs or create new revenue streams through these models.
💰 Revenue Streams
1. API Token Usage-Based Billing: Charging developers or application providers based on input/output token counts or query frequency. Revenue fluctuates in real-time with inference volume, forming 'consumption-based revenue,' typically measured by Annual Run Rate (ARR). 2. Enterprise Subscription Contracts: SaaS-style packages for large clients with fixed monthly or annual fees, including dedicated capacity, SLAs, and security compliance to ensure stable cash flow. 3. Cross-Modal Value-Added Services: Charging higher unit prices for video generation, agent orchestration, fine-tuning, or custom models, driving higher average revenue per user (ARPU) and revenue diversification. 4. Ecosystem Revenue Sharing: Leveraging model capabilities to drive partner development or cloud hosting, thereby gaining supplementary revenue from cloud consumption or application revenue sharing.
🧮 Cost Structure
The primary expenses are large-scale compute and inference costs: the cost of self-developed or leased GPU/TPU inference remains high, and these expenses tend to scale proportionally with token volume. Secondary costs include compensation for top-tier R&D talent, as well as ongoing expenditures for acquiring training data and building data centers and infrastructure. Marketing and customer acquisition costs are also significant, covering developer relations and enterprise sales operations.
🛡️ Moat
1. Technological Leadership and Model Performance Gap: Achieving or exceeding benchmarks like GPT requires massive technical and capital barriers, making it difficult to replicate quickly. 2. Scale Effects and Cost Advantages: As inference scale grows, unit inference costs decrease, providing room for pricing flexibility and market defense, narrowing the window for latecomers to catch up. 3. Lock-in Effects and Ecosystem Stickiness: Once enterprises embed APIs into production workflows, Agent processes, and internal tools, migrating to a new provider entails high switching risks and costs. 4. Brand and Trust: Enterprises prefer models with proven risk control and long-standing compliance records, creating upward pressure for new entrants.
🔑 Keys to Success
- Compound growth of enterprise-grade and API businesses
- Scaling consumption-based revenue (token usage)
- Driving volume through cross-modal capabilities (video/agents)
⚠️ Risks
- High inference costs and price wars
- Market share diversion by open-source and domestic models
🏢 Cases
- OpenAI, Anthropic (~$30B ARR)
📊 SWOT Analysis
Strengths
- Superior model performance and multi-modal capabilities, resulting in spillover productivity for end-users
- Consumption-based pricing is naturally tied to usage; higher volume leads to better unit cost efficiency
- Strong market share lock-in effect with high migration costs
Weaknesses
- Inference costs represent a heavy portion of revenue, posing a risk of losses amid price wars
- Revenue is highly correlated with the client's own AI usage, leading to significant volatility
Opportunities
- Surge in Agent use cases drives exponential growth in API calls, creating space for token volume expansion
- Enterprise AI adoption is shifting from experimentation to core business, opening a window for stable subscription revenue
Threats
- Substitution by open-source models and domestic models offering better pricing and flexibility
- Systemic price reductions or subscription discount wars driven by potential regulation and competition