Together AI Open-Source Model Inference Cloud Platform
1) Token-based billing: Server-side inference is billed separately for input and output tokens, with pricing per million
Key Fields
FIELD STAMPS📌 Background
Open-source large models continued to surge in 2026, with a sharp spike in developer demand for low-cost, high-performance inference services. Together AI aggregates over 200 open-source models and provides an OpenAI-compatible API, enabling developers to invoke models without building their own GPU clusters. Its ARR exceeded $300 million in 2025, and a pure consumption-based model drives continued revenue growth.
👤 Target Customers
High-scale inference developers, training/fine-tuning developers, production-grade SaaS, and AI startups
💰 Revenue Streams
1) Token-based billing: Server-side inference is billed separately for input and output tokens, with pricing per million tokens ranging from $0.06 to $4.50; 2) Dedicated endpoints: Dedicated inference endpoints are billed by GPU hours (hours/days/months) rather than token volume; 3) Cluster rental: GPU cluster rentals for training and custom inference are available on-demand or via long-term leases of up to 6 months, billed by lease duration; 4) Usage scaling: Overage fees are charged on a tiered basis when usage exceeds limits, with dedicated capacity quoting for expansion (though actual revenue from this item has not been disclosed).
🧮 Cost Structure
GPU procurement and cluster operations make up the bulk of costs, including power, data center facilities, and maintenance for high-performance GPUs such as H100 and B200. Server-side inference also incurs multi-tenant scheduling, model loading, and idle GPU costs.
🛡️ Moat
Aggregates over 200 open-source models and provides an OpenAI-compatible API, lowering migration costs for developers. The pure consumption model has no subscription barriers, and combined with dedicated GPUs and a 50% discount on Batch API, creates a cost advantage for high-token-volume developers.
🔑 Keys to Success
- Aggregates over 200 open-source models and provides a unified API
- Pure consumption model lowers experimentation costs for developers
- Dedicated GPUs and Batch API cater to high-scale customers
⚠️ Risks
- Cooling of the open-source model ecosystem leading to a decline in inference demand
- Price wars among cloud giants squeezing gross margins
- GPU shortages or price fluctuations impacting costs
🏢 Cases
- Llama 3.3 70B server-side inference pricing is $1.04 per million tokens
- Dedicated H100 GPU single-card pricing is approximately $6.49 per hour
- Batch API offers a 50% cost reduction to attract production-grade SaaS customers
📊 SWOT Analysis
Strengths
- Broad coverage of open-source models; developers do not need to build their own clusters
- Low barrier to entry with a pure consumption model; ARR has already exceeded $300 million
- Dedicated GPUs and Batch API meet high-concurrency and low-latency demands
Weaknesses
- Heavy reliance on the open-source model ecosystem, lacking differentiation from proprietary models
- High-token-volume developers may migrate to other cloud providers
Opportunities
- Continued boom in the open-source community drives growth in inference demand
- Multimodal models and training/fine-tuning customers bring in GPU-hour revenue
Threats
- Tech giants such as AWS and Google Cloud entering the open-source inference cloud space
- Rapid iteration of open-source models requires frequent platform updates and adaptations