Gunjo · Business Intelligence for the AI Era
← Sticker Wall MODEL · DETAIL

Together AI Open-Source Model Inference Cloud Platform

1) Token-based billing: Server-side inference is billed separately for input and output tokens, with pricing per million

MODEL

Key Fields

FIELD STAMPS
IndustryAI / LLM
RegionUS
ScaleMid-size
ChannelOnline

📌 Background

Open-source large models continued to surge in 2026, with a sharp spike in developer demand for low-cost, high-performance inference services. Together AI aggregates over 200 open-source models and provides an OpenAI-compatible API, enabling developers to invoke models without building their own GPU clusters. Its ARR exceeded $300 million in 2025, and a pure consumption-based model drives continued revenue growth.

👤 Target Customers

High-scale inference developers, training/fine-tuning developers, production-grade SaaS, and AI startups

💰 Revenue Streams

1) Token-based billing: Server-side inference is billed separately for input and output tokens, with pricing per million tokens ranging from $0.06 to $4.50; 2) Dedicated endpoints: Dedicated inference endpoints are billed by GPU hours (hours/days/months) rather than token volume; 3) Cluster rental: GPU cluster rentals for training and custom inference are available on-demand or via long-term leases of up to 6 months, billed by lease duration; 4) Usage scaling: Overage fees are charged on a tiered basis when usage exceeds limits, with dedicated capacity quoting for expansion (though actual revenue from this item has not been disclosed).

🧮 Cost Structure

GPU procurement and cluster operations make up the bulk of costs, including power, data center facilities, and maintenance for high-performance GPUs such as H100 and B200. Server-side inference also incurs multi-tenant scheduling, model loading, and idle GPU costs.

🛡️ Moat

Aggregates over 200 open-source models and provides an OpenAI-compatible API, lowering migration costs for developers. The pure consumption model has no subscription barriers, and combined with dedicated GPUs and a 50% discount on Batch API, creates a cost advantage for high-token-volume developers.

🔑 Keys to Success

  • Aggregates over 200 open-source models and provides a unified API
  • Pure consumption model lowers experimentation costs for developers
  • Dedicated GPUs and Batch API cater to high-scale customers

⚠️ Risks

  • Cooling of the open-source model ecosystem leading to a decline in inference demand
  • Price wars among cloud giants squeezing gross margins
  • GPU shortages or price fluctuations impacting costs

🏢 Cases

  • Llama 3.3 70B server-side inference pricing is $1.04 per million tokens
  • Dedicated H100 GPU single-card pricing is approximately $6.49 per hour
  • Batch API offers a 50% cost reduction to attract production-grade SaaS customers

📊 SWOT Analysis

Strengths

  • Broad coverage of open-source models; developers do not need to build their own clusters
  • Low barrier to entry with a pure consumption model; ARR has already exceeded $300 million
  • Dedicated GPUs and Batch API meet high-concurrency and low-latency demands

Weaknesses

  • Heavy reliance on the open-source model ecosystem, lacking differentiation from proprietary models
  • High-token-volume developers may migrate to other cloud providers

Opportunities

  • Continued boom in the open-source community drives growth in inference demand
  • Multimodal models and training/fine-tuning customers bring in GPU-hour revenue

Threats

  • Tech giants such as AWS and Google Cloud entering the open-source inference cloud space
  • Rapid iteration of open-source models requires frequent platform updates and adaptations