Gunjo · Business Intelligence for the AI Era
← Sticker Wall MODEL · DETAIL

Gradient Serverless Inference Per-Second Billing Service

1) Billing based on the actual GPU seconds consumed per request; 2) Monthly subscription tiers (such as Pro and Growth p

MODEL

Key Fields

FIELD STAMPS
IndustryCloud Computing
RegionUS
ScaleGiant
ChannelOnline

📌 Background

In 2026, AI applications entered an inference-dominated phase. Many small and medium-sized developers renting whole-card GPUs for model deployment found that actual request processing time was less than 10%, leading to a massive waste of budget on idle compute power. Following the integration of Paperspace, DigitalOcean introduced serverless inference under the Gradient product line, featuring per-second billing and zero cost during idle time, emerging as a cost-reduction option for small and medium AI teams.

👤 Target Customers

AI application developers and startup teams deploying open-source models with high traffic fluctuations

💰 Revenue Streams

1) Billing based on the actual GPU seconds consumed per request; 2) Monthly subscription tiers (such as Pro and Growth plans) to unlock higher quotas; 3) Driving incremental revenue for GPU Droplets and bare-metal GPU hourly rentals.

🧮 Cost Structure

GPU procurement (such as H100) and data center operation costs; R&D investments in platform scheduling and metering systems; subsidies for existing free allowances.

🛡️ Moat

DigitalOcean's existing customer base of small and medium developers, out-of-the-box user experience, integrated toolchain from training to inference, and low migration costs.

🔑 Keys to Success

  • Accurate metering and transparent billing to build a reputation for cost control
  • Cross-selling with DigitalOcean's overall cloud product portfolio

⚠️ Risks

  • Gross margins under the per-second model compressed during GPU supply tightness
  • Price adjustments triggering developer community churn

🏢 Cases

  • User practices show serverless inference significantly reduces monthly GPU expenditures for low-frequency dialogue models
  • Post-integration offering of two rental categories: H100 bare-metal reservations and on-demand GPU Droplets

📊 SWOT Analysis

Strengths

  • Per-second billing accurately matches sparse traffic scenarios with transparent costs
  • Integration with Notebook, training, and deployment toolchains

Weaknesses

  • Legacy Paperspace users face product migration and pricing adjustments following brand integration
  • Prices may not be the lowest compared to specialized GPU rental providers

Opportunities

  • Inference demand exceeds training, becoming the primary consumer of compute power
  • The pain point of idle costs is evident, making per-second billing effective for customer acquisition

Threats

  • Competition from low-cost GPU rental platforms like RunPod and Vast.ai
  • Pressure from similar serverless inference products by hyperscale cloud providers