Gradient Serverless Inference Per-Second Billing Service
1) Billing based on the actual GPU seconds consumed per request; 2) Monthly subscription tiers (such as Pro and Growth p
Key Fields
FIELD STAMPS📌 Background
In 2026, AI applications entered an inference-dominated phase. Many small and medium-sized developers renting whole-card GPUs for model deployment found that actual request processing time was less than 10%, leading to a massive waste of budget on idle compute power. Following the integration of Paperspace, DigitalOcean introduced serverless inference under the Gradient product line, featuring per-second billing and zero cost during idle time, emerging as a cost-reduction option for small and medium AI teams.
👤 Target Customers
AI application developers and startup teams deploying open-source models with high traffic fluctuations
💰 Revenue Streams
1) Billing based on the actual GPU seconds consumed per request; 2) Monthly subscription tiers (such as Pro and Growth plans) to unlock higher quotas; 3) Driving incremental revenue for GPU Droplets and bare-metal GPU hourly rentals.
🧮 Cost Structure
GPU procurement (such as H100) and data center operation costs; R&D investments in platform scheduling and metering systems; subsidies for existing free allowances.
🛡️ Moat
DigitalOcean's existing customer base of small and medium developers, out-of-the-box user experience, integrated toolchain from training to inference, and low migration costs.
🔑 Keys to Success
- Accurate metering and transparent billing to build a reputation for cost control
- Cross-selling with DigitalOcean's overall cloud product portfolio
⚠️ Risks
- Gross margins under the per-second model compressed during GPU supply tightness
- Price adjustments triggering developer community churn
🏢 Cases
- User practices show serverless inference significantly reduces monthly GPU expenditures for low-frequency dialogue models
- Post-integration offering of two rental categories: H100 bare-metal reservations and on-demand GPU Droplets
📊 SWOT Analysis
Strengths
- Per-second billing accurately matches sparse traffic scenarios with transparent costs
- Integration with Notebook, training, and deployment toolchains
Weaknesses
- Legacy Paperspace users face product migration and pricing adjustments following brand integration
- Prices may not be the lowest compared to specialized GPU rental providers
Opportunities
- Inference demand exceeds training, becoming the primary consumer of compute power
- The pain point of idle costs is evident, making per-second billing effective for customer acquisition
Threats
- Competition from low-cost GPU rental platforms like RunPod and Vast.ai
- Pressure from similar serverless inference products by hyperscale cloud providers