Gunjo · Business Intelligence for the AI Era
← Sticker Wall MODEL · DETAIL

Stepfun Low-Inference-Cost Document and Code API

1) Primary source: API call revenue billed by token, providing model interfaces and usage-based billing to developers an

MODEL

Key Fields

FIELD STAMPS
IndustryAI / LLM
RegionChina
ScaleMid-size
ChannelOnline

📌 Background

In 2026, the large model competition shifted from a parameter race to inference efficiency and cost control. With global computing resources remaining tight and cloud service costs rising, token cost has become a key factor in the commercialization of large models. Leveraging its Step series MoE architecture (flagship Step-2 with 2 trillion parameters and Step 3.5-Flash leading in inference efficiency) as well as the open-source model Step 3.7 Flash (196.5 billion total parameters, approx. 11 billion activated per run, 256K context), Stepfun reduces inference costs while maintaining high throughput. Its models rank among the top in call volume across multiple open platforms.

👤 Target Customers

Enterprise customers, including software development teams, enterprise IT departments, and developer platforms requiring intelligent document processing (contracts, reports, knowledge base Q&A) and code generation/completion/review, as well as developers paying-as-they-go via open APIs.

💰 Revenue Streams

1) Primary source: API call revenue billed by token, providing model interfaces and usage-based billing to developers and enterprises via open platforms; 2) Secondary source: Enterprise private deployment licensing fees, providing exclusive model deployment and customized services for data-sensitive large enterprises; 3) Usage scaling: Tiered unit pricing for usage exceeding package quotas, with dedicated reserved fees for clients needing exclusive capacity.

🧮 Cost Structure

Core costs include GPU computing power procurement and cloud resource leasing (whose costs are rising due to global computing constraints), followed by model training and iterative R&D talent costs, as well as open-source model ecosystem operations and technical support investments.

🛡️ Moat

Native multimodal end-to-end joint training enables foundational semantic alignment, giving the Step series a leading edge in cross-modal document processing (mixed text-image, tables, code screenshots). The inference cost advantage brought by the MoE architecture's low activation parameters creates strong pricing competitiveness. Furthermore, open-sourcing Step 3.7 Flash attracts a developer ecosystem, establishing technical influence and standard-setting authority.

🔑 Keys to Success

  • Continuously drive down inference costs via MoE low-activation parameters to maintain competitiveness in the token price war
  • Deeply cultivate enterprise document (contracts, reports, knowledge bases) and code scenarios to build industry-specific solutions
  • Attract a developer ecosystem and expand API call volume using open-source models to feed back into the commercialization of closed-source flagship models

⚠️ Risks

  • Hallucination issues in scenarios like code generation may impact enterprise trust and adoption rates
  • Rising computing costs and global resource competition squeeze gross margin space
  • Intensifying price wars among domestic top vendors put downward pressure on unit revenue per API call

🏢 Cases

  • Stepfun completed over 5 billion RMB in Series B+ financing in 2026, initiated plans for a Hong Kong IPO, and expects a valuation in the tens of billions of dollars
  • Step 3.7 Flash was open-sourced on May 29, 2026, adopting an MoE architecture with 196.5 billion total parameters and approx. 11 billion activated, reaching a maximum inference speed of 400 Tokens/s
  • In the 2026 call volume platform rankings, Stepfun's Step3.5Flash ranked second across the entire platform, with domestic vendors occupying six of the top ten spots

📊 SWOT Analysis

Strengths

  • Trillion-parameter MoE architecture excels in complex reasoning tasks
  • Step 3.7 Flash achieves an inference speed of 400-409 Tokens/s, driving inference costs down to edge-device levels
  • Open-source models maintain strong competitiveness in lightweight domains, ranking among the top in call volume across multiple open platforms

Weaknesses

  • Flagship model has a large parameter size, leading to higher time-to-first-token (TTFT) and impacting interactive real-time performance
  • Prone to hallucinations, occasionally and confidently fabricating non-existent function names, which affects reliability in coding scenarios
  • Operating for less than three years, making it weaker than leading vendors in brand trust and large-scale enterprise deployment case accumulation

Opportunities

  • Global computing constraints drive up cloud costs, making low-inference-cost models more price-competitive
  • Enterprise-grade document intelligence and code-assistance demand continues to grow as AI applications deepen
  • Plans for an IPO in Hong Kong provide capital backing, expected to enhance enterprise customer acquisition capabilities

Threats

  • Domestic vendors such as DeepSeek, MiniMax, and Zhipu are closing in on inference costs and scenario capabilities
  • Continuous iterations of open-source models like Meta Llama and Mistral put downward pressure on the pricing structure
  • Computing resource constraints and rising cloud service costs compress profit margins