StepFun Trillion-Parameter MoE Enterprise Agent
1) Enterprise private deployment and customized contracts (charged by seat, usage, or custom service); 2) Large model AP
Key Fields
FIELD STAMPS📌 Background
In 2026, the large model industry shifted from a technical race to commercial realization, leading to an explosion in demand for enterprise-grade documents, code, and Agents. StepFun leverages a trillion-parameter MoE architecture to reduce inference costs. In early 2026, it completed a B+ round of financing exceeding 5 billion RMB, setting a new single-deal record in the domestic large model sector, with cumulative financing reaching nearly 2.5 billion USD (company-disclosed figures, unverified). The company holds a leading market share in private deployments for finance, energy, and manufacturing, and is currently advancing its Hong Kong IPO.
👤 Target Customers
Large and medium-sized enterprises in finance, energy, manufacturing, and government sectors, as well as developers requiring API access and overseas voice model clients.
💰 Revenue Streams
1) Enterprise private deployment and customized contracts (charged by seat, usage, or custom service); 2) Large model API usage revenue (Step series models billed by token or prompt); 3) Licensing for terminal device models and revenue sharing from the core module ecosystem.
🧮 Cost Structure
Core costs include training compute for the trillion-parameter MoE model and maintenance of inference clusters. The company collaborates with domestic chip providers like Huawei Ascend and Enflame to reduce per-inference compute consumption, alongside costs for overseas commercial teams and delivery personnel.
🛡️ Moat
Engineering capabilities in cost reduction via sparse activation of trillion-parameter MoE models, with Step 3.7 Flash achieving inference speeds of 409 Tokens/s with only ~11B active parameters. This is bolstered by domestic chip ecosystem barriers and a first-mover advantage in private deployments for finance and manufacturing.
🔑 Keys to Success
- Reducing trillion-parameter model inference costs to an enterprise-acceptable range via MoE sparse activation
- Penetrating high-paying industries through benchmark private deployment cases in finance and manufacturing
- Building a compute cost barrier and promoting terminal distribution in partnership with domestic chip manufacturers
⚠️ Risks
- Continuous decline in API unit prices driven by open-source models makes maintaining high premiums for closed-source models difficult
- Risk of commercial viability being questioned post-IPO if phenomenal applications fail to emerge
- Overseas model API sales subject to geopolitical and compliance uncertainties
🏢 Cases
- StepFun Step-2 Trillion-Parameter MoE Model
- StepFun Step 3.7 Flash Open-Source Lightweight Model
- StepFun AI Xiao Cai Shen Financial Application
📊 SWOT Analysis
Strengths
- Trillion-parameter MoE architecture reduces inference costs by up to 9x compared to Dense models
- Established professional moat in the financial sector with the 'AI Xiao Cai Shen' application
- Deep integration with terminal manufacturers such as Honor, OPPO, and ZTE
Weaknesses
- Lack of phenomenal consumer-facing (C-end) applications, resulting in relatively limited market visibility
- Closed-source model faces cost-based pressure from open-source alternatives
- Overseas commercial team is still in the early stages of development
Opportunities
- Explosive growth in demand for enterprise-grade Agents and private deployments in 2026
- Hong Kong IPO provides capital injection and brand endorsement
- Domestic chip ecosystem alliance reduces compute costs
Threats
- Competition from domestic models like Moonshot AI and DeepSeek offering lower pricing
- Performance gap between frontier models narrowing to 2.7%, reducing the advantage of technical generation gaps
- High industry-wide compute bills compressing gross margins