Shandianshuo: Finding Success in Voice Input After Three Pivots
Founded: Yu Meng, Gong Zhen · Shandianshuo (Tanwei Lab)
Key Fields
FIELD STAMPSOrigin
After the ChatGPT boom in 2023, founder Yu Meng realized that in the era of large models, the only true differentiation is data—computing is commoditized, and only about a hundred people globally can truly innovate on model architecture. The team decided to build a 'data-centric memory tool' and set a rule: focus on one major direction, dedicating themselves exclusively to audio-related projects starting January 2024.
Milestones
Turning Points
- Early 2025 pivot from memory to voice input—Claude Code provided firsthand experience of the necessity of voice programming.
- Shift from 'active recording' to 'passive input'—recording things is counter-intuitive, but speaking is a human instinct.
- Moving the product focus from 'memory tool' to 'voice input' meant abandoning previous efforts to solve a more instinctive user need.
Failures & Pitfalls
- Conversational Memory Assistant: High cost ($10/month/user for GPT-4) + low frequency (counter-intuitive).
- 24/7 Recording: Captured data but the AI ecosystem couldn't utilize it; the ecosystem gap was beyond the team's capabilities.
- (Cognitive level) Big tech experience does not equate to startup ability—being a cog in a machine doesn't teach you how to earn your first million.
关键成功要素
- Find a long-term direction and brute-force it: Since Jan 2024, focus only on audio, compounding daily.
- Sell AI apps as consumer products: The 'milk tea' logic—procure large model tech + unique scenario insight + package it well to sell.
- Growth has no shortcuts: 100 micro-KOLs at a few hundred/thousand each (one viral hit covers the cost) + SEO/GEO + outsource to experts to learn how to do it yourself.
- Commercialization: Target those with money—top 3%-5% pay, 70% margin covers 100% of user costs.
- Prioritize T0 countries for expansion: US conversion rate 5%-7% with high ARPU; use open-source SenseVoice for Chinese to save costs.
Lessons
- The truth about technical moats: Sometimes the moat is simply that you survived; looking back, those who make it to the end are the ones with the moat.
- Innovation has no methodology: All methodologies should be questioned; what matters are new ideas and methods.
- Idea sourcing: Monitor top-rated GitHub projects (which reflect clear needs) + verify search volume via WeChat Index.
- Chinese users don't have poor payment habits, just different economic levels; giving physical goods (microphones) converts better than cash-back (200 RMB).
- Give yourself enough time: No one succeeds in their first year of a startup; once the direction is found, the rest is patience.
Core Data
- 智能纠错上线时间:50% (Public data, independent verification pending)
- 接口密钥用量:<20% (Public data, independent verification pending)
- 付费意向未转化数:30% (Public data, independent verification pending)
- 毛利率:70% (Public data, independent verification pending)
- 中国大陆付费用户数:1%-2% (Public data, independent verification pending)
- 台湾付费用户数:3%-5% (Public data, independent verification pending)
- 美国付费用户数:5%-7% (Public data, independent verification pending)
- 方言模型训练成本:Approx. 2 million/unit (mostly data) (Public data, independent verification pending)
- 主要语种准确率:Around 90% (Public data, independent verification pending)
Competitors / Peers
Shandianshuo's main competitors in the voice input and recording tool space include Wispr Flow, SenseVoice, and Monica. Wispr Flow uses a free strategy to capture individual users; SenseVoice leverages open-source solutions to cover over 90 dialects for Chinese recognition; Monica gains subscription revenue overseas through mobile feature differentiation. Shandianshuo's differentiation lies in binding voice input with local memory structures, using a 'data-centric' positioning to avoid a direct price war on pure recognition capabilities (Public data, independent verification pending).