Gunjo · Business Intelligence for the AI Era
← Sticker Wall JOURNEY · DETAIL

Shandianshuo: Finding Success in Voice Input After Three Pivots

Founded: Yu Meng, Gong Zhen · Shandianshuo (Tanwei Lab)

JOURNEY

Key Fields

FIELD STAMPS
IndustryAI / LLM
RegionChina
ScaleSME
ChannelOnline

Origin

After the ChatGPT boom in 2023, founder Yu Meng realized that in the era of large models, the only true differentiation is data—computing is commoditized, and only about a hundred people globally can truly innovate on model architecture. The team decided to build a 'data-centric memory tool' and set a rule: focus on one major direction, dedicating themselves exclusively to audio-related projects starting January 2024.

Milestones

2023
First Attempt: Conversational Memory Assistant Failure
A mobile memory tool using AI as the interface (for to-dos, cycle tracking, birthdays). Reason for failure: GPT-4 cost $10 per user/month (impossible to compete with free apps); recording things goes against human nature and frequency was too low.
2023
Second Attempt: 24/7 Recording Failure
Product evolved from PC client + voice recorder → mobile app → ear-clip headphones → app turning any headset into a recording device. Reason for failure: Once 24/7 audio was captured, the AI couldn't utilize it—snoring required specialized AI, coughing required medical capabilities, and emotion recognition was a separate stack. It was an ecosystem problem the team couldn't overcome, lasting from 2023 into 2024.
2025
Turning Point: Discovering the Necessity of Voice Programming Inflection Point
After the release of Claude Code, Yu Meng and the team used it all day and realized 'voice programming is incredible'—a real, high-frequency, high-willingness-to-pay scenario. This directly led to the pivot from memory to voice input.
2025
Product Finalization: Shandianshuo Voice Input PMF
Edge-first (local voice model with millisecond recognition), short-press to speak/long-press to assist, combining screen context and knowledge bases to refine output. AI correction toggle rate is ~50%, API key usage is <20%, and the top 3%-5% of paying users cover all costs, with a gross margin of ~70%.
2025
Commercialization and Global Expansion Growth
Positioned as a 'free Wispr Flow,' playing the free card in a market already educated by Wispr Flow's multi-million dollar spend. Followed a standard growth path: 100 micro-KOLs + SEO/GEO + paid ads. Commercialized via token-based distribution, prioritizing T0 countries for expansion (US conversion rate 5%-7% vs. China 1%-2%), continuing from 2025 into 2026.

Turning Points

  • Early 2025 pivot from memory to voice input—Claude Code provided firsthand experience of the necessity of voice programming.
  • Shift from 'active recording' to 'passive input'—recording things is counter-intuitive, but speaking is a human instinct.
  • Moving the product focus from 'memory tool' to 'voice input' meant abandoning previous efforts to solve a more instinctive user need.

Failures & Pitfalls

  • Conversational Memory Assistant: High cost ($10/month/user for GPT-4) + low frequency (counter-intuitive).
  • 24/7 Recording: Captured data but the AI ecosystem couldn't utilize it; the ecosystem gap was beyond the team's capabilities.
  • (Cognitive level) Big tech experience does not equate to startup ability—being a cog in a machine doesn't teach you how to earn your first million.

关键成功要素

  • Find a long-term direction and brute-force it: Since Jan 2024, focus only on audio, compounding daily.
  • Sell AI apps as consumer products: The 'milk tea' logic—procure large model tech + unique scenario insight + package it well to sell.
  • Growth has no shortcuts: 100 micro-KOLs at a few hundred/thousand each (one viral hit covers the cost) + SEO/GEO + outsource to experts to learn how to do it yourself.
  • Commercialization: Target those with money—top 3%-5% pay, 70% margin covers 100% of user costs.
  • Prioritize T0 countries for expansion: US conversion rate 5%-7% with high ARPU; use open-source SenseVoice for Chinese to save costs.

Lessons

  • The truth about technical moats: Sometimes the moat is simply that you survived; looking back, those who make it to the end are the ones with the moat.
  • Innovation has no methodology: All methodologies should be questioned; what matters are new ideas and methods.
  • Idea sourcing: Monitor top-rated GitHub projects (which reflect clear needs) + verify search volume via WeChat Index.
  • Chinese users don't have poor payment habits, just different economic levels; giving physical goods (microphones) converts better than cash-back (200 RMB).
  • Give yourself enough time: No one succeeds in their first year of a startup; once the direction is found, the rest is patience.

Core Data

  • 智能纠错上线时间:50% (Public data, independent verification pending)
  • 接口密钥用量:<20% (Public data, independent verification pending)
  • 付费意向未转化数:30% (Public data, independent verification pending)
  • 毛利率:70% (Public data, independent verification pending)
  • 中国大陆付费用户数:1%-2% (Public data, independent verification pending)
  • 台湾付费用户数:3%-5% (Public data, independent verification pending)
  • 美国付费用户数:5%-7% (Public data, independent verification pending)
  • 方言模型训练成本:Approx. 2 million/unit (mostly data) (Public data, independent verification pending)
  • 主要语种准确率:Around 90% (Public data, independent verification pending)

Competitors / Peers

Shandianshuo's main competitors in the voice input and recording tool space include Wispr Flow, SenseVoice, and Monica. Wispr Flow uses a free strategy to capture individual users; SenseVoice leverages open-source solutions to cover over 90 dialects for Chinese recognition; Monica gains subscription revenue overseas through mobile feature differentiation. Shandianshuo's differentiation lies in binding voice input with local memory structures, using a 'data-centric' positioning to avoid a direct price war on pure recognition capabilities (Public data, independent verification pending).