Stepfun
A Chinese frontier lab building the multimodal Step series, including a very large mixture-of-experts model.
Highlights
- Step-1, Step-2, and Step-Audio model series
- Step-2 — multi-trillion-parameter mixture-of-experts model, competitive on Chinese benchmarks
- Step-Audio for voice tasks including TTS and speech understanding
- Founded 2023 by Jiang Daxin (former CEO of Microsoft Research Asia)
- Approximately $1B+ raised; multi-billion-dollar valuation
- Investors include Tencent, Wuyuan Capital, and other Chinese strategic investors
- Shanghai HQ; consumer + enterprise API products including 跃问 (Yuewen) AI assistant
External link — opens platform.stepfun.com in a new tab. Stepfun is a third-party product; we are not affiliated with it.
About Stepfun
What it is
Stepfun builds the Step family of multimodal foundation models — Step-1, Step-2, a multi-trillion-parameter mixture-of-experts model, and Step-Audio for voice. It was founded by Jiang Daxin, formerly of Microsoft Research Asia, has raised around a billion dollars, and is backed by Tencent.
Why it's different
It belongs in a directory because the frontier is not only American, and pretending otherwise gives a distorted picture of where capability sits. Stepfun's multimodal and audio work is substantive and its scale is real. What to weigh before building on it: benchmark claims from any lab are best treated as marketing until independently reproduced, and this applies across the board rather than to Chinese labs specifically. More practically, jurisdiction matters — where requests are processed and under whose legal regime is a real consideration for business data, and regulatory and procurement restrictions in some countries affect whether this is usable at all. Documentation and support in English are also thinner than the Western providers.
How people use it
It is relevant to teams building for Chinese-language markets, where the domestic labs are generally stronger than the Western models, and to anyone tracking where capability actually sits rather than where the coverage is. For most Western teams it is context rather than a provider — but the audio work in particular is worth watching, since voice is where the quality gap between labs is narrowest.
Written by the n3os team. We are not affiliated with Stepfun.
This listing was written from public information, without Stepfun’s involvement. If you own it and something here is wrong — or you would rather not be listed at all — email us and we will correct or remove it.
Get the ones worth knowing about
We write one of these for every tool worth the trouble. Get the new ones, plus what we have found genuinely useful lately.
Your address goes to Buttondown, who send the emails on our behalf. One click unsubscribes, and the list is never sold or shared.