AI provider / StepFun
StepFun
A younger Chinese provider moving quickly across reasoning, audio and video, with Step 3.x now anchoring the company's public model story.
StepFun is building its public model story from two directions at once: official research releases such as Step3 and Step-Audio 2, and open checkpoints such as Step-3.5-Flash, Step-3.7-Flash and StepVideo on Hugging Face.
That makes the provider unusually readable for a younger company. You can see the frontier ambition in the Step reasoning line, but also the practical shipping motion across audio, video and adjacent experimental branches like NextStep.
Model catalog
Every release in this provider’s model line.
Each row is a distinct model release — its date, its type and the traits it is known for. Open a source to dig into any of them.
| Model | Released | Type | Key traits | Source |
|---|---|---|---|---|
| Step-3.7-Flash | May 2026 | Fast tier |
| Hugging Face |
| Step-3.5-Flash | Feb 2026 | Fast tier |
| Hugging Face |
| Step-Audio-R1.1 | Jan 2026 | Audio |
| Hugging Face |
| NextStep-1.1 | Dec 2025 | Experimental |
| Hugging Face |
| Step3 | Jul 2025 | Flagship |
| StepFun |
| Step-Audio 2 | Jul 2025 | Audio |
| StepFun |
| Step-Audio-AQAA | Jun 2025 | Audio |
| StepFun |
| Step-Audio-Chat | Feb 2025 | Audio |
| Hugging Face |
| StepVideo-T2V-Turbo | Feb 2025 | Video |
| Hugging Face |
Timeline
A practical timeline of model releases and company milestones.
This page tracks the moments that matter most: major model releases, version jumps and company moves that changed how this provider shows up in the AI market.
Step-3.7-Flash becomes the latest public checkpoint in the Step line.
The newest release suggests StepFun is iterating the reasoning family quickly while keeping a public open-checkpoint rhythm.
Step-3.5-Flash shows the reasoning line splitting into lighter serving tiers.
StepFun starts turning the Step family into a more deployable ladder instead of a single heavy model track.
Step3 becomes the public centerpiece of the reasoning line.
The multimodal reasoning release gives StepFun a clearer flagship story and a stronger frontier position.
Step-Audio 2 upgrades the company's speech stack.
The second-generation audio release turns the StepFun voice line into a more serious product family, not only an experiment.
Step-Audio-Chat makes voice interaction public from the start.
Audio arrives as an open-facing part of the StepFun story very early, not as a late add-on to the reasoning line.
StepVideo-T2V-Turbo gives StepFun an open video checkpoint early.
The release shows that StepFun is willing to ship public media models alongside its more research-heavy main line.