Back

AI provider / StepFun

StepFun

A younger Chinese provider moving quickly across reasoning, audio and video, with Step 3.x now anchoring the company's public model story.

StepFun is building its public model story from two directions at once: official research releases such as Step3 and Step-Audio 2, and open checkpoints such as Step-3.5-Flash, Step-3.7-Flash and StepVideo on Hugging Face.

That makes the provider unusually readable for a younger company. You can see the frontier ambition in the Step reasoning line, but also the practical shipping motion across audio, video and adjacent experimental branches like NextStep.

Model catalog

Every release in this provider’s model line.

Each row is a distinct model release — its date, its type and the traits it is known for. Open a source to dig into any of them.

ModelReleasedTypeKey traitsSource
Step-3.7-FlashMay 2026Fast tier
  • Latest Step line
  • Flash variant
  • Current public checkpoint
Hugging Face
Step-3.5-FlashFeb 2026Fast tier
  • Flash serving
  • Reasoning line
  • Lower latency
Hugging Face
Step-Audio-R1.1Jan 2026Audio
  • Reasoning audio tier
  • Current open release
  • Speech stack
Hugging Face
NextStep-1.1Dec 2025Experimental
  • Adjacent line
  • Public checkpoint
  • Portfolio expansion
Hugging Face
Step3Jul 2025Flagship
  • 321B total
  • 38B active
  • Multimodal reasoning
StepFun
Step-Audio 2Jul 2025Audio
  • Second generation
  • Speech reasoning
  • Voice quality
StepFun
Step-Audio-AQAAJun 2025Audio
  • Audio quality assessment
  • Specialized speech
  • Evaluation branch
StepFun
Step-Audio-ChatFeb 2025Audio
  • Audio chat
  • Open checkpoint
  • Voice interaction
Hugging Face
StepVideo-T2V-TurboFeb 2025Video
  • Text to video
  • Open checkpoint
  • Turbo tier
Hugging Face
Open the model line page