← Back to explorer

StepAudio 3

Type
model
Venue
StepFun blog / API release
Year
2026
Source
web
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

StepFun's StepAudio 3 is a five-model speech/audio family — Realtime, ASR, TTS, Gen, and Music — released September 15, 2026. Realtime tops the Artificial Analysis conversational-dynamics leaderboard at 98.9 with a 'Think-While-Speaking' design that reasons in parallel with speech; the family spans full-duplex voice agents to 5.5-minute music generation, though reported response latency (8.83s) lags competitors.

Keywords

audio · speech · tts · asr · music · voice-agents

Topics

audio, speech, tts, asr, music

Research notes

  • Method: Realtime rests on Deep Perception (acoustic cue modeling), Seamless Duplex (synchronized audio streams for pauses/backchannels/interruptions), and Think-While-Speaking (private chain-of-thought in parallel with speech). Music uses an MoE autoregressive model with ABC-notation arrangement planning followed by flow-matching diffusion to VAE latents decoded at 48kHz.
  • Key findings: StepAudio 3 Realtime: 98.9 on Artificial Analysis Conversational Dynamics (ranked #1), leads Speech Reasoning; Think-While-Speaking runs CoT in parallel with spoken output; StepAudio 3 ASR: 1.7% WER, tied #1 non-streaming recognition accuracy; StepAudio 3 Music: ABC-notation chain-of-thought before audio tokens; up to 5.5-min tracks; Quality Elo 1105 on Music Arena Vocals (behind Suno V5.5, Mureka); Five APIs: Realtime (full-duplex + tool calls), ASR, TTS, Gen (multi-element audio), Music
  • Limitations: Third-party coverage notes an 8.83-second response time on the same leaderboard where Gemini 3.8 Live manages 1.18s — a deployment-latency weakness. API-only release; no open weights. Music training corpus/licensing not described.
  • StepFun founded 2023 by Daxin Jiang (ex-Microsoft). Realtime technical report submitted to arXiv Sep 12, 2026. Released via StepFun open platform API.