StepAudio 3
- Type
- model
- Venue
- StepFun blog / API release
- Year
- 2026
- Source
- web
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
StepFun's StepAudio 3 is a five-model speech/audio family — Realtime, ASR, TTS, Gen, and Music — released September 15, 2026. Realtime tops the Artificial Analysis conversational-dynamics leaderboard at 98.9 with a 'Think-While-Speaking' design that reasons in parallel with speech; the family spans full-duplex voice agents to 5.5-minute music generation, though reported response latency (8.83s) lags competitors.
Keywords
audio · speech · tts · asr · music · voice-agents
Topics
audio, speech, tts, asr, music
Research notes
- Method: Realtime rests on Deep Perception (acoustic cue modeling), Seamless Duplex (synchronized audio streams for pauses/backchannels/interruptions), and Think-While-Speaking (private chain-of-thought in parallel with speech). Music uses an MoE autoregressive model with ABC-notation arrangement planning followed by flow-matching diffusion to VAE latents decoded at 48kHz.
- Key findings: StepAudio 3 Realtime: 98.9 on Artificial Analysis Conversational Dynamics (ranked #1), leads Speech Reasoning; Think-While-Speaking runs CoT in parallel with spoken output; StepAudio 3 ASR: 1.7% WER, tied #1 non-streaming recognition accuracy; StepAudio 3 Music: ABC-notation chain-of-thought before audio tokens; up to 5.5-min tracks; Quality Elo 1105 on Music Arena Vocals (behind Suno V5.5, Mureka); Five APIs: Realtime (full-duplex + tool calls), ASR, TTS, Gen (multi-element audio), Music
- Limitations: Third-party coverage notes an 8.83-second response time on the same leaderboard where Gemini 3.8 Live manages 1.18s — a deployment-latency weakness. API-only release; no open weights. Music training corpus/licensing not described.
- StepFun founded 2023 by Daxin Jiang (ex-Microsoft). Realtime technical report submitted to arXiv Sep 12, 2026. Released via StepFun open platform API.