On September 17–19, StepFun (Step) released the StepAudio 3 series—five models covering Realtime, ASR, TTS, full-audio generation and Music—with the Realtime model scoring 98.9% on dynamic-dialogue benchmarks and ASR word-error rate of just 1.7%, topping real-time speech leaderboards. Separately, Step 5 Preview arrived with 600B parameters, skipping the Step 4.x line; third-party benchmark Artificial Analysis scored it 44, on par with the 2.8-trillion-parameter Kimi K3.
Step 5 Preview prices at $1 per million input tokens and $2.7 per million output tokens at 100 tokens/s, with a Token Plan starting at 49 yuan. The rapid cadence of domestic model releases—Step, Qwen, Zhipu, Kimi—reflects an intensifying second-half race where price-performance and speed, not just raw scores, decide adoption.
Industry observers noted that China's open-source-weight models are climbing global rankings, with Kimi K3 named strongest open-weight model on LLM-Stats at 93.5 GPQA, signaling a shift from commercial export to ecosystem and standards export.
Source: StepFun / Artificial Analysis / AIBase




