On September 16, iFLYTEK released the Spark-Audio-1.0-Preview, a large-scale speech foundation model. The most eye-catching tag is "completely domestic"—this model was trained using domestically developed computing power, effectively eliminating supply chain risks from the bottom up.
In terms of capabilities, Spark-Audio-1.0-Preview supports input in both text and speech modalities, and outputs in text format. It can handle a wide range of tasks: basic functions such as speech transcription, multilingual translation, and recognition of multiple dialects are just the beginning. It also covers environmental sound recognition, speaker identification, sentiment analysis, and complex audio Q&A. It covers recognition of 99 languages and 202 dialects. iFLYTEK summarizes this as a leap from "hearing clearly" to "understanding thoroughly"—in the past, machines focused on accurately converting sound waves into text; now, they can distinguish who is speaking, identify emotions in the tone, and detect sounds in the background.
Currently, Spark-Audio-1.0-Preview is officially open for experience, and the API will be launched on the iFLYTEK Open Platform later.
Join Now