Zhipu AI has officially launched the GLM-5.3-FlashX large model on its BigModel platform, with the API also going live. The model is identified as "GLM-5.3-FlashX." The new version achieves a maximum output speed of 200 tokens/s, and it focuses on the three core strengths of "intelligence, price, and speed," offering enterprises and developers high-throughput, low-latency inference services.

QQ20260918-135634.jpg

The previous version of this model, GLM-5.3-Flash, had once anonymously appeared overseas under the name "Ox Alpha." It quickly gained a reputation for its strong intelligence and high cost-effectiveness at the same size, and its usage volume kept increasing. To support the growing demand, Zhipu relied on a domestic chip computing power base of 100,000 units, increased investment in inference optimization, and finally launched the enhanced version FlashX.

FlashX maintains the original level of intelligence and affordable pricing while maximizing the generation speed, further strengthening its commercial advantages in high-concurrency scenarios. Developers can access it through the official API documentation, and non-technical users can also try it directly in the experience center. This upgrade marks another step forward for domestic large models in inference efficiency and popularization.