According to the latest data from OpenRouter, the total global AI large model token usage for last week (August 31 to September 6) was 115 quadrillion tokens, representing a 1.77% increase compared to the previous week. Among them, the weekly token usage of the top Chinese AI large models reached 56.72 quadrillion tokens, an increase of 2.83% compared to the previous week; meanwhile, the weekly token usage of the US AI large models was 16.54 quadrillion tokens, declining by 3.1% during the same period. The weekly token usage of Chinese large models has exceeded that of the US for nineteen consecutive weeks and remains the highest in the world.
Four out of the top five are from China, with Hy4 preview leading
Four Chinese AI large models were among the top five in terms of weekly token usage last week. Tencent's Huan Yuan Hy4 preview ranked first, with a weekly token usage of 14.7 quadrillion tokens, up 379% compared to the previous week. The model was officially released and open-sourced on August 28, with a total parameter count of 770B, an activated parameter count of 49B, and a context length exceeding 1M. It is optimized for Agent, Coding, and productivity scenarios. On September 1, the Tencent Huan Yuan team also launched a lightweight version of Hy4 preview, reducing the weight from 1.5TB to about 214GB, further lowering the local deployment threshold. The long-text comprehension capability after compression is almost equivalent to the original model.
Following are: GPT-5.6 Luna ranks second with a weekly token usage of 12.9 quadrillion tokens, up 66% compared to the previous week; Zhipu GLM-5.3 Flash climbed to third place with a weekly token usage of 12.4 quadrillion tokens, up 101% compared to the previous week; DeepSeek-V4-Flash-0731 (official version) ranks fourth with a weekly token usage of 12.4 quadrillion tokens; DeepSeek-V4-Flash-0423 (preview version) ranks fifth with a weekly token usage of 5.19 quadrillion tokens.
MiniMax M3 returns to the list, becoming the foundation for overseas indigenous model development
After nearly a month, MiniMax M3 returned to the list, ranking sixth with a weekly token usage of 5.02 quadrillion tokens, up 95% compared to the previous week. MiniMax previously announced that developers could use models such as MiniMax M3 for free through GMI Cloud and OpenRouter from August 24 to September 6. The model was released and open-sourced this June, focusing on Coding, Agent, and native multimodal scenarios, and it performed multimodal mixed training from the beginning of the training process, supporting a context length of up to one million tokens.
At the same time, last week, the Saudi Public Investment Fund (PIF)-affiliated AI company HUMAIN launched its first Arabic large model, HUMAIN M3, based on MiniMax M3 for Arabic and localization capabilities training. Unlike past Chinese large models primarily going abroad via API or application products, MiniMax M3 directly became the foundational model for overseas indigenous model development in this collaboration.
Notably, Xiaomi MiMo-V2.5, which ranked third the previous week, and Gemini 3.7 Flash, which ranked ninth, have fallen off the list.
Join Now