A large model had to suspend new user access within three days of its launch, which is rare in the industry, but Moonshot just did so. On July 19, Kimi officially announced that due to the rapid growth of user requests after the release of Kimi K3, to ensure the experience for existing subscribed users, new consumer-side subscriptions have been temporarily suspended starting today. The existing computing power will be prioritized for serving subscribed users, ensuring their membership benefits are not affected.
This announcement was backed by numbers that exceeded the team's expectations. Kimi's team stated that since the release of Kimi K3 on July 16, the product has received far more user support than expected. However, within the past 48 hours, user request volume has significantly exceeded previous estimates and is approaching the capacity limit of the current computing cluster. Before the computing power is expanded, the platform will temporarily stop offering new subscription slots to consumer users. The announcement also made a commitment: Moonshot is currently accelerating computing power expansion. As additional computing power comes online, the platform will gradually resume new user subscriptions until full normal operations are restored.
Along with the suspension of subscriptions, a rights segmentation plan was also introduced. Kimi announced that for future new subscribers, the main Kimi rights and Kimi Code rights will be separated. The main Kimi rights cover Kimi Web, Kimi App, and Kimi Work. The platform claims this move aims to better match computing resources and ensure the user experience for different products. In other words, the increasingly resource-intensive programming capabilities will be separately accounted for.
Going back to the day of the release three days ago. Kimi K3 was officially launched on July 16, and it is currently the most powerful large model from Moonshot, featuring 2.8 trillion parameters and a 1 million Tokens context window. It mainly targets long-range programming and end-to-end knowledge work. According to information published by Kimi's open platform, K3 uses a pay-as-you-go pricing model, with input costs at 2 yuan per 1 million Tokens (cache hit) or 20 yuan (cache miss), and output costs at 100 yuan per 1 million Tokens.
More intriguingly, there were commercial feedbacks. On July 18, Zhang Yuting, President of Kimi, stated that K3's model weights would be opened soon. She also revealed that data showed that on July 17, the day of K3's release, the company's annual recurring revenue (ARR) achieved the largest single-day increase in its history. On one side, the revenue curve was suddenly spiked by demand, while on the other side, the computing infrastructure was pushed to its limits by real traffic — this suspension of subscriptions exposed the most acute contradiction of large models moving from a PPT presentation to real production environments.
When a 2.8 trillion parameter giant collides with users' genuine enthusiasm, the speed of computing power expansion has become the most urgent question, surpassing even the model parameters themselves.
