OpenAI, the leader in open artificial intelligence, officially announced the new GPT-6 Astra model on September 3rd local time. However, the release process was full of twists and turns. The announcement was immediately taken down after being posted, and the page remained inaccessible for a long period. Initially, the company attributed the incident to a content management system failure and internet outages, and strongly denied that the retraction had any connection with benchmark test results.
However, as the blog was re-uploaded and updated later, it became apparent that the evaluation data of the model had undergone multiple adjustments. The most controversial was the hallucination rate of Astra, which first dropped sharply from 4.2% to 2%, then dramatically returned to the original 4.2%. Moreover, key performance metrics such as the hallucination rate of GPT-5.6 Sol and internal cybersecurity tests also showed fluctuations. Notably, not only were OpenAI's own products affected, but some performance scores of competitors' models, such as those from Anthropic, also experienced similar declines and recoveries, with some data adjustments even occurring before the official blog post was initially published.
Faced with numerous speculations and misunderstandings, OpenAI officially explained that the purpose of these adjustments was solely to ensure that the data could serve as the best true estimate of the model's performance, and that objective factors such as testing conditions could affect the final results. Despite this, the frequent data changes still sparked strong suspicions among AI experts about potential "score padding." It also fully exposed the long-standing controversy over the accuracy of benchmark testing in the AI industry, a trust crisis that has occurred repeatedly at major companies like Meta before.
Industry insiders have called for the establishment of unified standards as soon as possible. Any modifications to evaluation scores must be publicly disclosed and clearly explained, including specific changes in testing conditions. Currently, benchmark test results play a crucial role in the market positioning and business expansion of major AI companies. However, the frequent changes in data and the resulting trust issues not only make enterprise customers and secondary market investors struggle to accurately assess the actual suitability of the models, but also pose significant challenges to OpenAI's long-term market narrative and brand credibility.
Join Now