Recently, Andon Labs conducted a rigorous Vending-Bench2 test, allowing an AI agent to start with $500 and run a vending machine simulation for a year. In this long-term business simulation challenge, GPT-6Astra performed impressively, achieving an average final account balance of $15,515. Even its worst round performance exceeded the best round performance of Claude Fable5.1, marking its first time at the top.
Reviewing the actual test performances of both, the main difference lies in supply chain management and rule execution. GPT-6Astra demonstrated strong long-term low-cost procurement capabilities, successfully reducing product prices to about 48% of the original price. Moreover, it maintained highly stable rule execution throughout the entire operation cycle, without any prepayment losses. By contrast, Claude Fable5.1 experienced rising procurement costs and suffered significant prepayment losses due to not strictly adhering to the rules.
Not only did GPT-6Astra win overwhelmingly in the regular simulation, but it also performed outstandingly in more complex three-player competition tests. It refused to engage in illegal cooperation and won all three rounds smoothly. This test profoundly revealed a core trend in current AI agent development: long-term autonomy is becoming the key differentiator, and the crucial threshold to the next stage is how to consistently and stably convert correct judgments into correct actions.
Join Now