If you've been using Gemini since last year, the latest popular joke is: it's like watching your own brother gradually develop Alzheimer's. Normally, when a company releases a new model, it should be used to refute such jokes. But just now, after Google unveiled three new models, instead of retracting that statement, netizens laughed even louder, giving a sense of "they didn't believe in you, yet you turned out to be the most disappointing." These three models are 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber for cybersecurity. The wording in the official blog is still familiar: more efficient, smarter, built for large-scale AI agents, without missing a beat. However, the actual experience is drastically different.
Starting with the flagship 3.6 Flash. It's a minor version upgrade of the 3.5 Flash from this May's I/O conference, and the official positioning is as a workhorse model. The biggest selling point of this generation is saving tokens: according to Artificial Analysis Index data, 3.6 Flash uses 17% fewer output tokens than 3.5 Flash, and up to 65% less in DeepSWE scenarios. When running multi-step tasks, the number of reasoning steps and tool calls is also reduced. The price has indeed dropped, at $1.5 per million input tokens and $7.5 per million output tokens, while the previous 3.5 Flash had an output price of $9. This directly cuts down to $7.5, making it faster and more cost-effective, with the cost of a single agent task significantly reduced.

In benchmark tests, code ability was a focus, with DeepSWE rising from 37% to 49%. The official claims there are fewer unnecessary code changes, shorter execution loops, and generated code closer to production requirements; the MLE Bench for machine learning research saw a more significant increase, from 49.7% to 63.9%; OSWorld-Verified for computer operations rose from 78.4% to 83%, and computer operations have become a built-in tool in the Gemini API and enterprise edition, ready to use right out of the box. A minor but practical update is that the knowledge cutoff date has finally been pushed from January 2025 to March 2026, finally moving on from the old calendar. In terms of security, the official said that 3.6 Flash has an upgraded Frontier Safety protection, focusing on two types of abuse: chemical, biological, radiological, nuclear, and cyber attacks, with stronger resistance to jailbreaking, while trying not to harm normal needs and minimizing unnecessary rejections.

The next tier down is 3.5 Flash-Lite, which focuses on speed and affordability, targeting high-throughput and low-latency tasks, such as agent search and document processing. It's the fastest in the 3.5 series, capable of generating 350 tokens per second according to Artificial Analysis measurements. The price is extremely cheap, at $0.3 for input and $2.5 for output per million tokens, with quality significantly better than the 3.1 Flash-Lite released in March. It supports adjusting the reasoning level, where lower-end tasks can run at the lowest level for speed and efficiency, and when facing complex sub-tasks requiring multiple steps, it can adjust the thinking level upward. Computer operations have also been made into a built-in tool.
Significant improvements: Terminal-Bench2.1 for code and agent tasks jumped from 31% to 54%, GDM-MRCR v2 for long context improved from 60.1% to 72.2%, and GDPval-AA v2 for real-task execution increased from 642 to 1140. Most interestingly, this lower-tier model actually surpassed its higher-level counterpart 3Flash in many agent and code tasks, such as SWE-Bench Pro (54.2% vs. 49.6%) and OSWorld-Verified (74.0% vs. 65.1%), meaning that tasks once handled by 3Flash now have a faster and stronger alternative. The official provided several use cases, including extracting product features from massive e-commerce data, acting as a sidekick for 3.6 Flash to generate 25 web design concepts at once, batch translating and summarizing receipts, and playing and iterating to create games. Ashler, Palo Alto Networks, and Ramp customers also praised its speed, intelligence, and cost-effectiveness.
The last one, 3.5 Flash Cyber, takes a different approach, specifically designed for finding and fixing security vulnerabilities in code. Google's logic is clever: AI's speed in finding vulnerabilities has already surpassed existing systems' speed in fixing them, leading to more vulnerabilities piling up. Instead, they're using the affordable and efficient Flash to batch fix them. This model was fine-tuned from the 3.5 Flash and paired with their own CodeMender tool, where multiple Flash Cyber agents work together and then summarize into a report. On the industry-known CyberGym benchmark, it reached first-tier levels and used fewer tokens compared to larger models. However, this model is currently unavailable, considering vulnerability-finding is a double-edged sword. Google has followed in Anthropic's footsteps by adopting a security narrative, locking it tightly and only offering limited access to trusted partners for internal testing. The goal is to give frontline defenders time to fix vulnerabilities before they are exploited, while preventing people from misusing it.
At this point, you might wonder, where is the true flagship, 3.5Pro? According to the official explanation, it's still being tested with partners and will be launched once ready. However, the story behind this statement may not be as smooth. According to reports from Bloomberg and others, the code generation capability of 3.5Pro has never met internal expectations. Google updated its training data again in late June to make up for the code shortcomings, but the results still didn't improve. There are even rumors that Google completely discarded the original training plan and started over. Meanwhile, Google AI spokesperson Logan Kilpatrick openly skipped 3.5Pro in public statements, shifting the focus to Gemini4, claiming the largest and most ambitious Gemini4 pre-training has already started, with progress exciting. Given the backdrop of the flagship delay, this feels somewhat like a distraction or a promise to prolong the illusion.
Looking at the Flash models released today alone, netizens don't fully buy into them. Third-party evaluation agency Artificial Analysis stated that the two new models indeed cut the time per task in half and improved token efficiency, with the intelligent index of Flash-Lite increasing by 11 points. However, the intelligence level of 3.6 Flash compared to 3.5 Flash remained basically unchanged. The price hasn't dropped enough to ignore the capability gap, nor has the capability risen enough to justify the cost. On X platform, some users criticized that 3.6 Flash scored exactly the same as 3.5 Flash on Artificial Analysis, and it's worse than competitors like Meta Spark1.1, GLM-5.2, 5.6Luna, Sonnet5, Grok4.5, 5.6Terra, calling it simply terrible.
User Angel approached from the perspective of value for money, pointing out that the usage cost of 3.6 Flash is higher than GPT-5.6Sol medium, but its intelligence level is lower, creating an awkward combination. After all, the foundation of the Flash series is value for money, but now it's pointed out that its value is worse than competitors, equivalent to self-sabotage. As a comparison, OpenAI's Tibo has also recently spoken up, stating that due to the weekly active user reaching 10 million, the paid users of Codex and ChatGPT Work have received a new round of quota reset. This contrast and disparity, do people still remember Google Gemini3 by the lake?
Real-world testers also added fuel to the fire. User Balder directly criticized, saying the performance of these three models is worse than the previous 3.5 Flash, with problems including but not limited to basic decoding errors, strange Chinese word choices, poor image and video quality, mismatch with instructions, severe limitations, frequent memory confusion, and incorrect tool calls. He even raised a conspiracy theory, suggesting that Google did this to free up computing power for sale to Anthropic, sarcastically saying this is what Pichai wants - Cloud First. X user Conor Dart ran an AI-generated game test using Google's own Antigravity, resulting in worse wood textures and marble that looked almost similar but still couldn't be used. He directly called it another failure, saying he would rather wait for Gemini3.5 or 3.6Pro.
On one side, the official accounts highlight more efficiency, cheaper prices, and fewer tokens. On the other side, most netizens gave negative feedback, along with a series of bad news behind the models, such as the flagship delay, top talents leaving, and shrinking market value. It's hard not to wonder if Google really messed up its technology path, focusing too much on cost-saving and forgetting about improving intelligence, only relying on a few Flash models to stabilize the situation. Regardless, the pre-training of Gemini4 has already started. If this attempt also fails, the joke about Alzheimer's could truly become a prophecy.
