AI evaluation organization Artificial Analysis released a blog post on August 24, collaborating with Liquid AI to launch a local AI performance benchmark for mobile devices, focusing on the real-world performance of models on the Apple iPhone 17 Pro. The two parties have clear divisions of labor: Artificial Analysis is responsible for testing intelligence levels, while Liquid AI is in charge of testing reasoning performance. Five types of benchmarks were selected: function calls, instruction following, knowledge understanding, difficult questions and mathematics.
Small models can also demonstrate high intelligence
The test objects are models with a size no more than 8GB after 4-bit quantization, which can run on the edge. Examples include Gemma 4 E2B and LFM2-2.6B-Exp. When the context limit is 16K tokens, the highest average scores were achieved by Nanbeige4.2-3B from Nanhua Lab and LFM2.5-2.6B from Liquid AI. If the response time is limited to one minute, LFM2.5-8B-A1B takes the top spot, followed by LFM2-2.6B-Exp, Gemma 4 E4B, and Granite 4.1 8B.
Nanbeige4.2-3B was developed by BOSS Zhipin's laboratory. It has only 3 billion non-embedded parameters but uses the Looped Transformer architecture to reuse layers cyclically, increasing the effective computational depth without increasing parameters, achieving scores comparable to its competitors. This benchmark brings the intelligence ceiling of "edge small models" to the forefront, also indicating that the competition for AI on mobile devices is shifting from simply being able to run to running faster and answering more accurately.
Join Now