The test report released by the Institute of Artificial Intelligence of China Academy of Information and Communication Technology recently showed that the local large model StartLux-V1.0-27B-Preview developed by Shanghai YuanDian Starlight Science and Technology Co., Ltd. (StartLux) performed outstandingly in the Trusted AI Large Model Benchmark Test (MCP Special Test). With 27B parameters, the model ranked second in comprehensive performance, surpassing DeepSeek-V4-Flash. Its capabilities have officially entered the trillion-parameter model capability range represented by DeepSeek-V4-Pro.

This MCP special test covered six specialized tasks including location navigation, web search, browser automation, financial analysis, code repository management, and 3D design, as well as comprehensive evaluation. It focused on evaluating the comprehensive performance of large models in multi-tool collaboration, complex task execution, and real-world interaction. The test results showed that StartLux-V1.0-27B-Preview achieved a comprehensive score of 39.25, exceeding DeepSeek-V4-Flash-0731 with 284B parameters and Step-3.7-Flash with 198B parameters. In terms of the same parameter scale, it was 5.34 percentage points higher than Qwen-3.6-27B. In individual tasks, the model ranked first in location navigation and tied or led in browser automation and financial analysis with the 1.6T parameter DeepSeek-V4-Pro.

image.png

This breakthrough came from technological innovation. StartLux conducted post-training directional enhancement based on Qwen3.6-27B, and the team independently designed new, multi-dimensional, verifiable, and scalable model optimization and upgrade technology. During training, the team used the AI Training AI (Auto Research) method to conduct training experiments independently and continuously optimized strategies through feedback, making it the first local Agent model in China to use this method for post-training.

image.png

This development reflects a deep industry shift: when the capabilities of foundational models have crossed the practical threshold, the arms race centered on parameter scale over the past two years has entered a stage of homogenization. Compared to benchmark scores, enterprises and users are more concerned about actual problem-solving capabilities, data security, and controllable costs, and these needs are accelerating the rise of local large models. Currently, global tech giants such as Google, Meta, and NVIDIA are accelerating their layout of 30B-level local models, and the industry generally predicts that local large models will occupy an important market share in the coming years.

At present, StartLux-V1.0-27B-Preview can already run on consumer-grade personal computers. According to reports, the company plans to launch its first local intelligent solution within the year, while the team is also steadily advancing subsequent training and conducting in-depth research and optimization on new architectures such as diffusion language models.