Microsoft announced at its Windows and Surface launch event on October 7, 2026, that it will introduce local AI model support for GitHub Copilot by the end of this month, allowing developers to freely switch or automatically schedule between cloud and device-side models.
To address memory bottlenecks and long context resource consumption in edge-side inference, Microsoft introduced the efficient MAI Code1.1Flash Mixture of Experts (MoE) model. This model has a total of 137 billion parameters and 6.8 billion activated parameters, combined with quantization and speculative decoding technologies, significantly reducing memory usage while improving device-side encoding response speed. It is first integrated into the new Surface Laptop Ultra equipped with NVIDIA RTX Spark hardware.

In practical applications, users can independently choose the inference mode through GitHub Copilot CLI, Copilot app, and Visual Studio Code. They can either rely on Copilot to automatically coordinate resources or specify calling local models such as MAI Code1.1Flash via Windows ML providers or compatible OpenAI endpoints. In terms of security, the system integrates Microsoft Executable Container (MXC) built by the Windows team, using native container isolation mechanisms of each operating system to ensure code execution security.
This upgrade marks the accelerating evolution of AI coding assistants from pure cloud dependency toward a hybrid architecture of "cloud + edge." This move not only simplifies the deployment threshold of local models but also provides more efficient implementation solutions for development scenarios requiring high privacy and low latency.
Join Now