Artificial intelligence

GitHub Copilot Will Automatically Decide Between Local and Cloud Execution

GitHub plans to expand Copilot so it can automatically route coding tasks to local or cloud models, but Microsoft has not yet clarified how much repository context might leave the developer’s device or how this routing can be restricted.

2026-10-08
5 min read
1 views
certi.news Editorial Team
GitHub Copilot Will Automatically Decide Between Local and Cloud Execution

GitHub is preparing to add automatic routing to GitHub Copilot to determine whether a coding task will be executed on a local model or sent to a cloud model, with the feature expected to arrive by the end of October 2026. The plan leaves a fundamental question unanswered: how much of the conversation history and repository context might move to the cloud when the cloud route is selected?

Microsoft disclosed the plan in a post co-written by Patrick Nikoletich, a product manager at GitHub, and Stuart Schaefer, a Windows platform partnerships engineer. The announcement coincided with the public availability of new sandboxing controls in GitHub Copilot, but the level of protection varies depending on the tools used.

Task- and Context-Based Routing

GitHub is expanding Project HydraFusion, which selects suitable models for coding tasks, so that it also determines where inference runs. According to the company, Copilot will consider task context and cache state when switching between local and cloud inference, including during multi-turn sessions.

Developers using Copilot CLI, the Copilot app, and Visual Studio Code will be able to use Auto mode or manually select a local model. The options include the MAI Code 1.1 Flash model through Windows ML, along with local OpenAI-compatible endpoints.

What Remains Unclear?

Microsoft acknowledges that “local inference does not make the session offline.” The company has not specified how much repository context or conversation history Auto mode sends to cloud models, nor whether developers will be able to inspect routing decisions or completely prevent cloud use.

This ambiguity matters to teams that enforce strict policies for handling code and data. Selecting a local model keeps the inference process on the device, but it does not prevent the agent from accessing external services or making network requests through its tools. To obtain a fully local session, developers will also need to restrict those tools’ permissions.

A Large Local Model and High Hardware Requirements

The local path relies on MAI Code 1.1 Flash, a mixture-of-experts model with 137 billion parameters in total, of which 6.8 billion are activated. Microsoft used mixed-precision quantization at approximately 3.3 bits per weight, reducing the model to 53 gigabytes, or 80% less than the cloud version in bfloat16 format. It also used speculative decoding to accelerate local inference.

The initial rollout targets Windows devices equipped with NVIDIA RTX Spark processors, such as the Surface Laptop Ultra, which provide up to 128 gigabytes of unified memory. Peak memory consumption was measured at 75.5 gigabytes with a 256,000-token context. This figure does not include the model alone; the system, applications, runtime environment, and KV cache also require additional space, while the cache grows as files are read and tool results are received.

Performance and Security Isolation

The quantized model scored 70.8% on SWE-Bench Verified, compared with 72.6% for the full-precision version. On Terminal-Bench 2.1, it scored 66.29%, compared with 62.9% for the original model, across a set of 89 tasks; the difference in this small test therefore amounts to approximately three tasks.

Copilot uses the open-source Execution Containers (MXC) library to apply isolation policies: the BaseContainer layer from ProcessContainer on Windows, Seatbelt on macOS, and bubblewrap on Linux. Shell commands and some local MCP servers and language servers, where supported, are subject to operating-system-enforced restrictions whether a local or cloud model is used.

Built-in file tools, meanwhile, run inside the agent process and rely on runtime-environment checks, while remote MCP servers remain outside local isolation. Even the demonstration Microsoft described as an offline workflow did not establish whether the GitHub data used in it was retrieved over the network or stored locally.

Why Does This News Matter?

The change is not limited to running a smaller model on the developer’s device; it moves a sensitive decision about where code is processed to an automatic routing mechanism. In practice, developers will gain greater flexibility between the speed of cloud models and the privacy of local execution, but teams will not be able to fully assess the risks until GitHub clarifies the data Auto mode sends, whether its decisions can be audited, and whether local-only execution can be enforced. The announcement therefore does not yet establish that Copilot provides a truly offline local session, nor that the current memory requirements are suitable for most development devices.

News source
The New Stack - Software Development
Open original source ↗
c
Author

certi.news Editorial Team

In the same category

You may also like

View all news