Microsoft Brings MAI Code 1.1 Flash Local Inference to GitHub Copilot

Microsoft is introducing MAI Code 1.1 Flash for local inference in GitHub Copilot, with a limited rollout planned by the end of October 2026.

Microsoft Brings MAI Code 1.1 Flash Local Inference to GitHub Copilot
MAI Code 1.1 Flash Brings Local Coding to GitHub Copilot

Microsoft plans to bring MAI Code 1.1 Flash to local coding workflows in GitHub Copilot, allowing supported Copilot experiences to run the coding model on a device instead of relying solely on cloud inference. The rollout is planned to begin with a limited audience by the end of October 2026, covering Copilot CLI, the Copilot app, and IDE integrations such as Visual Studio Code.

The development matters because it gives developers another route for AI-assisted coding: local model execution for suitable tasks, alongside cloud-hosted models when those are more appropriate. For teams that use Copilot in daily development, the practical value will depend on compatible hardware, the quality of routing between local and cloud models, and the final commercial terms. Microsoft has not disclosed pricing for local inference in its official announcement.

According to Microsoft's official overview of local models and sandboxed tools in GitHub Copilot, MAI Code 1.1 Flash is designed as a local coding model with a 256,000-token reference context window. It has 137 billion total parameters, with 6.8 billion active parameters, and uses quantization and speculative decoding to reduce its on-device footprint.

How local MAI Code 1.1 Flash works in Copilot

Microsoft describes two ways to use local inference. In Auto orchestration, Copilot determines whether a request should be handled locally or in the cloud. This routing is based on Microsoft’s HydraFusion orchestration approach, which is intended to coordinate models across edge and cloud environments.

The second option is explicit local-model selection. Users can choose MAI Code 1.1 Flash through the Windows ML provider or connect Copilot to local endpoints. That distinction is important for developers who want direct control over where a coding request is processed, rather than leaving the decision to automatic routing.

Usage mode How it works What it gives developers
Auto orchestration Copilot routes requests between local and cloud inference. A single workflow that can use the most suitable environment for a request.
Explicit local selection The user selects MAI Code 1.1 Flash through Windows ML or a local endpoint. Direct control over using a local coding model.

Microsoft is also pairing local model execution with sandboxed tool use. Its Microsoft Execution Containers, or MXC, are intended to isolate tool execution. This matters for agentic coding workflows, where a model may need to use tools as part of completing a development task rather than simply generating text in an editor.

The company had already expanded MAI-Code-1-Flash to more Copilot surfaces in June 2026, according to a GitHub Changelog entry. MAI Code 1.1 Flash extends that broader Copilot presence with a specifically announced local-coding path.

Hardware requirements and performance context

Local inference is not equivalent to running a lightweight autocomplete tool on any laptop. Microsoft’s published reference measurements use a Surface Laptop Ultra with NVIDIA RTX Spark. In that setup, MAI Code 1.1 Flash reaches approximately 75.5GB of peak memory usage at a 256,000-token context.

Microsoft reports the following decode throughput in that reference configuration:

  • 923.5 tokens per second at 64,000 tokens of context.
  • 769.8 tokens per second at 128,000 tokens of context.
  • A 256,000-token reference context window with about 75.5GB of peak memory use.

These figures show why hardware will be central to adoption. The official announcement does not provide a general minimum specification or a complete supported-device list. Businesses should therefore avoid assuming that every existing developer machine will be able to run the model locally at the reported context sizes or performance levels.

What the rollout could mean for developer workflows

For developers already working in Copilot CLI, the Copilot app, or supported IDE integrations, local inference could become part of the existing Copilot workflow rather than a separate coding assistant to install and learn. The explicit selection option may be useful when a developer wants to work with a local endpoint, while automatic orchestration is designed to reduce the need to decide manually for every task.

For businesses, the immediate question is less about replacing cloud AI entirely and more about matching work to the right environment. Local processing may be relevant when teams have compatible machines and want a local option within their existing development tools. Cloud inference remains part of Microsoft’s stated routing model, so the announcement points to a hybrid workflow rather than a fully local-only Copilot experience.

Cost is still an open detail. Although an originating social post described local calls as having no inference charge, Microsoft’s official announcement does not disclose local-inference pricing, subscription treatment, usage limits, or whether every local workflow will carry no additional charge. Teams should wait for rollout and product documentation before making budget assumptions based on that claim.

For companies evaluating where AI coding assistance fits into their software delivery process, implementation details matter as much as model performance. Scalevise can help map practical development tasks, connect AI tools to reliable workflows, and reduce repetitive operational work through AI workflow automation services. A focused assessment can identify where local and cloud AI capabilities may fit your current tools without disrupting how developers already work. Request a consultation to discuss an AI automation project.

Frequently Asked Questions

What is MAI Code 1.1 Flash in GitHub Copilot?

MAI Code 1.1 Flash is Microsoft’s coding model introduced for local inference in GitHub Copilot. Microsoft says it will work across Copilot CLI, the Copilot app, and IDE integrations.

When will local MAI Code 1.1 Flash be available?

Microsoft plans to begin a limited rollout by the end of October 2026. The official announcement does not provide a broader availability date.

Can users choose local inference instead of cloud inference?

Yes. Microsoft describes explicit local-model selection through the Windows ML provider or by connecting to local endpoints. Copilot can also automatically route requests between local and cloud inference.

Does Microsoft disclose the price of local Copilot inference?

No. Microsoft’s official announcement does not disclose pricing, usage limits, or subscription treatment for local inference. Claims that local calls have no inference charge are not confirmed in the official announcement.


Conclusion

MAI Code 1.1 Flash gives GitHub Copilot a planned local-inference option built around hybrid routing, explicit local selection, and sandboxed tool execution. The announced performance figures are substantial, but so are the reference memory requirements. Its real-world value will depend on the rollout, supported hardware, and commercial details Microsoft has yet to publish.