From the source
Microsoft announced that GitHub Copilot will gain the ability to automatically route coding tasks between local on-device models and cloud-scale models, starting with the quantized MAI Code 1.1 Flash model on NVIDIA RTX Spark Windows PCs.
The local model achieves 70.80% on SWE-Bench Verified and 66.29% on Terminal-Bench 2.1 with a peak memory usage of 75.5GB at 256k context.
Developers can let Copilot orchestrate model placement automatically or explicitly select a local model.






