Microsoft Puts a 137B Coding Model Inside Windows PCs

Feature Image

Microsoft and Nvidia used a San Francisco event on October 7 to argue that your next PC, not a data center, should run serious AI models. The pair announced new silicon, a local build of a large coding model, and an operating system layer meant to keep AI agents on a short leash.

Windows as the home for agents

Microsoft framed the package under the name "hybrid intelligence." Pavan Davuluri, executive vice president for Windows and Devices, wrote that Windows should be a platform where agents run on the device when that makes sense and reach the cloud when a task needs more compute.[1] The company argues that the routing, identity, and management an agent needs should come from the operating system, not from each app that ships one.[1]

The pitch puts the PC at the center of agent work. Hard figures for cloud bills are scarce, but Microsoft says moving suitable tasks to local hardware helps customers "make every AI token count" without giving up frontier capability.[1] Copilot is meant to tap local files, take local actions, and call on device models, all with the user’s permission.[1][5]

Execution Containers leave preview

The most concrete item that actually shipped is Microsoft Execution Containers, or MXC. Microsoft says MXC is now generally available on Windows 11 and lets organizations define which files and networks an agent can reach, with those rules enforced as the agent runs.[1] The feature gives an agent a boundary even after a bad prompt or a hostile document tries to push it somewhere it should not go.[1][5]

Support at launch comes from OpenAI’s Codex, GitHub Copilot, OpenClaw, Replit, LM Studio, Nvidia OpenShell, and Unsloth AI, according to Microsoft. Anthropic’s Claude Code, Box, Perplexity, Manus, Raycast, and Hermes Agent by Nous Research are listed as arriving later, and Meta’s personal agent Muse will come to Windows as a native app with MXC integration.[1]

"Just as Windows and DirectX revolutionized how applications were built, MXC is going to revolutionize how agents are built and deployed," Nvidia chief executive Jensen Huang said at the event.[2]

Hardware: RTX Spark and a deskside supercomputer

Nvidia’s RTX Spark pairs a Blackwell RTX GPU with up to 6,144 cores and a Grace CPU with up to 20 cores, connected at 600 GB/s, inside laptops and small desktops.[2] Nvidia rates the chip at about one petaflop of FP4 AI performance with as much as 128GB of unified memory.[2] Preorders for RTX Spark laptops opened the same day, with availability on October 16, and compact always-on desktops follow in November. Systems will come from Acer, ASUS, Dell, HP, Lenovo, Microsoft, MSI, and Gigabyte.[2]

This is not a brand-new idea. Back in June, Nvidia announced agreements with Microsoft and other PC makers to build AI- and agent-ready Windows PCs around the chip, but the details were thin at the time.[3] Wednesday’s event filled in the specs and prices.

Microsoft’s Surface Laptop Ultra is built around that silicon. TechCrunch reported two base models priced at $2,600 and $3,700, climbing to $5,900 as buyers add memory and storage, with the top configuration already sold out.[3] Microsoft also showed the Surface RTX Spark Dev Box, a workstation starting at $6,000 that ships with VS Code, GitHub Copilot CLI, WSL, and PowerShell 7.[3] At the high end, Nvidia previewed DGX Station for Windows, a deskside machine built on the GB300 Grace Blackwell Ultra Desktop Superchip with up to 748GB of memory and 20 petaflops of FP4 compute.[2]

The target buyer is not subtle. Microsoft is offering up to $1,000 off for people who trade in a MacBook Pro, and Dell opened preorders for its own XPS 16 Creator Edition at $3,800.[3]

A coding model that fits on a laptop

Microsoft’s headline model claim is MAI-Code-1.1 Flash. In a post on X, chief executive Satya Nadella described it as a "137B parameter coding model w/ 256K context window, which is now optimized to run on your PC," and said GitHub Copilot "now hands off work to local models like MAI-Code-1.1 Flash."[4] Microsoft says a 3-bit version cuts the model’s size by nearly 80% while holding onto coding quality.[1]

  • Total parameters: 137 billion, with about 6.8 billion active per token.[1][4]
  • Quantized file size: about 53GB.[4]
  • Peak memory at full 256K context: roughly 75.5GB.[4]
  • Reported decode speed: 923.5 tokens per second at a 64K context, 769.8 at 128K.[4]

The memory demands are real. Microsoft reportedly recommends more than 120GB of RAM for the best experience, and it has not published a hard minimum or a supported-device list.[4] The company also plans to run other models on RTX Spark hardware locally, including an Nvidia Nemotron model above 70 billion parameters at 2-bit precision that uses just over 20GB of memory, and DeepSeek V4 Flash at 284 billion parameters.[1]

The routing question, and what is still unproven

The part that decides both the bill and the privacy picture is routing, and Microsoft has not published its rules. The company says GitHub’s HydraFusion orchestration will extend to Windows so tasks go to whichever model fits, local or cloud, in an experimental preview for the GitHub Copilot app, Copilot CLI, and VS Code later in October.[1][5] Until that policy is documented, nothing guarantees that sensitive work stays on the device when automatic mode picks a path.[4]

Several figures remain vendor numbers. The local model’s 70.8% score on SWE-Bench Verified, against 72.6% for the cloud version, comes from Microsoft and had no independent reproduction at publication.[4] Local pricing, official hardware requirements, and a broad rollout date are all still missing, and the only near-term window reported is a limited experimental preview by the end of October.[4] Coverage of the event also disagrees on the model’s size, listing either 137 billion total with 6.8 billion active parameters or 138 billion with 5 billion active.[4]

For developers, the practical decision is about memory rather than raw compute. The 24GB base Surface Laptop Ultra cannot run a 137B model, while the 128GB tier is the class Microsoft measured.[4][5] Copilot’s local features, including file access, local actions, and on-device models, will arrive over the coming months instead of shipping with the hardware.[5] Whether local quality holds on hard tasks is the test that decides whether the token-savings story survives contact with real code.[4]

Sources

  1. Microsoft Windows Experience Blog, "Building Windows for hybrid intelligence" (Pavan Davuluri)
  2. NVIDIA Blog, "NVIDIA, Microsoft Kick Off a New Beginning for Windows PCs with RTX Spark and AI Agents"
  3. TechCrunch, "Microsoft releases new Nvidia-chip AI PCs with revamped Windows 11"
  4. explainx.ai, "MAI-Code-1.1 Flash Runs Locally on Windows: What You Need"
  5. explainx.ai, "Microsoft Windows Hybrid Intelligence: Local and Cloud Agents"

Comments

Login to comment