7 min readRishi

GitHub Rewrote the Copilot Runtime in Rust With Agents. The Playbook Is the Story.

GitHub Rewrote the Copilot Runtime in Rust With Agents. The Playbook Is the Story.

On September 16, 2026, Stephen Toub published the account of porting the GitHub Copilot agent runtime — the engine behind Copilot CLI, the Copilot app, and the Copilot SDK — from TypeScript on Node.js to Rust. By August 21 it was 832,378 lines of production Rust plus 468,689 lines of Rust unit tests, written mostly by coding agents across 128 pull requests, primarily by one developer, in roughly fourteen and a half weeks. The attributed token spend was about $120,000 plus about three weeks of that developer's time.

The numbers are the headline. The method is the part worth copying, and it has little to do with Rust.

Why a runtime, not a CLI, needed the port

The runtime started life as the engine of a terminal UI. TypeScript on Node with V8 is a fine choice for that. It is a poor choice for a component meant to be embedded by six SDKs (C#, TypeScript, Python, Rust, Go, Java). Every SDK consumer shipped a second language runtime, on the order of 100 MB of working set minimum, for an engine their application otherwise had no use for. Every message and every session file-system read crossed a process boundary. A crash in Node took the session with it. Operators had two processes to supervise.

The requirement was an engine that embeds through a C ABI, starts fast, and uses memory predictably. Rust met that. The post is explicit that this is not a claim every large TypeScript program should become Rust. It is a claim about those requirements.

The scoping number was right and misleading

The May 2026 plan measured the runtime at about 130,000 lines of TypeScript. That was accurate for scoping and wrong as a picture of the work. During the port, tens of agent-assisted developers were merging hundreds of pull requests a week, adding roughly 300,000 lines of production TypeScript while the port removed about 430,000. About 1.2 million lines of Rust entered and about 365,000 left. The TypeScript line count on the chart looked flat for months. It was hiding churn in both directions.

If you are estimating a rewrite from cloc on the day you start, add the lines the team will write while you port. On a busy repo that is not a rounding error.

The decision that made it shippable: in place, atomic, no shadow copies

The post lays out four options. Big bang in main (stop the world). Big bang on a branch (rewrite tries to keep up with main). In place with atomic replacement (each PR swaps one component from TypeScript to a thin shim over Rust and deletes the old code). In place with A/B (keep both implementations hot-swappable until confidence plateaus).

They chose atomic replacement. The reasons are worth quoting in spirit:

  • main is always shippable. Each PR removes a TypeScript component and lands a shim that calls Rust in one change. The new code runs in situ immediately.
  • Each PR is one component or slice, small enough to review by a person, an agent, or both.
  • Small ports drift less against concurrent PRs. Components too large to port were refactored into portable pieces first.

They rejected A/B deliberately. Keeping two implementations in two languages with two dependency sets, in a repo merging hundreds of PRs a week, is its own source of regressions. The subsystems that would benefit most from a parallel cutover — session orchestration, which owns mutable state and drives callbacks in both directions — are exactly the ones that cannot be shadowed with an if/else on a flag. The coupling that makes a component hard to port is the coupling that makes it near impossible to run twice.

Releases were the test harness

Over the fourteen-and-a-half-week window, main shipped 135 releases: 100 pre-release and 35 stable, about 1.3 a day, roughly matching 1.3 port PRs a day. Each release carried a small, knowable set of ported components. Pre-release versions were about 10.5% of downloads in a trailing seven-day sample, so first exposure was limited and mostly first-party inside Microsoft and GitHub while feedback channels were watched.

That is the alternative to A/B. Confidence came from shipping small increments to a small audience often, not from running two engines side by side.

The regressions, by category

Dozens of regressions were traced and fixed by September 14. The post groups them, and the groups are a checklist for any port:

  • Serialization at the boundary. A repository id emitted as a float; a timestamp with a trailing .0 that every strongly typed SDK on the other end rejected. The compiler cannot know the wire contract.
  • Lifecycle and ordering. A queue protected by locks where two senders each concluded the other would drain it. Events moved safely between threads and arrived in the wrong order. A synchronous napi function that was memory-safe and blocked Node's main thread for a minute.
  • Ships passing in the night. Rebase drift. Thousands of rebases in a repo with hundreds of weekly PRs; a very high success rate still leaves failures, including a rebase that quietly removed a guard and its test.
  • Slowpokes. Functionally correct, slower. One read-only scan deep-copied a 260 MB event log instead of borrowing it. Another retained an async handle per event until V8 exhausted its heap. Redundant serialization, locking, polling, and native-to-host crossings at the Rust-TypeScript boundary.

Every regression compiled. The post makes the point directly: "if it compiles, it's correct" is a meme, and the known-regression list is the rebuttal. A compiler checks that the program you wrote is internally coherent. It cannot check that you wrote the whole program, preserved the old contract, or did the work at an acceptable cost.

Public issue trackers for copilot-cli and copilot-sdk, classified by bug-like labels and titles from January through August, showed no quality spike during the migration. That is not an availability metric, and the author says so. It is a useful check that the product-facing channels did not notice.

What it bought

End to end through the C# SDK, against a pre-port build with the TypeScript runtime reached over stdio: creating a client, a session, one turn, and teardown dropped to about 55 milliseconds in-process. A pressure test of 1,000 one-turn session lifecycles went from 7.55 per second on the TypeScript CLI to 57.45 with Rust out of process and 120 in process. Measured ten-client memory delta fell 91%. Other changes landed in the same period, and the post says to take the numbers with salt. It also says this is the baseline port, with TypeScript-shaped algorithms rendered faithfully in Rust and the redesign work still ahead.

The parts you can reuse without Rust or agents

  1. Pick in-place atomic replacement over a long-lived branch when the repo is busy. Keep main shippable every day.
  2. Refactor before porting when a component is too big to swap in one PR.
  3. Make releases small and frequent enough that each one carries a knowable set of changes, and route early exposure to an audience that will report back.
  4. Write the boundary contract tests the compiler cannot: serialization shapes, ordering, blocking behavior, and a performance budget per operation.
  5. Count rebases as a risk, not an inconvenience. A guard that disappears in a rebase is a regression with no diff anyone meant to write.

The full write-up is on the GitHub Blog. Read the regressions section twice. It is the honest half of the story and the half most rewrite posts leave out.

Keep reading

Newsletter

New posts, straight to your inbox

One email per post. No spam, no tracking pixels, unsubscribe anytime.

Comments

  • No comments yet. Be the first.