How 2,000 agents rewrote their own harness without breaking it

Prime Intellect's coding agent rebuilt itself in Rust using a swarm of 2,000 sub-agents, then cut cold start 13.2x and memory 82.5%. The parity checks, not the language, are the part worth copying.

How 2,000 agents rewrote their own harness without breaking it

Prime Intellect shipped a Rust rewrite of Prime Agent, and the method matters more than the language. Over two weeks, the coding agent orchestrated a swarm of more than 2,000 sub-agents across 10,000+ sandboxes to port its own TypeScript harness end to end, spending over 200 billion tokens of GLM-5.3 inference on the job. Prime Agent launched in August and has since been downloaded 300,000+ times while processing more than 8 trillion tokens of user work. Now the tool that runs agent swarms is itself the product of one.

The rewrite log reads like a list of standard complaints about JavaScript tooling. TypeScript types are optional and disappear at runtime, errors travel as unchecked exceptions, and CPU-heavy work like parsing large sessions fights keyboard input on a single event loop. Every process also pays for a JavaScript runtime and a garbage collector. Rust answers each point with a named mechanism: native code without a collector accounts for most of the memory and startup gains, Send and Sync let the compiler check which data can move between threads, and exhaustive enums with ownership rules eliminate whole bug classes before the code runs. When agents write most of the code, compile-time checks stand in for the judgment a human reviewer would otherwise spend.

The orchestration deserves the closest read. A single root agent sorted the port into a topological ordering of tasks and wrote none of the product code, keeping itself free to monitor progress and merge finished work. Each task then moved through four roles: a planner wrote the machine specification plus the parity check, an implementer built the Rust code in a dedicated worktree, a reviewer attacked the pull request adversarially on a different model in a separate context, and a verifier compiled and ran the checks in a fresh sandbox. A failed review bounced the work back with findings attached, and the change merged only after both passed. Generation and verification stayed separate on purpose, because an agent that wrote code is biased when judging it.

Parity checks are what made the autonomy safe. Four kinds ran continuously: a differential suite executed both binaries against the same scripted model and diffed the terminal frames they rendered, harness checks compared session transcripts and the requests each binary sent to model providers, protocol checks validated every daemon message type, and feature audits classified each product component as matching, partial, or missing. Objective checks let agents measure their own progress and catch regressions before merge, which is what allowed human review to shrink. The team is candid about the limit: the bugs that survived were the ones no test exercised, found by people using the build daily.

The numbers justify the effort. Cold start time to type dropped from 737.8 milliseconds in TypeScript to 55.8, a 13.2x improvement, and warm starts run 13.0x faster. Memory across the whole process tree fell from 607 MB to 106 MB, about 82.5% less, and the installed footprint shrank from 172 MB to 60 MB. First visible output now lands in 23.6 milliseconds. After parity held, a three-day hillclimb let agents profile their own benchmarks and merge 69 performance changes across 144 logged experiments, with two reviewer models confirming behavior never drifted.

The restructure may outlast the speed gains. The codebase splits into nine crates under a one-way dependency graph Cargo enforces, the largest source file sits near 2,500 lines down from about 15,000, and every session runs in its own worker process, so one crash leaves the rest running. Windows support fell out of clean platform abstraction, and the model catalog loads at runtime, so new models ship without a release. Prime Intellect is now building the same four-role pattern into the product as reusable workflows, on the theory that a defined swarm factory beats improvising one per project.

Two honesty notes belong in any read of this work. The cross-tool comparison table came from Prime Intellect's own runtime suite, and the post says plainly that without a common benchmark standard those numbers should be interpreted with caution; the before-and-after on Prime Agent itself is the solid part. And autonomous never meant unsupervised. Humans built the verification, set priorities, and reviewed the results while the swarm did the typing.

The transferable lesson has little to do with Rust. The team invested in objective checks, differential tests, and isolated sandboxes first, then set the agents loose with a score to optimize and no numeric stopping target. Tooling that survives agent-scale rewrites gets built around verification loops, and the Prime Agent paper plus the repo document the whole loop in unusual detail, parity checks included. The next codebase worth a swarm rewrite is probably already sitting in a monorepo near you.