[KongchangAI]
· 3 min read· 1,950 words

2,500 PRs a Month: The 'Michelin Kitchen' Workflow Behind AI-Powered Engineering

2,500 PRs a Month: The 'Michelin Kitchen' Workflow Behind AI-Powered Engineering

How SpaceX AI engineer Lauren Tan ships 2,500 PRs/month using verification, environment constraints, and coordinator agents.

This article documents a deep conversation between SpaceX AI engineer Lauren Tan (ex-Meta React team) and Matt Pocock on shipping 2,500 production PRs in a single month. Lauren's methodology rests on three pillars: making verification the core skill so agents can self-check and close the loop without human intervention; refining the environment with type constraints, lint rules, and convention-driven architecture to make bad code structurally hard to write; and connecting the inner loop (agents writing code) to the outer loop (Slack/Linear context) via coordinator agents, with sampling-based review replacing manual PR-by-PR oversight.

Can AI write code — but can it actually finish the job? In a widely discussed conversation, Lauren Tan (known online as Potato) — a former Meta React team member and now an engineer at SpaceX AI (the Cursor team) — sat down with Total TypeScript founder Matt Pocock to explore a topic that set the community ablaze: how to ship 2,500 PRs to production in a single month. The number racked up roughly 3 million views on X, but the real value isn't in the figure itself — it's in the reusable engineering methodology behind it.

From "Meat Proxy" to a Ladder of Trust

Lauren's story starts with a period of burnout after leaving Meta. To decompress, she started side projects — only to find herself trapped in a strange loop: spending enormous amounts of time micro-managing individual AI agents. This was in the early days when the community was obsessed with "orchestration," and people were still hacking together custom orchestrators in the terminal.

After joining Cursor, she was tasked with optimizing performance issues in the agents window. While poring over flame graphs and heap snapshots, she realized she had become the bottleneck — "I was the meat proxy between the agent and Chrome DevTools." That frustration of being stuck in manual operations became the starting point for her entire methodology.

She introduced a core concept: the trust ladder. As your trust in an agent grows incrementally, you can delegate increasingly complex tasks — and schedule more and more agents in parallel. The key to climbing this ladder isn't writing more code; it's building an environment where agents can verify their own work.

The trust ladder analogy

One particularly interesting point: Lauren believes domain expertise doesn't lose value in the AI era — it becomes more important. As models grow smarter, the bottleneck shifts from the agent to whether you can articulate intent clearly. "A doctor or lawyer who knows a little tech and can use agents can build great products — as long as they have a clear vision in their head and can express it in a way agents understand."

Verification: The Most Important Skill in Your Toolbox

Throughout the conversation, Lauren kept coming back to one point: verification is the most important skill in your toolbox.

Her analogy is vivid — verification is like giving an agent "a pair of hands and a pair of eyes." The agent can actually run code, interact with it like a real user, debug it, and capture traces and snapshots. Without verification, an agent can't form a real loop.

"People always talk about loops, but the most important part of what makes a loop a loop is verification — because the agent can check its own work, which removes me from the equation."

This lets her practice what the labs call "hill climbing": given a rubric and a loop, an agent can continuously try to improve. She mentioned that Andrej Karpathy's auto research release follows a similar idea. Today, every application at SpaceX AI has an automatically maintained verification skill — a piece of critical team infrastructure.

Dividing Work: Deterministic vs. Non-Deterministic

Lauren also built a custom CLI to strengthen verification. At its core, this CLI extracts the deterministic parts of a skill and encodes them as scripts.

She sees an agent's work as a spectrum: on one end are tasks that require judgment (thinking through and integrating context from many sources); on the other end are mechanical, deterministic tasks (like refactoring code from one pattern to another). "You don't need the agent to think through it in a novel way every single time."

Before the CLI, every agent would "rebuild the entire world" when verifying its work — each doing its own thing, wasting context and slowing everything down. Encoding this into a CLI lets all agents share the same tooling. Matt captured it precisely: this is essentially "hiding information from the skill," keeping skills lean while offloading complex deterministic work to scripts.

The same thinking applies to tech migrations — using code mods and AST traversal for mechanical code transformations, letting scripts do the work instead of agents.

The Michelin Kitchen: A Better Metaphor Than "Software Factory"

Lauren admitted she dislikes the popular term "software factory" because "factory" tends to evoke something at odds with quality and craft. She prefers the metaphor of a Michelin kitchen.

This metaphor maps perfectly onto the trust ladder: as a "home cook," you do everything yourself — chopping, prepping, cleaning. It's a solo act. When more and more family members crowd into the kitchen, most people crack under the pressure. A head chef, by contrast, no longer cooks every dish personally; instead, they operate as the kitchen's "tech lead" or "CEO" — thinking about when to order ingredients, how to store them, when to prep.

"That's exactly how engineers write code now: you're not writing the code yourself, you have agents — but as a human, you're still responsible for the final result, and your name and reputation are still on the work." So how you set up your kitchen, configure skills, build the environment, and shape the codebase becomes the "new ingredients" for building products.

TypeScript and type constraints

Using Environment Constraints to Lift the Memory Burden off Agents

Lauren invested enormous effort into the "environment" — a sharp contrast to the attitude many people have of trusting in model magic and not daring to touch the environment. She believes that refining the environment is the core work of engineers in this new era.

Both Lauren and Matt come from TypeScript backgrounds, and she highlighted her favorite technique: type narrowing — starting from a broad type and narrowing it down to a very specific type through type guards and runtime checks. Setting constraints in a codebase works the same way: you're narrowing the space of possibilities, making it so "there's only one way to do a given thing."

Her team built an internal framework called Dune (think of it as an internal Next.js for their Electron-style app), packed with strict lint rules and convention-driven patterns — every feature lives in its own directory and is auto-discovered by a registry. This environment "makes it hard to write bad code," freeing both humans and agents from having to think about it.

The earliest versions of Grokbot consisted of 8 "god files," each at least ten thousand lines long. Lauren refined the codebase by watching how agents failed and converting those failures into lint rules. "Every time I saw an error, I stepped back and asked: how do I make this error impossible in the codebase?"

Inner Loop and Outer Loop: How the Factory Runs Itself

So where do those 2,500 PRs actually come from? Lauren clarified that she obviously didn't manually start 2,500 conversations.

The key is connecting the inner loop and the outer loop. The inner loop is her agent engineers doing work on the code. The outer loop is the stream of context flowing in from external systems — Slack, Linear, X, email — bug reports, feature requests, infrastructure constraints, and more. In the past, she had to personally ferry all of this information to her agents. Now, tools like Grokbot handle the outer loop automatically.

A question during the review segment

In practice, she has multiple Grokbot instances subscribed to various Slack channels and feeds. When a bug surfaces, it gets forwarded to Cursor's "projects" feature. A project is essentially a cloud-based coordinator agent — think "chief of staff" or "executive chef" — that doesn't do the hands-on work but rather delegates, orchestrates, and manages sub-agents. When 30 issues arrive simultaneously, the coordinator agent determines the optimal topology and dispatches tasks efficiently.

Her management philosophy draws from Netflix's "context, not control" — teach agents to be self-sufficient rather than micro-managing them. She constantly asks herself: "Where in this process am I the bottleneck? Why does the agent still need me?"

Worth noting: a large portion of those 2,500 PRs is "gardening" work rather than new features. She even has dedicated agents continuously scanning for React pitfalls — but rather than rushing to fix them, she appends findings to documentation first. "Sometimes a buffer queue is more effective than immediately dispatching an agent to fix something, because it forces you to see the big picture."

From "Tasting Every Dish" to Sampling and Autonomous Merges

To the inevitable question of how she reviews 2,500 PRs, Lauren's answer is: sampling, not tasting every dish. Just as a QA manager can't inspect every single product, she spot-checks PR quality daily. The focus isn't on correcting individual agents — it's that when she notices multiple agents making the same mistake or taking the same shortcut, she corrects the environment itself.

Fuzzing validation mechanisms

PSAC includes a "full autopilot" mode that triggers an extremely rigorous validation loop, spinning up a batch of verifier agents to fuzz each PR — actually running the app, simulating real user clicks, hunting for regressions and bugs, then self-correcting until the PR is ready to merge. This is token-intensive, but tunable (for example, dropping from 10 verifiers to 1).

It's precisely this combination of verification and environment that gives her the confidence to let agents merge code while she sleeps. "The first day I turned on the 'lights-out factory' mode was terrifying — I was worried about causing production incidents. But now I sleep much better. Agents merge their own code while I'm asleep, and in the morning I review through the commit history and roll back or fix things if I find issues."

She's also honest about the limits: this depends heavily on how verifiable your domain is. Software engineering is largely verifiable, and certain areas of mathematics (like formal proofs) are too. But for domains where it's hard to programmatically verify work and PRs are "one-way doors" (causing irreversible data loss) — certain medical, legal, or financial applications — she frankly admits "I don't have an answer." This is a problem the industry needs to explore together. She's hoping for more agent-oriented programming languages to emerge, and mentioned Bend, a language that combines programming with proofs.

What Skills Really Are: Materialized Process

When asked how PSAC and Matt's skill library fit together, both arrived at the same conclusion: a skill is fundamentally a process — a workflow translated into language (Markdown). There's no magic to it. "If there's any magic, it's only in the word choice, the phrasing, and the thinking that goes into converting an abstract process into language."

Lauren offered an extremely practical piece of advice: mine your own conversation history. Have an agent comb through all your past prompts and the moments where you had to step in and correct things, then distill those high-level learnings into reusable skills or lint rules. "Past conversations are a treasure trove of context, because they're materialized process — real workflow — not abstract ideas floating in your head."

She has a skill in PSAC called Re-Call that does exactly this: it compresses the workflow of "go look through all previous conversations and extract context" into a single skill, so you don't have to write a lengthy explanation every time. She also predicts that as model capabilities improve, skills will get smaller and more compact — last year's skills were full of implementation details (specific script commands), whereas today you can strip those out and focus purely on the workflow itself.

As Matt put it: "Every chef changes jobs and takes their knives with them." Trust, at its core, is trust in your own tools. Spend the time to sharpen them and understand them deeply — that's what makes great work possible.

Share:

Related articles