2026-10-04AITao
Poteto's 2,500 PRs: Building AI Agents That Verify Their Work
Lauren Tan and Matt Pocock examine the engineering behind high-volume AI coding: runtime verification, enforceable conventions, coordinated feedback, and the human judgment, costs and limits that remain when agents merge their own work.
Contents8 sections
- Trust grows from running the software
- Turn recurring actions into reusable tools
- Encode recurring mistakes as engineering constraints
- Connect external feedback to development
- Collect issues before deciding how to fix them
- People still inspect results after agents can merge
- Verifiability sets a limit on autonomy
- Build personal skills from repeated interventions
Original video: LIVE: Poteto (creator of pstack) on shipping 1,000’s of PR’s a month at SpaceX
Matt Pocock channel · Livestream date: 2026-10-02 · Duration: 1 hour 5 minutes 36 seconds
Host: Matt Pocock · Guest: Lauren Tan, known as poteto, creator of pstack, discussing her engineering work at Cursor / SpaceXAI.
Editorial note: An original synthesis of the complete English automatic captions, with names and terminology checked against official material. PR counts and internal workflows are the guest's account and have not been independently audited.
Shipping 2,500 pull requests in a month invites an immediate question: who read all that code?
In her conversation with Matt Pocock, Lauren Tan traces the answer back to work done before code generation. She invested in agents that could run software, inspect outcomes and retrieve external context. She also kept changing the codebase to make recurring mistakes harder to produce. Sustained parallel work and autonomous merging followed as that environment matured.
The 2,500 PRs are her reported total for the previous month, including substantial maintenance, refactoring and code gardening. They do not represent 2,500 new features. The interview supplies no independently inspectable PR inventory, defect rate or cost comparison. The figure describes the scale of her activity; it cannot establish efficiency or quality on its own.
Tan prefers a Michelin kitchen to a software factory as an analogy. A home cook can handle purchasing, preparation, cooking and cleanup alone. Adding cooks requires equipment, division of labor and dependable processes. The chef takes responsibility for the kitchen as a whole while remaining accountable for the food it serves. She sees similar demands in managing multiple coding agents.
Trust grows from running the software
After joining Cursor, Tan worked on performance in its agents window. Initially, she opened Chrome DevTools herself, inspected flame graphs and heap snapshots, then relayed what she found to the coding agent.
That made her a human relay. The agent changed code; she launched the app, observed its behavior and explained the result. Every iteration passed through her, limiting how much work could proceed in parallel.
An early improvement gave the agent ways to inspect the real application: launch it, operate it as a user would, collect performance evidence and judge whether a change helped. An observed failure could lead directly to another diagnosis, edit and verification attempt.
Completion needs an observable result. Passing tests can contribute evidence, while the intended flow in the running application also needs checking. Both speakers criticize tests that merely repeat implementation logic without detecting meaningful errors. More tests do not automatically create more confidence.
The public pstack implementation makes that approach concrete. Its verification skill generator documents how to launch and operate a particular project, creates a feature map and requires an actual verification run. The maintenance skill keeps that map aligned with the changing application. Verification instructions can become stale too.
Tan says these capabilities became shared infrastructure for her team. They reduce the amount each agent must rediscover about an application and give engineers a basis for gradually delegating more. Relevant video segment
Turn recurring actions into reusable tools
Once agents could verify their work, Tan noticed another source of waste. Each task created fresh scripts to control and inspect the application. A working method from an earlier run was discarded, and the next agent rebuilt it differently.
She extracted stable, mechanical operations into command-line tools. Repeated interactions through Playwright and the Chrome DevTools Protocol could become reusable steps. The surrounding skill could focus on when to invoke them and how to interpret the result.
The distinction is practical. Combining evidence, choosing an approach and weighing tradeoffs require judgment. Fixed conversions, repeated operations and established migration rules can live in scripts. Tan points to syntax-aware code transformations as a way to separate judgment from mechanical execution during migrations.
A solved recurring operation should become a tool the next task can use. That saves context and avoids repeatedly writing and debugging disposable scripts. Relevant video segment
Encode recurring mistakes as engineering constraints
Tools improve execution. The structure of the codebase influences what agents are likely to produce.
Tan recalls early Grok Bot versions consisting of roughly 8 enormous files, each at least 10,000 lines long. This internal example comes from her account and helps explain her later emphasis on organizing code by feature.
She describes an internal framework called Dune as something like Next.js for Electron applications. Conventions, feature directories, automatic discovery and restrictive lint rules narrow the choices available when implementing a feature. The framework is not open source; the interview offers its design rationale.
When agents repeatedly append code to oversized files, engineers can change module boundaries. When a bad pattern keeps appearing, types, static checks or build rules can block it. Requirements that previously needed repeated reminders become part of the codebase.
This is what Tan means by sharpening tools. Supervising every edit consumes the time that could improve the environment. Investing in that environment can reduce repeated rework later. A recurring failure warrants examining the conditions that keep producing it. Relevant video segment
Connect external feedback to development
After an agent can complete a task reliably, another question arises: where do its tasks and context come from?
Tan separates her workflow into an inner and an outer loop. Inside are agents interpreting a goal, editing code and checking results. Outside are bug reports, requests and newly discovered constraints in Slack, Linear, X and email.
An agent working from an initial snapshot of intent can miss later information. A human then has to carry those updates back into the task. In the setup Tan describes, Grok Bot gathers external context and passes related information to Cursor Projects.
A project's coordinator organizes goals, delegates work, supplies context and tracks progress. It need not implement every change itself. Tan likens that role to an executive chef who understands the orders, knows who is handling them and can spot connections among them.
She therefore did not manually start 2,500 separate chats. Incoming feedback and existing goals became work through coordination. This describes her setup at the time of the interview; it does not establish that every user has the same product capabilities or configuration. Relevant video segment
Collect issues before deciding how to fix them
Assigning an agent as soon as a bug report arrives seems straightforward. But several reports can describe the same underlying performance problem. Treating each separately can produce duplicate patches and obscure the cause at a higher level.
A coordinator can group related reports, identify commonalities and then decide how to divide the work. Tan also describes a more deliberate routine: an agent scans for problematic React patterns and appends its findings to a document.
Every few days, she reviews that record. Several findings often turn out to be instances of the same problem. The buffer gives both people and agents a chance to recognize a pattern before immediately changing code in response to each symptom.
Faster execution still needs room for identifying shared causes. Scanning, aggregation, judgment and repair can happen at different stages. A structural problem may deserve a common fix rather than a succession of isolated patches. Relevant video segment
People still inspect results after agents can merge
Pocock eventually returns to review: did Tan examine every one of those PRs?
She describes rigorous sampling and review after changes have landed. She inspects selected PRs and code for shortcuts, workarounds and poor patterns. Recurring problems prompt changes to skills, rules and the environment. Faulty merged changes can also be reverted or modified.
According to Tan, she had more than 10 coordinators acting like chiefs of staff, covering areas such as performance and user-reported bugs, with work continuing 24/7. Some projects were exploratory, including imagining a desktop application in another language. Their activity should not all be counted as shipped product functionality.
Her autonomous merge workflow assigns verification agents to run the application, exercise user behavior and look for regressions. They repair discovered problems and repeat the checks. Once the workflow's requirements are satisfied, agents can proceed to merge, and she subsequently examines the commit history.
This has a cost. Tan calls repeated verification with multiple agents token-intensive. She gives the example of reducing 10 verifiers to 1, or asking the implementing agent to verify its own work. The amount of verification changes with that configuration.
Her confidence rests on sustained improvements to verification and the environment. She explicitly cautions that installing pstack will not immediately reproduce the setup. Reaching that point takes considerable time and deliberate adjustment. Relevant video segment
Verifiability sets a limit on autonomy
The conversation addresses a difficult boundary. Pocock distinguishes changes that can be merged and reverted from changes that might destroy data or cause consequences that are hard to undo. How should the latter be handled?
Tan leaves uncertainty in her answer. Domains with clearly verifiable outcomes make this approach easier. Work that is difficult to verify programmatically makes comparable autonomy much harder. She discusses formal verification and languages designed with agents in mind, while acknowledging that she has no general answer.
An editorial implication follows: autonomous merging needs to be considered alongside verifiability and recoverability. Passing a set of tests cannot make an irreversible action reversible, and a finite collection of checks cannot cover every consequence in a system.
PR throughput is therefore a poor basis for granting authority. More useful questions concern the evidence available for a particular task, the consequences still outside those checks and the existence of a dependable recovery path. Relevant video segment
Build personal skills from repeated interventions
The conversation closes with how the speakers' skill libraries might fit together. Tan encourages developers to assemble tools they understand and trust. Other people's skills can be studied and combined, while the resulting process still needs practical use and evaluation.
Her suggested starting point is concrete: revisit earlier agent conversations and find repeated corrections, explanations of context and manual interventions. Those records contain both failures and processes that eventually worked. They can inform new skills or improvements to types and checks.
pstack's recall skill grew from that need. While handling virtualization bugs in Cursor, Tan repeatedly wanted useful context from an earlier conversation in a new task. She turned the work of finding and recovering that context into a reusable skill.
She also expects skills to become shorter as models improve, retaining workflow and judgment while shedding some detailed command instructions. That is her view of their likely evolution. Project-specific knowledge and clear completion criteria still need to be expressed. The official pstack overview likewise emphasizes less code, better quality and parallel work supported by trust.
This connects to her earlier argument about expertise: people who understand a domain and know what a good outcome looks like are better equipped to communicate their intent. As models become more capable executors, goals, taste and acceptance criteria need greater clarity.
The interview suggests a practical sequence: identify a recurring manual intervention, give the agent a way to complete and verify that work, then preserve the proven process in tools and rules. Expand the workload as results become dependable. The chef still chooses the menu and remains responsible for what reaches the table. Relevant video segment
- Published from
- atlasnote-editorial
- Published
- 2026-10-04
- Tags
- AIAgentsengineeringinterview