Where the agent runs

DoorDash built an internal platform called Flux that runs its engineering agents in cloud sandboxes instead of on developer machines. By the company's own count Flux ran 130,000 automated engineering tasks in a single month and handles more than 25,000 automated code reviews a week, with more than 300 playbooks drawing more than 10,000 invocations a week. No independent verification accompanies those numbers, and adoption at that size says nothing on its own about whether the output was right.[1]

The architecture is where the interesting part sits. Playbooks written in YAML declare the task, the tools, the permissions and the safety boundaries before a run starts; an agent gateway speaking Model Context Protocol holds scoped access to internal systems; the sandboxes themselves are Firecracker micro virtual machines that DoorDash brings up at a 95th percentile of under 5 seconds, counting the machine start, the repository clone, the build tools and the agent configuration. Runs begin from Slack, GitHub, cron, the command line or a conversation. The stated reason for leaving the laptop was processor and memory limits, a job tied to a developer's machine staying awake, and weak control over the credentials an agent was able to reach.[1]

The other half of the boundary

AWS made Agent Registry generally available, and it takes on the inventory half of the same problem. The service runs on two planes: a governance plane that keeps the full store together with compliance signals, discovery policies and custom metadata schemas, and a discovery plane that exposes only approved resources to the teams searching it. It holds Model Context Protocol servers, Agent2Agent agents, skills and custom descriptors, so one list covers what exists and another covers what a team is allowed to find.[2]

Put the two side by side and the same requirement shows through from opposite ends. DoorDash declares permissions and safety boundaries inside a playbook before the agent starts; AWS Agent Registry separates an approved discovery plane from its full store once the agent exists. Both answer the need for a scoped execution boundary and an authoritative list of what an agent is allowed to reach. There is a plainer reading too: this might be ordinary platform engineering, in which any internal service acquires a gateway and a catalogue once enough teams depend on it, with nothing specific to agents involved.[1], [2]

The number nobody published

For a team deciding whether to copy any of this, the missing figure matters more than the impressive one. DoorDash's task and code-review counts and AWS Agent Registry's compliance signals report no review burden, no defect rate and no before-and-after baseline. So 130,000 automated engineering tasks remains a measure of volume, and neither account measures what checking that work cost the engineers who did the checking. That is the figure to ask a vendor for, and a team can produce its own: hold the task set constant, count review time and reverted changes over one month with the agents and one month without, then compare those two numbers instead of the task total. If review time per accepted change holds steady while the task count climbs, the platform is doing the work DoorDash describes; if that time climbs with the task count, the work moved rather than shrank.[1], [2]