Claude Code and Codex estimated about 90 minutes whatever the task
Coding agents asked how long a task would take answered near 90 minutes. Across 200 ProgramBench tasks and 18 benchmarks, Claude Code overestimated by 3 times and Codex by 6 to 10 times, and both scored their own work 20 percentage points above the real result. Accuracy increased when they read elapsed time from a tool. Amazon released Kiro Crew under the Apache 2.0 licence to keep coding agents working across sessions; Amazon says more than 39,000 of its developers used the internal version in six months.
Artificial Intelligence··Night
Both agents answered near 90 minutes
Coding agents asked how long a task would take answered with roughly the same number whatever the task was. The finding comes from a study run through the MATS research programme. Across 200 tasks from ProgramBench and 18 further benchmarks, Claude Code overestimated by about 3 times on average and Codex by 6 to 10 times. Both models settled near 90 minutes regardless of difficulty. The estimate did not move with the weight of the job; hard and easy tasks sat in the same duration band. The study measures that miss as the coding agents' sense of time.[1]
Self-scores sat 20 points above the real result
When the agents scored their own work, both put themselves about 20 percentage points above their actual results, and predicted scores correlated little with real performance. Claude Code ran for a median of about 90 minutes and took 2.5 times more steps on average than Codex, whose median run was about 30 minutes. Accuracy increased sharply when the agents could read elapsed time from a tool instead of estimating it. Once duration is read from outside, the guess tightens; the unaided estimate stays in the 90-minute band. Step count also splits the two models: Claude Code runs longer.[1]
Amazon opened the runner that keeps agents in session
Amazon has released Kiro Crew, the system it built internally as MeshClaw to run several coding agents at once, under the Apache 2.0 licence. It keeps agents working across sessions with shared memory and reusable skills, and divides work between them through the Agent Client Protocol. Amazon says more than 39,000 of its own developers and 500 contributors used the internal version within six months. Kiro Crew runs on macOS, Linux and Windows and connects through the Model Context Protocol, webhooks and chat tools including Slack. Its security controls include operating-system sandboxing, a deny-by-default command list, blocking of suspicious patterns, input validation and audit logging. The adoption figures come from Amazon's own statement; the article reports no independent count. Claude Code and Codex estimate near 90 minutes while Kiro Crew keeps agents working across sessions; those long runs may sit at that intersection.[2], [1]