OpenAI proposes cutting Cursor off its models as a study shows the harness sets benchmark scores
On 28 August OpenAI proposed ending Cursor's access to its models on 12 November 2026 after SpaceX bought Anysphere. A preprint published on arXiv in the same period held the model fixed and reported that changing only the coding-agent harness raised complete solutions on SWE-bench Verified from 43 to 72. The licensing dispute and the measurement finding both point to how much of a coding agent stack lives in the harness, not only the model.
Artificial Intelligence··Midday
OpenAI announced a November cut-off
OpenAI said on 28 August that it is ending the agreement under which the coding tool Cursor serves its models, and proposed 12 November 2026 as the date access would stop. SpaceX completed its purchase of Anysphere, Cursor's parent company, earlier in the month, in an all-stock deal announced in June at 60 billion dollars. OpenAI says it cannot be confident SpaceX would keep to its terms of service, citing its experience with Elon Musk's companies. The contract's change-of-control clause gave OpenAI a limited window to cancel once the acquisition was complete. Michael Truell, a Cursor co-founder now at SpaceX, said the two teams are talking to resolve the notice, and put OpenAI models at roughly 5 percent of Cursor's user traffic. OpenAI's account of earlier contract trouble with Musk-owned companies is reported as its own claim, while Truell frames the models as a small share of overall traffic. The reported findings come from the authors' own runs, and the work is presented as a preprint that has not been peer-reviewed.[1]
The preprint jumped scores by changing the harness
A preprint published on arXiv on 29 August held the model fixed and changed only the harness around it: as the context window filled, the experimental configuration abbreviated older tool outputs instead of carrying them whole. On SWE-bench Verified, run over 169 tasks in a 20,480-token window, mean per-task fail-to-pass rose from 28 percent to 49 percent and complete solutions from 43 to 72. Under a constrained context the experimental harness improved the fail-to-pass fraction in all three comparisons and raised complete solutions on both SWE-bench Verified and SWE-bench Pro. With a wide window the two configurations came out close to each other, except on FeatureBench, where the experimental one stayed ahead. The author reports the same effect across several models without model-specific reconfiguration and concludes that a coding-agent evaluation should treat the model and the harness together as the solver under test. The paper is a preprint on arXiv and has not been peer-reviewed.[2]
Cursor's small traffic share carries a larger measurement debate
The roughly 5 percent share Truell gives for OpenAI models shows much of Cursor still rests on other providers, yet OpenAI's proposed November cut-off targets one of the coding-agent market's most visible model supply lines. The preprint finding shows that, with the same model held constant, scores depend heavily on how the harness manages context, so licensing fights and benchmark arguments can meet on the same product surface. OpenAI's account of earlier contract trouble with Musk-owned companies is reported as its own claim, while Cursor says talks continue. For readers the concrete development is a proposed access date arriving alongside a measurable harness effect in the same week. The reported findings come from the authors' own runs, and the work is presented as a preprint that has not been peer-reviewed.[1], [2]