OpenAI releases Astra to cyber defence customers first, with its opaque-reasoning concern unresolved
OpenAI released Astra on Thursday, sending it first to customers in its Daybreak cybersecurity programme before opening it to paid plans and the API a week later. The rollout follows two days of safety-researcher alarm over a technique that loops the model's reasoning inside itself, leaving the chain of thought harder to read from outside. OpenAI says it added extra monitoring around the release but has not addressed the underlying architecture.
Artificial Intelligence··Night
Astra ships to Daybreak customers first, on OpenAI's own benchmark claims
TechCrunch reports that OpenAI released Astra on Thursday, sending it first to customers in the Daybreak cybersecurity programme, with the Pro, Plus, Enterprise and Business plans and the API following a week later. The outlet reports that OpenAI presents Astra as scoring higher than its own Sol model and Anthropic's Fable on cybersecurity evaluations built around bug finding and terminal tasks, while noting that the comparison and the scores come from OpenAI itself and that no independent measurement has been published. Chief executive Greg Brockman calls the model the company's most capable yet and, by his account, its most aligned one to date.[1]
Two days earlier, researchers had already sounded the alarm
Two days before the release, TechCrunch and The Verge both reported that Astra loops its reasoning inside itself through a technique known as opaque recurrence, which removes the readable chain of thought that safety researchers rely on. TechCrunch quoted Redwood Research's Buck Shlegeris and Ryan Greenblatt describing the report as alarming and warning that scaling the technique further would be a natural next step. The Verge reported that OpenAI, according to The Information, limited the technique's use in Astra so researchers could keep monitoring the reasoning, and that the company's Tuesday blog post said it was deploying Astra with additional chain-of-thought monitoring without addressing the underlying architecture.[2], [3]
The launch reaches customers before the concern is settled
Astra reached customers before the safety community's concern was resolved. TechCrunch reports that OpenAI is keeping access limited to Daybreak's cybersecurity customers first, which puts the model under extra oversight before it reaches a wider audience, and that Anthropic and Google DeepMind are said to be discussing the same technique. The Verge reports that chief scientist Jakub Pachocki has acknowledged monitorability grows harder as model capabilities increase, while saying Astra's computation depth stays within a factor of two of GPT-4.[1], [2], [3]