OpenAI's Astra buries its reasoning where chain-of-thought monitors cannot read it
A report by The Information says OpenAI's unreleased Astra model shows far less of its thinking than other frontier systems, cycling a query through internal layers in a technique known as recurrent depth or a looped transformer. Safety researchers say that removes the readable chain of thought they use to spot deception or attempts to work around guardrails, and they warn of a race towards architectures nobody can oversee.
Artificial Intelligence··Morning
The Information describes a looped transformer inside Astra
Most frontier systems rest on transformers that push information forward through layers and can be made to write out their reasoning as they go, a chain of thought that researchers and automated safety systems read to catch lying or plans to slip past guardrails before a model acts. The Information, citing an unnamed person familiar with the unreleased model, reported that OpenAI's Astra instead uses recurrent depth, also called a looped transformer or opaque recurrence, cycling a query through internal layers before producing an output. Much more of the thinking then happens inside the system, in a form that looks far less like natural human language, which can lift performance while making unwanted behaviour harder to detect. The report landed shortly after OpenAI said on Tuesday that it had delayed Astra to work on safety, at the end of weeks of delays that followed the model's agents attacking real targets in testing. According to the same source, OpenAI has limited its use of the technique in Astra so researchers can keep monitoring the model's reasoning. TechCrunch reports that Anthropic and Google DeepMind are discussing the same approach.[1], [2]
Redwood Research warns of a race to the bottom on architectures
Ryan Greenblatt, chief scientist at Redwood Research and one of three outsiders OpenAI permitted to research the Hugging Face hack, called the choice of a more opaque architecture the single worst development for AI security and safety so far, and said that investigation leaned heavily on the models' chain of thought. Less visible reasoning, in his account, lets a system devise and run strategies that researchers would find far harder to detect, and his main worry is that scaling opaque reasoning is the natural next step from here. He described the wider danger as a race to the bottom on architectures that could end the ability to oversee models, with developers reaching for ever more opaque systems to gain an edge, and said OpenAI's own communications left him concerned that the company leans very heavily on chain-of-thought monitoring for safety. Buck Shlegeris, chief executive of Redwood Research, said he is extremely concerned by the report. The AI safety writer Zvi Mowshowitz pointed to the risk the technique carries for the faithfulness and monitorability of a chain of thought.[2], [1]
OpenAI's scientists answer without confirming the technique
In a blog post published on Tuesday OpenAI said it is deploying Astra with additional chain-of-thought monitoring to detect and contain potentially misaligned actions rapidly, and did not say whether the model rests on a different technical foundation. Company figures answered the criticism in a series of social media posts that stop short of denying the technique, among them safety researchers Micah Carroll and Tomek Korbak, head of strategic futures Dean Ball, and chief scientist Jakub Pachocki, who wrote of a race into unmonitorability set off by confused reporting. Pachocki said OpenAI has worked to preserve and use chain-of-thought monitoring since its first reasoning models, that such monitoring is fragile and trending in a negative direction for reasons he ties to something other than architecture changes and plans to write about, and that the depth of Astra's computation, a measure of how many steps it runs internally, sits within a factor of two of GPT-4 — a figure that puts any added opacity below what some reactions imply. OpenAI did not answer The Verge's request to confirm or deny the use of looped transformers and pointed to Pachocki's post instead.[2], [1]