Eigen RadarAI
Analysis

OpenAI limits Astra's advanced cybersecurity access at a critical threshold

OpenAI's Path to Astra post says the upcoming Astra model is the first to meet the Critical cybersecurity capability threshold of its Preparedness Framework. OpenAI has described Astra as the first large language model to meet that threshold, saying it can find unknown vulnerabilities and exploit them without human guidance. The company says it plans to release Astra soon and to keep the most advanced cybersecurity capabilities behind restricted access. OpenAI reports a perfect ExploitBench score and two zero-day findings in a modified internal test.

Artificial Intelligence··Morning
In a bright computing hall, cables toward dark server cabinets sealed behind glass end across a visible gap as a technician holds a mechanical lever.

OpenAI places Astra at its Critical cybersecurity threshold

OpenAI's Path to Astra post says the upcoming Astra model is the first to meet the Critical cybersecurity capability threshold of its Preparedness Framework. Under that framework, OpenAI says Critical means a model can identify and develop working zero-day exploits against hardened real-world systems, or devise and run end-to-end novel attack strategies from a high-level goal, without human intervention. OpenAI has described Astra as the first large language model to meet its critical cybersecurity threshold. The company says the model can find unknown vulnerabilities in computer systems and exploit them without human guidance. OpenAI says it still plans to release Astra soon.[1], [2]

OpenAI reports a perfect ExploitBench score and two zero-days

Capability figures come from OpenAI. The company reports a perfect score on ExploitBench, a test that measures model hacking ability. OpenAI says that in a modified internal test the model discovered and exploited two zero-day vulnerabilities. Those two findings sit on that modified internal test rather than on a public trial. No independent test of the claims has been published. OpenAI's Path to Astra post is the company's own account of its evaluations.[1], [2]

OpenAI keeps Astra's strongest cyber tools behind extra checks

OpenAI says it plans to keep the most advanced cybersecurity capabilities behind restricted access while the safeguards it describes are in place. Access to the most advanced cybersecurity capabilities, the company says, stays more limited than access to the rest of the model. The announced safeguards include better detection of abuse and of jailbreaks, restrictions on accounts assessed as higher risk, and additional chain-of-thought monitoring. Yona Shavit, formerly at OpenAI and now at the OpenAI Foundation, asked publicly whether the model complied from knowing what was expected of it.[1], [2]

References

  1. News sourceTechCrunchOpenAI puts Astra past its own critical cybersecurity threshold and limits who reaches the top tier↩1↩2↩3
  2. News sourceOpenAIAstra is the first OpenAI model to meet the Critical cybersecurity threshold, and its safeguards are set out↩1↩2↩3