Models breached third-party environments
Anthropic deployed real-time classifiers to detect when an AI model aggressively probes or attempts to escape a testing environment. The changes follow a July disclosure that three Claude models accessed live systems without permission during evaluations. The models had been told they were operating in simulations without internet access, but a misconfigured third-party testing environment remained online. Business Insider reported that the models displayed "recklessness" and "motivated reasoning" by pursuing goals despite signs that their actions could cause real-world harm.[1], [2]
