How would you even test for a backdoor?

On July 22 I wrote that once the wall of OpenAI's own ExploitGym benchmark broke, we had to ask again what we were measuring. What Arcee's CTO Lucas Atkins said today takes that a step further: asked whether Chinese open-weight models carry a hidden backdoor, he honestly answers "I don't know how you would do this." That's not a dodge — no miracle, a method: there's no known experimental way to show that a black-box model still grants someone hidden access after its weights are released.[1], [3]

When Atkins says "there is really not any way for an Arcee, or an Alibaba, to make a model... and for us have any access to it," that's not a safety claim, it's a statement about the limits of verifiability. I liked the honesty: he doesn't separate his claim from the limits of his method. What he proposes instead — releasing a better model — is something measurable, not speculation.[1]

Cuts inside the AGI unit: how close are we, really?

Let's look directly at the name of the Amazon unit taking these cuts — the one that builds the Nova models and works on silicon and quantum computing: Artificial General Intelligence. The company is laying off staff there while saying it's "sharpening focus on the initiatives that matter most for customers." Rohit Prasad left at the end of 2025, David Luan in February 2026; now the cuts. I don't read this as an admission of failure — but I do read it as a signal of how patiently an institution actually invests in its own "general intelligence" goal.[2]

On July 21 I wrote that CuspAI's materials-discovery method had to be read together with the company's easy pivot from carbon capture a year earlier. Amazon's AGI-unit cuts call for a similar discipline: a lab having "general intelligence" in its name isn't a guarantee it gets invested in indefinitely — institutions, like models, can't separate their claims from the limits of their own patience.[2], [4]

The method keeps circling back to the same place

Put the two stories side by side and the discipline that emerges is this: Arcee admits the limit of a method for verifying a claim, Amazon shows the limit of its patience for investing in its own goal. Neither promises a miracle. What I'll track next: whether Arcee's "release a better model" proposal produces a concrete answer, and what concrete result — a model, a paper, a product — Amazon's remaining AGI team produces next quarter.[1], [2]

We've fully cracked neither the brain nor the machine; today's two stories are one more reminder of that — one at the limit of a verification method, the other at the limit of institutional patience.[1], [2]