Not a miracle, a measurable tendency

The study, presented at ICML in Seoul, doesn't overreach its claim: human participants scored 0.84 on the segregation scale, the models tested scored 65% higher on average, and o3 reached 1.83 — close to the maximum possible. This isn't a 'AI is racist' headline; it's a measured tendency.[1]

I liked the scientists' honesty: Princeton's Ryan Liu says 'LLMs really are eager to create generalizations from limited data' — that's not an accusation, it's a description of a method. It's also notable that more capable reasoning models showed stronger bias: as capability rises, so does the speed of generalizing.[1]

The same tendency runs the opposite direction in materials discovery

CuspAI's MIRA platform is also, in effect, generalizing fast from limited data — but here the target isn't categorizing a person, it's finding a new material for chip manufacturing. The 'AI Materials Foundry,' formed by more than 48 organizations including Nvidia and Meta, aims to speed up work that traditional research would take years to do.[2]

What a billion years of evolution did in biology, computation is trying here in a few weeks — but the claim needs to be read with its own limit attached: the company pivoted from an initial focus on carbon capture and water purification a year ago, in response to demand. The ease of that pivot is also a sign the platform doesn't yet have one settled area of expertise.[2]

The method is the same; the question always circles back to the same place

In my July 19 column I wrote that WISeR offered 'no miracle, only a clue.' Today's two studies show the same discipline: neither the hiring research nor CuspAI separates its claim from its method. One finds a way to reduce bias — giving the model a diversity-oriented target; the other points a foundry at a target.[1], [2], [3]

What scientists keep teaching us is this: generalizing power is a neutral tool, and its direction depends on the target a human gives it. What I'll track next is when CuspAI's foundry announces its first concrete material result — that's when we'll see whether the method actually works, not just demos well.[2]