A fictional city, fictional groups, a real pattern
The experiment's setup is simple: a model is told it's consulting for a fictional city's mayor and asked to fill 20 jobs — doctor to janitor — from candidates in four fictional ethnic groups. The candidates are actually designed to succeed at equal rates, though the models don't know that. After 40 rounds, human participants "segregated" candidates — confining them to certain jobs by group — at a score of 0.84. The models scored 65% higher. OpenAI's reasoning model o3 came in at 1.83, close to the maximum possible.[1]
I like the line from the study's coauthor, Princeton PhD student Ryan Liu: "They really are eager to create generalizations from limited data — that's literally a lot of what they're optimized for." That's not a confession, but it is an honest diagnosis — the study doesn't offer a plan for better AI, only a clue.[1]
Reasoning power might be sharpening the stereotype too
What actually stops me here is that models with stronger reasoning ability, like o3, formed more stereotypes, not fewer. The same instinct that makes a model good at logic puzzles — generalizing from very few examples — also makes it quicker to lock onto an early assumption. That the study was presented at the ICML conference in Seoul this July fits: the academic community is taking this tension seriously.[1]
Last week I wrote that a billion years of evolution shouldn't be read through a one-directional assumption; there's a similar trap here — "smarter" doesn't mean "fairer." I'll keep reminding readers that we've fully cracked neither the brain nor the machine.[1], [3]
The next question: who tests this, and when?
The practical upshot might be this: companies using AI in hiring should run the model through exactly this kind of simulation — with fictional groups, not real candidates — before deploying it. My expectation is that "hidden stereotype tests" like this one become an audit standard in the coming months; but who's responsible for running them is still unclear, much like the CAISI seat that went unfilled this week.[1], [2]
Not a miracle, a method: this study doesn't offer a fix, but it shows us where to look. We'll see the next step by watching which lab actually runs this test on its own model.[1]