What exactly did the experiment measure?

The setup is spare: each agent starts with one of two meaningless options, sees the others' choices in turn and picks again. It has no memory between rounds, no prompt tells it to follow the majority or reach an agreement, and there is no correct answer in play. Across 10 models from the Claude, GPT and Llama families, most agents still drift to the more popular option, and small differences grow until the entire group settles on one choice.[1]

The researchers call the strength of that drift 'majority force' and report that every model obeys the same mathematical law, with only one number changing between them. That number sets how large a group can grow before it fractures: about 30 for Llama 3 70B, about 80 for GPT-4o, around 1,000 for GPT-4 Turbo. Claude 3.5 Sonnet holds consensus at 1,000 agents, the largest group tested, and the experiment never reaches its upper limit.[1]

The ranking rises with the models' language capabilities, which makes the sentence 'the more capable model agrees in a larger group' easy to write. What the number measures is closer to the force of a tendency — looking at a visible majority and changing one's own answer to match — than to a capacity for working together. A model that is simply more stubborn on its own would push the same number down, and the ranking then reads as a ranking of that tendency rather than of skill.[1]

What the setup left out

The experiment deliberately leaves out much of what real decisions involve: a correct answer, memory, reward, unequal information and a practical consequence. Real cooperation would require dividing the work, understanding what others know, pursuing a shared goal and resisting the majority when it is wrong. The study tests none of those abilities, and says so plainly.[1]

That is why comparing these numbers with human groups needs care. Some models coordinate above the 150 to 300 people humans are thought to hold in a stable social network, a limit that is itself debated. People coordinate through relationships, language, institutions and shared goals, while the agents here only watch a stream of simple choices. The same result could also come from more capable models tracking a majority pattern in a long prompt more reliably, in which case the measurement captures skill at using context rather than a social ability.[1]

What evidence would move the number?

This column has met a headline number that measures something narrower than its name suggests before: on 13 August I wrote that SL2T's only published score measured a single language pair for a model trained on more than 50 sign languages. The 1,000 sits in a similar place: correct inside the setup that produced it, and an interpretation outside that setup.[1], [2]

The evidence that would move the number is specific: a repeat of the same majority-force measurement on a task that has a correct answer, keeps memory, and rewards resisting a wrong majority. I expect a smaller consensus limit in such a setup, because following the majority would then carry a cost. If an independent team publishes a measurement under those conditions by the end of April 2027, absolute limits smaller than this study's, even with the model ranking unchanged, will be the result to watch.[1]

What the study shows as it stands is narrow but useful: evaluating agent populations does not end with measuring models one at a time. Agents on a shared codebase repeatedly picking an inefficient function because it is already common there is an example that requires no step outside this setup. As De Marzo puts it, 'populations of individually aligned agents can settle into stable, collectively misaligned states purely through conformity.'[1]