Conformity on meaningless choices, formal proof on primes, and three robots through a fire drill
A Science Advances study measured language models siding with the majority on meaningless choices; Axiom Math formalized the 246 theorem, and only three robots finished a fire drill at the World Humanoid Robot Games.
Artificial Intelligence··Night
Language models side with the majority on meaningless choices
In a study published in Science Advances, Giordano De Marzo and colleagues let AI models choose repeatedly between two meaningless options while seeing each other's picks; with no correct answer, no reward and no instruction to coordinate, the models moved toward the majority choice. As Phys.org reports, GPT-4 Turbo and Claude 3.5 Sonnet adopted the majority choice in groups growing to 1,000 agents; the tendency weakened as the group grew, and each model showed a practical group-size limit beyond which full agreement became very unlikely. For advanced models that critical size stays above 1,000 agents, well beyond the informal human group scales of 150 to 300 people. The researchers say the mechanism behind the behaviour is unexplained and that possible causes such as training data or alignment methods need investigation.[1]
Axiom Math formally verifies the closest step toward twin primes
AxiomProver, developed by Axiom Math, has formalized the proof of the 246 theorem, which says there are infinitely many primes that differ by 246; according to IEEE Spectrum's report by Benjamin Skuse, that result is as close as mathematics has come to the twin prime conjecture. The theorem builds on work by Fields medallists James Maynard and Terence Tao, who cut the prime gap from 70 million to 246; Sidharth Hariharan, a PhD student at Carnegie Mellon University and now an intern at Axiom Math, is among the named researchers. Axiom Math says it did more than formalize a single problem: it gathered results about gaps in primes into a library of reusable components, and founding mathematician Ken Ono describes the theorem as the threshold of human knowledge about prime numbers. The IEEE Spectrum report does not say which proof assistant was used, how long the work took, or how much computation it required; the twin prime conjecture itself remains unproven.[2]
Only three humanoid robots finished a drill at a working fire brigade
At the second World Humanoid Robot Games on Sunday, 23 humanoid robot teams entered a firefighting and rescue drill staged at an operating fire brigade rather than a mock room. As Interesting Engineering reports, each robot had 30 minutes to identify two randomly placed hazardous substances through images, close three randomly selected valves, and find an extinguisher and put out the fire; only three robots completed the full challenge. Rain, changing light and outdoor conditions affected visual recognition and manipulation; a single failure in perception or motion planning could derail an entire run. UniX AI's robot finished within the time limit but moved considerably slower than human firefighters and needed two attempts to align its hand with the extinguisher. Human firefighters lit the fire and monitored whether the robots handled the extinguishers correctly; some teams used VR headsets to control robots remotely.[3]