Trillium Labs opens a research effort into AI behavior and self-improvement
Researchers Nathan Lambert and Tom Zick have launched Trillium Labs, a nonprofit seeking to open the methods behind high-stakes AI experiments. Its first work centers on post-training and reinforcement learning, including how rewards shape model behavior. The founders plan to share experimental details so others can scrutinize and reproduce the work, with support from Schmidt Sciences and Halcyon Futures.
Artificial Intelligence··Night
Trillium begins with open post-training research
Nathan Lambert and Tom Zick launched Trillium Labs, a nonprofit research organization seeking to open the science behind advanced AI. Its initial work includes post-training, the methods that adapt a model after its basic training, and sharing methods and experimental details for outside scrutiny.[1], [2]
Rewards and user agreement are research subjects
The laboratory’s agenda includes reinforcement learning, where rewards shape behavior, and self-improving AI systems. It also focuses on sycophancy, a model’s tendency to agree excessively with a user, and the model character traits studied during training. The founders intend to publish research details that allow other researchers to examine findings and repeat experiments. They seek access to the training and experimental process beyond testing a finished model through a remote interface.[1]
The founders bring experience from open models and responsible AI
Lambert previously worked on open models at the Allen Institute for AI and at Hugging Face. He also founded American Truly Open Models, an initiative encouraging US organizations to release more open models. Zick’s background includes responsible-AI work at Harvard and Charles Schwab. They met during graduate studies at Berkeley. The laboratory’s supporters include Schmidt Sciences and Halcyon Futures; the total support provided was not disclosed. The organization’s initial activity is research, with substantial computing resources needed for its experiments.[1]