Amazon researchers account for correlated errors among judges that share prompts, training lineage or model families, beating weighted majority voting across three tasks. Goodfire's newly public Silico takes a question in natural language, plans interpretability experiments, and runs them in parallel across a model's weights, activations and attention patterns, offering another route past black-box evaluation.
Artificial Intelligence··Night
Consensus may not be independent evidence
Amazon Science researchers argue that matching verdicts from several models may contain less independent evidence than the vote count suggests. Judges that share prompt examples, training lineage or a model family can fail together at the same blind spot. Their aggregation method therefore goes beyond weighting each judge by individual reliability: it also learns pairwise dependence through an Ising model. Human reference labels are not used for training and enter only later to measure experimental performance.[1]
The relationships between votes are measured
The method was tested on relevance classification, toxicity detection and summarisation assessment. Dependence-aware variants beat majority voting weighted by historical accuracy across all three tasks. Patterns of agreement and disagreement in evaluation logs can reveal which judges merely duplicate one another and where a task creates a shared blind spot. Teams can use that distinction to test whether another judge adds genuinely different evidence or merely amplifies bias already present in the panel.[1]
Silico looks inside the weights
Goodfire's newly public Silico moves evaluation from model outputs into the model itself. A user asks in natural language when and why a model hallucinates; the platform plans experiments with sparse autoencoders and probes, then runs them in parallel across weights, activations and attention patterns. IEEE Spectrum reports that Prima Mente used these tools to find its Alzheimer's detection model was relying on DNA fragment-length patterns the team had not known about. For now, that finding remains a company-described use case.[2]