How much evidence does agreement actually provide?
Alignment pipelines increasingly use panels of LLM judges to decide which of two responses is preferred, and keep an example only when enough judges agree. The intuition is that agreement among many judges is stronger evidence than any single judge. That intuition assumes the judges make their errors independently. LLM judges share pretraining data, post-training recipes, evaluation templates, and blind spots, so they often make the same mistakes. This project measures how much that dependence weakens consensus.
i
How much independent information?
We estimate the pairwise correlation of judge errors and convert it into an effective number of independent judges, so that a bank's evidence is reported in the same units as its size.
ii
Can shared errors change conclusions?
We compare a pooled-vote test, which treats every judge vote as independent, with an item-level test that uses each preference pair as one observation, and count how often they disagree.
iii
Does the pattern of errors matter?
We separate errors shared by most judges from errors concentrated in a smaller group, and test which way of combining judge decisions works under each pattern.
Findings
Agreement overstates evidence
The number of judge votes should not be read as the amount of independent statistical evidence. Correlated errors shrink a bank's evidence well below its nominal size, can change statistical verdicts, and are stronger, not weaker, among the most accurate judges.
10 → 3.5
judges → effective judges
Shared errors shrink the bank
In the main ten-judge bank, the mean pairwise error correlation is 0.21. This inflates the variance of the mean error rate by a factor of 2.85, so the ten judges carry about as much information as 3.5 independent ones. The same holds on two more datasets and 73 alternative bank compositions.
up to 28%
of comparisons flip
Conclusions change
On resampled 100-pair panels, a pooled-vote test finds a significant system difference in 79% of panels, while an item-level test does so in 52%. In 28% of panels the verdict depends only on whether shared errors are ignored.
0.56
frontier error correlation
Accuracy is not independence
Frontier judges from three different providers reach 91–93% accuracy, yet make the same error together at 7.7× the rate expected under independence. Cross-provider judges are about as correlated as judges from one provider. Adding open-weight judges to the frontier bank raises the effective number of judges from 1.88 to 3.83.
2 patterns
different best filters
The pattern of errors matters
When many judges fail together, measured correlation rises. When failures are concentrated in a vulnerable subgroup, correlation can fall, from 0.22 to 0.07 in our experiments, even as majority decisions get worse. The two patterns favor different filters, so correlation magnitude alone is not enough.
Recommendation. Report error dependence alongside judge accuracy, treat the item rather than the vote as the unit of inference, and use a small trusted set of about 100 labelled examples to estimate judge accuracy, identify shared mistakes, and choose the voting rule before applying it to new data.
Setup
Judges, data, and protocol
Open-weight and frontier judges are evaluated on the same preference pairs, so that dependence can be compared across banks, datasets, and capability tiers on identical items.
Filters are compared on held-out splits at matched retention, and no threshold or other parameter is tuned on the test split.
Scope
What the paper does not claim
The contribution is a measurement and validation protocol, not a new aggregation method that replaces majority voting. The negative results are reported with the same prominence as the positive ones.
×
No universal filter
No single fixed rule works best under both error patterns. None of the 11 screened factuality and code settings met the criteria for a filter to help, and on all eight LLM-AggreFact components neither Automatic-Jury-Consensus nor Bias-Cluster beats plain majority.
≈
A bounded safeguard
The learned selector helps mainly when the best rule shifts between datasets. It guards against transferring an unsuitable filter, and its gains over a fixed rule are small.
!
Controlled regimes
The global and subgroup patterns are created by controlled interventions and do not establish how often each occurs naturally, although subgroup failures do appear in the unmodified data.
Code
Use the work
Code and paper
The repository contains the code for the experiments in the paper. The method-level recipe is simple enough to apply to any judge bank:
Estimate error correlation and the effective number of judges for your own bank
Validate the voting rule on a small trusted set before applying it to new data
@article{hossain2026agreement,
title = {Agreement Overstates Evidence: Error Dependence
in LLM Judge Consensus},
author = {Hossain, Elias and Yousefi, Niloofar and Lim, Ser-Nam},
year = {2026}
}