Preprint · 2026

Automatic-Jury-Consensus

Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus

Authors Elias Hossain, Niloofar Yousefi, Ser-Nam Lim University of Central Florida

Code Contact
0.21
Mean error correlation
3.5
Effective judges of 10
28%
Significance flips
0.56
Frontier correlation

The Question

How much evidence does agreement actually provide?

Alignment pipelines increasingly use panels of LLM judges to decide which of two responses is preferred, and keep an example only when enough judges agree. The intuition is that agreement among many judges is stronger evidence than any single judge. That intuition assumes the judges make their errors independently. LLM judges share pretraining data, post-training recipes, evaluation templates, and blind spots, so they often make the same mistakes. This project measures how much that dependence weakens consensus.

i

How much independent information?

We estimate the pairwise correlation of judge errors and convert it into an effective number of independent judges, so that a bank's evidence is reported in the same units as its size.

ii

Can shared errors change conclusions?

We compare a pooled-vote test, which treats every judge vote as independent, with an item-level test that uses each preference pair as one observation, and count how often they disagree.

iii

Does the pattern of errors matter?

We separate errors shared by most judges from errors concentrated in a smaller group, and test which way of combining judge decisions works under each pattern.


Findings

Agreement overstates evidence

The number of judge votes should not be read as the amount of independent statistical evidence. Correlated errors shrink a bank's evidence well below its nominal size, can change statistical verdicts, and are stronger, not weaker, among the most accurate judges.

10 → 3.5
judges → effective judges

Shared errors shrink the bank

In the main ten-judge bank, the mean pairwise error correlation is 0.21. This inflates the variance of the mean error rate by a factor of 2.85, so the ten judges carry about as much information as 3.5 independent ones. The same holds on two more datasets and 73 alternative bank compositions.

up to 28%
of comparisons flip

Conclusions change

On resampled 100-pair panels, a pooled-vote test finds a significant system difference in 79% of panels, while an item-level test does so in 52%. In 28% of panels the verdict depends only on whether shared errors are ignored.

0.56
frontier error correlation

Accuracy is not independence

Frontier judges from three different providers reach 91–93% accuracy, yet make the same error together at 7.7× the rate expected under independence. Cross-provider judges are about as correlated as judges from one provider. Adding open-weight judges to the frontier bank raises the effective number of judges from 1.88 to 3.83.

2 patterns
different best filters

The pattern of errors matters

When many judges fail together, measured correlation rises. When failures are concentrated in a vulnerable subgroup, correlation can fall, from 0.22 to 0.07 in our experiments, even as majority decisions get worse. The two patterns favor different filters, so correlation magnitude alone is not enough.

Recommendation. Report error dependence alongside judge accuracy, treat the item rather than the vote as the unit of inference, and use a small trusted set of about 100 labelled examples to estimate judge accuracy, identify shared mistakes, and choose the voting rule before applying it to new data.

Setup

Judges, data, and protocol

Open-weight and frontier judges are evaluated on the same preference pairs, so that dependence can be compared across banks, datasets, and capability tiers on identical items.

ComponentCoverage
Main open-weight bank10 judges: Llama-3.1-8B, Qwen-2.5-7B, Gemma-2-9B, Mistral-7B-v0.3, Phi-3.5-mini × 2 prompt styles
Frontier judges6: three Gemini models, plus GPT, Claude, and Grok
Preference datasetsRewardBench v2 · UltraFeedback · PKU-SafeRLHF
Calibration set1,195 stratified RewardBench pairs (1,133 answered by all ten judges)
Robustness73 alternative bank compositions; label-noise sweep of 1–10%
Training-regime controlDPO- and GRPO-trained open-weight judges
Transfer screen8 LLM-AggreFact components, plus factuality and code judging
Filters comparedmajority · 0.75 supermajority · Automatic-Jury-Consensus · Bias-Cluster · learned selector

Filters are compared on held-out splits at matched retention, and no threshold or other parameter is tuned on the test split.


Scope

What the paper does not claim

The contribution is a measurement and validation protocol, not a new aggregation method that replaces majority voting. The negative results are reported with the same prominence as the positive ones.

×

No universal filter

No single fixed rule works best under both error patterns. None of the 11 screened factuality and code settings met the criteria for a filter to help, and on all eight LLM-AggreFact components neither Automatic-Jury-Consensus nor Bias-Cluster beats plain majority.

≈

A bounded safeguard

The learned selector helps mainly when the best rule shifts between datasets. It guards against transferring an unsuitable filter, and its gains over a fixed rule are small.

!

Controlled regimes

The global and subgroup patterns are created by controlled interventions and do not establish how often each occurs naturally, although subgroup failures do appear in the unmodified data.


Code

Use the work

Code and paper

The repository contains the code for the experiments in the paper. The method-level recipe is simple enough to apply to any judge bank:

github.com/eliashossain001/Automatic-Jury-Consensus
@article{hossain2026agreement, title = {Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus}, author = {Hossain, Elias and Yousefi, Niloofar and Lim, Ser-Nam}, year = {2026} }