Manuscript and results held privately during review
16
Judges per Bank
5
Model Families
3
Preference Datasets
4
Frontier Providers
The Question
What does agreement actually buy you?
Alignment pipelines increasingly delegate data curation to panels of LLM judges: a preference pair is kept when enough judges agree. Every such rule rests on an assumption about how those judges fail. This project tests that assumption directly, in the filtering setting where it has consequences, rather than inferring it from headline judge accuracy. The work is supervised by Prof. Ser-Nam Lim, who led AI research teams at Meta from 2018 to 2023.
◉
Where it bites
Consensus rules decide which preference pairs survive to train a policy. Whatever they let through becomes the reward signal, so a filtering assumption that quietly fails does not stay contained at the filtering stage.
⇄
Measured, not assumed
Judges are run over fixed item sets so that banks of different sizes, families, and prompt conditions are directly comparable, and every panel-level claim is checked against a matched statistical test rather than a pooled vote count.
≡
On real data
The behaviour is characterised on public preference data as it is, including a human-labelled safety set. No synthetic corruption is required to make the effect appear, and label-noise sensitivity is swept rather than assumed away.
Scale of the Evidence
How far the study reaches
The experimental grid is the part of this work that can be described in full without pre-empting the results. It spans bank size, model family, parameter scale, prompt condition, dataset, judge training regime, and capability tier.
Dimension
Coverage
Judge bank sizes
4 → 16 judges
Open-weight model families
5
Parameter range
1.5B – 14B
Prompt conditions per model
2 styles
Preference datasets
3 public, incl. human safety labels
Judge training regimes
base · DPO · GRPO, matched bank
Frontier capability tier
4 commercial providers
Label-noise robustness
40-seed sweep
Pipeline
32 staged drivers, SLURM cluster runs
The grid is deliberately wider than one bank on one dataset: several of the claims only become checkable when bank size, family, training regime, and capability tier can each be varied while the others are held fixed.
Findings
Held until publication
The manuscript reports three principal findings, a diagnostic that separates the regimes they imply, and a mitigation for each. Those are the contribution, so they stay off this page while the paper is in review. Collaborators and hiring committees are welcome to the full draft on request.
01Withheld
02Withheld
03Withheld
What is not withheld: the manuscript also documents three places where the proposed mitigations do not yet win, including one where no downstream claim is made at all. Those negative results are reported in the paper with the same prominence as the positive ones, and they are the first thing shared with anyone who asks about the work.
Access
Requesting the work
Available on request
The paper is under double-blind review, so the manuscript stays private for now. I am glad to share it directly with people who have a reason to read it.
Full manuscript, including the negative results and limitations
Experimental grid, judge-bank configurations, and evaluation protocol