Under double-blind review

CorrFilter

A large-scale audit of the reliability assumptions behind LLM-judge consensus in alignment-data filtering.

Supervised by Prof. Ser-Nam Lim University of Central Florida · formerly led AI research at Meta

Request the manuscript Other work

Manuscript and results held privately during review

16
Judges per Bank
5
Model Families
3
Preference Datasets
4
Frontier Providers

The Question

What does agreement actually buy you?

Alignment pipelines increasingly delegate data curation to panels of LLM judges: a preference pair is kept when enough judges agree. Every such rule rests on an assumption about how those judges fail. This project tests that assumption directly, in the filtering setting where it has consequences, rather than inferring it from headline judge accuracy. The work is supervised by Prof. Ser-Nam Lim, who led AI research teams at Meta from 2018 to 2023.

Where it bites

Consensus rules decide which preference pairs survive to train a policy. Whatever they let through becomes the reward signal, so a filtering assumption that quietly fails does not stay contained at the filtering stage.

Measured, not assumed

Judges are run over fixed item sets so that banks of different sizes, families, and prompt conditions are directly comparable, and every panel-level claim is checked against a matched statistical test rather than a pooled vote count.

On real data

The behaviour is characterised on public preference data as it is, including a human-labelled safety set. No synthetic corruption is required to make the effect appear, and label-noise sensitivity is swept rather than assumed away.


Scale of the Evidence

How far the study reaches

The experimental grid is the part of this work that can be described in full without pre-empting the results. It spans bank size, model family, parameter scale, prompt condition, dataset, judge training regime, and capability tier.

DimensionCoverage
Judge bank sizes4 → 16 judges
Open-weight model families5
Parameter range1.5B – 14B
Prompt conditions per model2 styles
Preference datasets3 public, incl. human safety labels
Judge training regimesbase · DPO · GRPO, matched bank
Frontier capability tier4 commercial providers
Label-noise robustness40-seed sweep
Pipeline32 staged drivers, SLURM cluster runs

The grid is deliberately wider than one bank on one dataset: several of the claims only become checkable when bank size, family, training regime, and capability tier can each be varied while the others are held fixed.


Findings

Held until publication

The manuscript reports three principal findings, a diagnostic that separates the regimes they imply, and a mitigation for each. Those are the contribution, so they stay off this page while the paper is in review. Collaborators and hiring committees are welcome to the full draft on request.

01Withheld
02Withheld
03Withheld
What is not withheld: the manuscript also documents three places where the proposed mitigations do not yet win, including one where no downstream claim is made at all. Those negative results are reported in the paper with the same prominence as the positive ones, and they are the first thing shared with anyone who asks about the work.

Access

Requesting the work

Available on request

The paper is under double-blind review, so the manuscript stays private for now. I am glad to share it directly with people who have a reason to read it.

Email me about CorrFilter