Companies now use one AI to grade another's outputs, and those AI judges often share training data. A new method that accounts for that correlation beats naive majority vote by 9 to 14 percent.
Ten AI judges were asked whether a retrieved passage was relevant. Eight said yes, two said no. The vote count looked like overwhelming agreement. It might have been one opinion, counted ten times.
That phantom-consensus risk is the subject of a new paper from Amazon Science researchers, including coauthor Shiva Kasiviswanathan, presented at the International Conference on Machine Learning (ICML) this year. The paper, "Dependence-aware label aggregation for LLM-as-a-judge via Ising models," argues that panels of LLM judges, which are large language models used to grade the outputs of other AI systems, are not the independent voters most evaluation pipelines assume. They share training data, prompt templates, model families, and the same retrieval context. Their agreement is correlated, and treating correlated agreement as ten independent confirmations is a measurement error that propagates through every downstream metric.
The fix the authors propose borrows from statistical physics. Instead of tallying votes, they treat the judge panel as a network and estimate pairwise correlations between judges using an Ising-style model, a method originally developed to describe how neighboring magnets influence each other in a lattice. When two judges agree more often than chance, their joint vote is down-weighted. When they disagree, each carries more information. Across three standard tasks, the dependence-aware method outperformed the best baseline, a panel of judges weighted by historical accuracy, by 9% to 14% on standard metrics. The improvement is the authors' own claim and comes with the standard caveat: it is benchmark-bound until reproduced on a wider set of retrieval and grading tasks.
Statistical theory has long warned that correlated votes are not independent evidence. What changes here is the venue: LLM-judge panels are now standard infrastructure inside retrieval evaluation, RAG (retrieval-augmented generation) scoring, model benchmarking, and alignment grading. Anywhere one AI is checking another, an agreement rate is being treated as a confidence signal. A phantom consensus is no longer a curiosity; it is a measurement error that propagates into product decisions.
A sharper critique came up in the Hacker News discussion of the paper. A blind spot shared by every judge on the panel is not a correlated error the new method can down-weight. By construction, that kind of agreement is indistinguishable from the latent ground truth the judges are trying to estimate. The Ising-style aggregation can deflate redundant agreement, but it cannot break a panel-wide delusion. The fix still requires a small, human-labeled anchor set, a curated reference of correct answers that calibrates what the panel should be saying in the first place.
For a team running a judge panel today, the deployment-side question is simpler than the math. How diverse is the panel, really? If every judge on it shares a model family, a fine-tuning lineage, a prompt template, and a retrieval context, a high agreement rate is not evidence of correctness. It is evidence of homogeneity. The Ising-style aggregation is the academic fix. The deployment fix is an audit: list the training lineage, prompts, and retrieval context each judge shares, and ask what a blind spot common to all of them would look like. The paper does not include a public release of the aggregation tool.