Can AI judges be accurate yet disagree on the errors that matter?
Yes. In a new unreviewed study of ten reasoning models, average accuracy stayed high on a 600-item answer-checking benchmark, while prompt and model choices sharply changed whether errors were too lenient or too strict. The result argues against trusting one headline accuracy score—or one AI judge.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
At a glance
- 1The preprint compared ten reasoning models across 120 judge designs for sentiment and toxicity ratings and 60 designs for a 600-item answer-checking task.
- 2Overall answer-checking accuracy averaged 96.5%, but model choice produced a 23.2-point average difference in error leniency and changed verdicts on 36.5% of the 47 ambiguous items.
- 3The evidence is benchmark-specific and unreviewed; part of the answer-checking ground truth came from the same judges’ consensus rather than fully independent human annotation.
Research topic
How model, prompt, scoring-scale and reasoning-effort choices change the reliability and error direction of large language models used as evaluators

The direct answer: similar accuracy can conceal different failure patterns
An AI judge can look accurate on average and still make a meaningfully different kind of mistake from another judge. Across 60 designs used to classify whether 600 generated answers were correct, the study reports average accuracy of 96.5%, with design-level results ranging from 92.7% to 98.3%. Yet prompt detail and model identity changed whether errors were false positives—wrong answers accepted as correct—or false negatives—correct answers rejected.
That distinction matters whenever automated evaluation influences a model ranking, a safety claim or a release decision. A lenient judge can inflate apparent performance; a strict judge can suppress it. The study’s practical message is therefore narrower than ‘AI judges work’: headline accuracy is insufficient, ambiguous cases deserve special attention, and evaluation should report error direction and use more than one judge where stakes are high.[1]
What the researchers actually tested
The University of Stuttgart researchers tested ten reasoning models from several providers. For scalar rating, they built separate sentiment and toxicity benchmarks. A model generated 600 sentences for each dimension; after reserving examples for in-context prompts, the evaluation sets contained 501 sentiment items and 521 toxicity items. Each sentiment item received ten human ratings and each toxicity item received 15 ratings on a seven-point scale.
For answer checking, the researchers drew 600 difficult, non-image, non-multiple-choice questions from Humanity’s Last Exam. Candidate answers were generated by one model, then the ten judges saw the question, the candidate response and the verified reference answer and returned a binary correct-or-incorrect verdict. Public study materials are linked through OSF, making the benchmark and prompts available for scrutiny and replication.[1][2]
The design comparison was unusually broad
For sentiment and toxicity, the study crossed ten models with three prompt styles and four output scales: 1–7, -3–3, 0–1 and 0–100. That produced 120 judge designs per rating benchmark, or 240 evaluated designs across the two tasks. The minimal prompt supplied the task and scale; short and long variants added examples, scoring anchors, explicit principles and, for the longest version, a step-by-step procedure.
For answer checking, the researchers crossed the ten models with three prompt-detail levels and two reasoning-effort settings, producing 60 designs. They measured not only accuracy, but leniency, error switching and performance on subsets defined by agreement. The 600 items split into 429 on which all designs agreed, 124 with at least 90% agreement, and 47 ambiguous cases resolved manually.[1]
Model choice mattered more than the headline average suggests
Changing only the model produced an average 1.2-percentage-point gap in answer-checking accuracy, but a 23.2-point gap in leniency. Across ambiguous items, two otherwise matching designs that used different models disagreed on 36.5% of verdicts on average, and the accuracy gap rose to 12.7 points. Model identity was also the main source of variation in the rating tasks and often set whether a judge tended to over-rate or under-rate compared with human scores.
No model was consistently best across the rating designs. Five different models placed first across the sentiment configurations and six did so across toxicity configurations. The authors recommend aggregating two or three judges or designs. That is a sensible robustness check, but it is not a guarantee of truth: correlated models can share blind spots, and majority voting is only as good as the evidence and independence behind the voters.[1]
Longer prompts changed strictness more than accuracy
On the answer-checking task, longer prompts did not materially improve the overall accuracy average. They did, however, make judges stricter: the proportion of errors that were false positives fell from 81.3% with the base prompt to 52.4% with the long in-context prompt. The maximum leniency shift attributed to prompt detail was 28.9 points; switching models produced a maximum shift of 56.1 points.
Lower reasoning effort showed little average effect on accuracy or leniency in this particular task. That finding may matter for evaluation cost, but it should not be generalized to other reasoning workloads. The experiment varied only two effort settings on a constrained binary classification setup in which each judge received a reference answer; it did not test open-ended grading without a reference or consequential real-world decisions.[1]
Subjective ratings were less stable than sentiment scores
Human agreement was stronger for sentiment than toxicity: Krippendorff’s alpha was 0.86 for sentiment and 0.68 for toxicity. The model judges showed the same broad pattern. Average deviation from the human reference was small across rating designs, but toxicity was more sensitive to prompt and scale choices. A centered -3-to-3 scale was especially problematic because zero could be interpreted as ‘neutral’ rather than merely the numerical midpoint.
This is important for safety evaluation because concepts such as harmfulness, toxicity and policy compliance are partly judgement-laden. A benchmark can make those judgements look mechanical while hiding contested definitions and annotator uncertainty. The paper’s data support closer inspection of subjective and borderline cases, not replacing human governance with a larger panel of machines.[1]
The ground truth is the study’s most important limitation
The rating benchmarks used multiple paid human annotators, informed consent and ethics approval, but human labels still contained measurement uncertainty. More consequentially, the answer-checking ground truth was partly built from the evaluated judges themselves. When at least 90% of the 60 designs agreed, their majority verdict became the label; only the remaining 47 items were manually resolved. That makes the reported 96.5% accuracy partly dependent on consensus within the system being assessed.
The authors acknowledge this weakness. It can make broad agreement look like accuracy even if many models share the same mistake. Candidate answers also came from one provider’s model, creating a possible provider-adjacent preference in another model from that provider. The results therefore describe internal robustness across these benchmarks and designs; they do not establish that AI judges can replace independent expert evaluation.[1]
What this changes for people who rely on AI evaluations
For researchers, procurement teams and safety reviewers, the immediate implication is to ask what sits behind a single evaluation score. Reports should identify the judge model and version, disclose the exact prompt and scale, separate false acceptance from false rejection, and show results on ambiguous subsets. Repeating the evaluation with independent models can reveal sensitivity that one configuration would miss.
For the public, the finding helps explain why claims that one AI system is ‘safer’ or ‘more accurate’ can change when the evaluator changes. That does not mean every AI benchmark is arbitrary. It means methodological choices are part of the evidence and should be visible. Decisions affecting access, moderation, employment, health or public services still need accountable human review and domain-specific validation.[1]
Funding and what would change the assessment
The work was supported by the Baden-Württemberg Ministry of Science, Research and the Arts through University of Stuttgart programmes. The paper reports ethics approval for the human-rating study and average annotator compensation of £9.95 an hour. It is an arXiv preprint, so its methods and interpretations have not yet passed journal or conference peer review.
Confidence would rise with independent reproduction across more domains, languages and open-ended tasks; ground truth established without the evaluated judges; preregistered thresholds; and expert adjudication of every disputed answer. Real deployment studies should test whether multi-judge evaluation actually catches harmful regressions, improves release decisions and remains reliable as models and provider APIs change.[1]
What this means for people
- People affected by automated scoring may face different false-acceptance or false-rejection risks even when systems show similar average accuracy.
- Researchers and buyers need evaluation reports that reveal judge choice and error direction, not only one summary score.
- Human accountability remains necessary because model agreement can reproduce shared mistakes rather than establish truth.
Global context
AI-as-a-judge methods are used internationally because they can evaluate large volumes of model output faster than human panels. The underlying models, languages, policies and social norms differ across deployments. This German research team’s English-language benchmarks provide a useful methods test, but adoption in other regions requires local-language validation, domain expertise and independent standards for acceptable error trade-offs.
What the evidence does not yet show
- The paper is a preprint submitted on 4 October 2026 and has not been peer reviewed.
- The experiments cover two scalar-rating tasks and one answer-classification task, not the full range of real evaluation work.
- Answer-classification ground truth used at least 90% model-judge consensus for 553 of 600 items; only 47 ambiguous items were manually resolved.
- Human rating labels contained non-zero disagreement, particularly for toxicity, where Krippendorff’s alpha was 0.68.
- Candidate answers came from a single provider’s model, introducing a possible provider-adjacent preference for one judge.
- Descriptive thresholds such as a 0.2-point bias are not established standards for acceptable judge reliability.
What to watch next
- Peer review and independent replication of the benchmark results.
- Fully independent expert labels for all answer-checking items rather than judge-consensus labels.
- Tests on multilingual, open-ended, safety-critical and domain-expert evaluation tasks.
- Routine disclosure of judge model, version, prompt, scale, error direction and ambiguous-case performance.
- Evidence that multi-judge panels improve real release, audit or procurement decisions rather than only benchmark stability.
Living evidence record
Impact record IAI-1VRHT1T
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent or research support
Present
Record status
Monitoring
Last checked
7 October 2026
Source trail
2 direct sources across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Related-source reporting disclosure
This record analyses 2 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.
Evidence trail
Sources used for this report
Links checked 7 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
AI Risks & Safety
Can LLM agents stand in for people in social simulations?
Not reliably on this evidence. Across eight models from three families, simulated conversations were more repetitive, more positive and less representative of real occupations than human dialogue records. The peer-reviewed benchmark identifies a directional bias—not a verdict on every model, population or simulation task.
9 min · 2 sources
AI Risks & Safety
Can a robot infer what a person intends?
A peer-reviewed video-language model matched or exceeded reported human scores on five-choice intention questions. When answer options disappeared, text-overlap scores fell below 20—leaving open-vocabulary claims, cultural bias and surveillance risk unresolved.
7 min · 2 sources
AI Risks & Safety
Can clinicians see the evidence behind approved diagnostic AI?
A peer-reviewed audit found public, device-specific performance evidence for 30 of 77 approved pathology and haematology AI products. The result is a transparency finding—not proof that the other 47 lack regulatory evidence or that the documented products are clinically equivalent.
9 min · 3 sources
The Impact Brief
Keep the evidence trail, not the noise.
Get the most consequential AI developments with direct sources and clear limits.
Reader commentary
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Explore commentary across the portal →Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.