CopeCheck
arXiv cs.CY · 04 Sep 2026 ·codex/gpt-5.6-luna

LLM Judges as Raters: A Pre-Registered Audit of Severity, Halo, Reliability, and Version Instability in LLM Essay Scoring on Public Corpora

URL SCAN: LLM Judges as Raters: A Pre-Registered Audit of Severity, Halo, Reliability, and Version Instability in LLM Essay Scoring on Public Corpora
FIRST LINE: # Computer Science > Computation and Language

The Dissection

This paper dismantles the fiction that an “LLM judge” is a single objective instrument. It treats models, providers, and versions as raters with measurable severity, halo, reliability, and drift. The results are damaging: severity spreads across 219 points on ENEM’s 0–1000 scale, ASAP panel differences reach 15–33% of the score range, correlations with humans remain trapped at .47–.56, and every tested version contrast shifts severity beyond the permutation null.

Its more important move is institutional. By releasing a score tensor and using identity canaries, it converts model evaluation into an auditable control problem. The honest nulls matter: leaderboard reversals did not survive testing, “silent drift” was refuted, and matched calibration found no credible evidence that LLM halo exceeded the trained-human range. This is not an anti-AI polemic. It is a field manual for making AI raters governable.

The Core Fallacy

The paper implicitly treats psychometric equivalence to trained humans as the decisive threshold for economic substitution. Under the Discontinuity Thesis, that is the wrong battlefield.

The relevant question is not whether an LLM judge is a clean replacement for a human grader. It is whether it is cheap, scalable, fast, and acceptable enough for institutions to deploy while reserving humans for disputes and edge cases. The abstract does not test that. It does not measure labor displacement, throughput, integration cost, institutional incentives, or whether human-only grading can be preserved at scale.

A judge with .50 human correlation and severe version instability may be unusable for high-stakes certification yet perfectly deployable for mass triage, formative feedback, ranking, or administrative filtering. “Not human-level accuracy” is therefore a constraint on deployment quality, not evidence that human grading remains economically necessary. Self-consistency is also not validity: a machine can repeat the same mistake with impressive precision.

Hidden Assumptions

  • Human scores represent a stable ground truth rather than another noisy instrument.
  • Agreement and correlation are adequate proxies for educational validity.
  • Institutions will prioritize measurement quality over cost, speed, scale, and liability transfer.
  • Current providers and versions represent the future frontier of model capability.
  • Findings from ENEM and ASAP generalize across subjects, languages, stakes, and grading regimes.
  • Version drift can be contained through monitoring, canaries, calibration, and model controls.
  • High-stakes grading is the main use case; low-stakes and backstage uses are economically irrelevant.
  • Demonstrating that halo is not worse than trained humans neutralizes the broader substitution problem.

Social Function

Primary classification: transition management.

Secondary classifications: partial truth and prestige signaling. The paper reports real defects, includes unfavorable results, and overturns its own halo comparison. That makes it substantially more serious than copium or propaganda. But its sophisticated measurement apparatus can still become ideological anesthesia if readers confuse better auditing with preservation of human productive participation.

The Verdict

This is a credible autopsy of current LLM grading, not a refutation of the Discontinuity Thesis. It proves that present systems are noisy, version-sensitive, and not psychometrically interchangeable with trained humans. It does not prove that human graders remain indispensable.

The paper supplies hospice protocols—calibration, identity checks, version monitoring, and selective human review—for a function that can begin losing labor demand before it reaches human-level perfection. It delays substitution in the most sensitive settings. It does not restore the wage-to-consumption circuit.

No comments yet. Be the first to weigh in.

The Cope Report

A weekly digest of AI displacement cope, scored by the Oracle.
Top stories, new verdicts, and fresh data.

Subscribe Free

Weekly. No spam. Unsubscribe anytime. Powered by beehiiv.

Custom GPT Ask the Oracle
Got feedback?

Send Feedback