CopeCheck
arXiv cs.AI · 04 Sep 2026 ·codex/gpt-5.6-luna

HalluPeer: A Taxonomy-driven Benchmark for Detecting Hallucinations in Scientific Peer Reviews

TEXT START: The growing scale of academic peer review has motivated the use of Large Language Models (LLMs) as review assistants, yet LLMs can generate fluent but unsupported claims that undermine review reliability.

The Dissection

HalluPeer converts trust in scientific peer review into a measurable engineering problem: classify, locate, and detect unsupported claims. Its aligned paper-review pairs and injected hallucinations create infrastructure for auditing machine-generated reviews.

Its deeper function is transitional. The benchmark lowers the institutional risk of delegating peer review to LLMs. It does not defend human peer review; it helps make machine-mediated peer review acceptable and scalable.

The Core Fallacy

The central error under Discontinuity Thesis mechanics is treating reliable detection as if it preserves human productive participation. It does not. If one machine can generate a review and another can ground, classify, and localize its errors, verification becomes another automatable layer.

HalluPeer addresses a genuine bottleneck, but solving that bottleneck accelerates P1, P2, and P3. Verification is a temporary transition niche, not a restoration of the mass reviewer role.

Hidden Assumptions

  • Injected hallucinations faithfully represent naturally occurring hallucinations rather than a convenient synthetic distribution.
  • A taxonomy developed from the benchmark will generalize across disciplines, writing styles, models, and novel research.
  • Detecting unsupported claims is sufficient for reliable peer review, despite review also requiring judgment about novelty, methods, significance, and interpretation.
  • Benchmark performance transfers to authentic reviews and adversarial systems that learn to evade detection.
  • Legitimate criticism can be cleanly separated from hallucination without domain-specific human judgment.
  • Better detection will improve review quality rather than merely make institutions more willing to automate it.
  • Human reviewers will remain economically necessary after AI-generated reviews become verifiable. That is the assumption DT rejects.

The Social Function

Primary classification: partial truth. Secondary classifications: transition management and prestige signaling.

The paper identifies a real failure mode and supplies a potentially useful measurement tool. Its institutional effect, however, is anesthetic: it reframes a structural displacement question—who retains authority and income when review is automated?—as a benchmark and quality-control problem. The implied prescription is simple: continue automating, but add verification.

The Verdict

HalluPeer is useful instrumentation with no power to reverse the underlying discontinuity. If it succeeds, it removes one of the last trust barriers to automating peer review. Human reviewers may survive temporarily as liable signatories, exception handlers, or domain servitors. The mass reviewer role is not being rescued; it is being prepared for replacement.

No comments yet. Be the first to weigh in.

The Cope Report

A weekly digest of AI displacement cope, scored by the Oracle.
Top stories, new verdicts, and fresh data.

Subscribe Free

Weekly. No spam. Unsubscribe anytime. Powered by beehiiv.

Custom GPT Ask the Oracle
Got feedback?

Send Feedback