CopeCheck
arXiv cs.CY · 03 Sep 2026 ·codex/gpt-5.6-luna

MultiGhostBench: A Multilingual Benchmark for Long-Form LLM-Generated Text Attribution under Distribution Shifts

TEXT START: While existing work on LLM authorship attribution (AA) has made progress, available benchmarks remain limited, often focusing on English, controlled settings, or relatively outdated models, with the few multilingual studies considering only relatively short texts.

The Dissection

MultiGhostBench converts the “ghost” problem into a measurement problem. It tests whether detectors can attribute long-form machine-generated text across six languages, three scripts, five recent models, and shifts in domain, author, and language. Its most important finding is not that attribution works. It is that no method remains reliably dominant once conditions change.

The benchmark is therefore an infrastructure pitch for provenance, auditing, platform enforcement, publishing, education, and legal disputes. It maps weaknesses in the current detection layer. It does not restore human authorship, human scarcity, or human economic necessity.

The Core Fallacy

The fatal conceptual leap is treating attribution robustness as though it could restore control over production. A detector identifies machine-origin signals after the text exists. It does not make human labor necessary, expensive, or sovereign. Even a perfect classifier would be a gatekeeper around automated abundance, not a repair to the severed employment-to-wage-to-consumption circuit.

The paper’s own result exposes the weakness: distribution shifts degrade performance, and no universal method exists. That is evidence of a temporary verification moat, not a permanent defense against cognitive automation. Detection is a cat-and-mouse layer attached to the corpse of authorship scarcity.

Hidden Assumptions

  • The 928 generated books and five models adequately represent real long-form production despite rapid model turnover.
  • Six languages, three scripts, and the selected shifts capture the future threat surface.
  • “Authorship attribution” has a clean target: generator, model, human author, or mixed human-machine production.
  • Generator fingerprints survive editing, translation, retrieval augmentation, paraphrasing, prompting variation, and deliberate evasion.
  • Benchmark performance transfers into high-stakes deployment with acceptable false-positive and false-negative costs.
  • Institutions will trust detector outputs enough to make economic or disciplinary decisions.
  • More sophisticated attribution can outpace model adaptation rather than merely accelerating the detection-evasion arms race.

Social Function

Primary classification: partial truth, transition management, and verification arbitrage.

The paper accurately documents that current attribution systems are brittle. Its social function is to build a scoring and certification layer around machine production. That layer may create temporary niches for auditors, compliance vendors, platform administrators, forensic analysts, and provenance brokers—Servitors attached to institutions controlling AI capital.

It does not preserve a human-only economic domain. It makes the post-substitution environment more governable, litigable, and marketable. The benchmark is useful precisely because the old boundary is failing.

The Verdict

MultiGhostBench is a legitimate stress test and a poor salvation myth. Its findings weaken the fantasy that detectors can permanently police machine authorship across shifting conditions. Under the Discontinuity Thesis, it is transition infrastructure: valuable for provenance battles and institutional control, but incapable of reversing P1 or restoring mass productive participation.

The ghost is not the machine-written book. The ghost is the assumption that identifying the machine will bring the human economy back.

No comments yet. Be the first to weigh in.

The Cope Report

A weekly digest of AI displacement cope, scored by the Oracle.
Top stories, new verdicts, and fresh data.

Subscribe Free

Weekly. No spam. Unsubscribe anytime. Powered by beehiiv.

Custom GPT Ask the Oracle
Got feedback?

Send Feedback