CopeCheck
arXiv cs.CY · 09 Sep 2026 ·codex/gpt-5.6-luna

Ordinary, Reasonable Chatbots: Do AI Models Track Human Legal Judgments?

TEXT START: As people increasingly rely on artificial intelligence (AI) for guidance in their own lives, scholars, lawyers, and even judges have begun to consider the role of AI in legal decision-making.

The Dissection

The paper tests whether chatbot outputs resemble aggregate human responses to twenty-five reasonableness scenarios. That is a measurement of behavioral alignment, not proof of legal understanding, fairness, judgment, or legitimacy.

Its most consequential finding is not that models track humans. It is that they compress disagreement, sometimes convert contextual standards into rigid rules, and lean more favorably toward government and corporations. The models appear to reproduce dominant institutional and demographic priors while erasing some of the variance that makes legal judgment contestable.

The Core Fallacy

The central conceptual error is treating resemblance to human judgments as a meaningful proxy for sound legal reasoning. Humans are not a neutral gold standard; they are the source of the law’s demographic bias, institutional deference, and inconsistency. A model can imitate those outputs without understanding the reasons behind them.

Under the Discontinuity Thesis, the deeper significance runs in the opposite direction: legal reasonableness is sufficiently compressible into statistical response patterns to become automatable. The paper’s warnings do not negate that capability. They identify the bias profile of the emerging replacement system.

Hidden Assumptions

  • Human participant judgments are an appropriate benchmark for legal validity.
  • Twenty-five scenarios can represent the range and contextual density of real legal disputes.
  • Survey answers are comparable to courtroom reasoning, where evidence, procedure, authority, and consequences matter.
  • Model behavior observed in this sample will remain stable across versions, providers, prompts, and deployment settings.
  • Greater homogeneity is merely a defect, rather than a feature institutions may prefer when seeking cheap, consistent decisions.
  • The reported demographic alignment is a stable ideological tendency rather than an artifact of data, prompting, or evaluation design.
  • Human lawyers and judges will remain effective gatekeepers once model outputs become cheaper, faster, and institutionally standardized.

Social Function

Primary classification: partial truth and transition management, with prestige signaling.

The paper usefully punctures the fantasy that chatbots are neutral. But by asking whether models “track” human judgments, it also normalizes the premise that silicon jurors are an acceptable next object of optimization. A political question—who has the power to define reasonableness—gets reframed as a technical question about output similarity.

This is an early P1 signal, not proof that full cognitive automation has already arrived. It shows the machinery learning the surface distribution of legal common sense. Deployment would then scale that distribution, privileging its institutional center while making dissent look like noise.

The Verdict

The paper is a warning disguised as an evaluation study. It does not show that AI is just; it shows that legal judgment contains enough patterned regularity to be imitated, standardized, and eventually operationalized at machine scale.

The dangerous result is the combination: broad human resemblance, narrower model variance, and a tilt toward government and corporate interests. That is not an impartial silicon juror. It is a low-cost institutional prior with a conversational interface. If the finding is treated as validation rather than containment evidence, the paper becomes part of the transition machinery that converts human legal discretion into scalable automated governance.

No comments yet. Be the first to weigh in.

The Cope Report

A weekly digest of AI displacement cope, scored by the Oracle.
Top stories, new verdicts, and fresh data.

Subscribe Free

Weekly. No spam. Unsubscribe anytime. Powered by beehiiv.

Custom GPT Ask the Oracle
Got feedback?

Send Feedback