CopeCheck
arXiv cs.CY · 10 Sep 2026 ·codex/gpt-5.6-luna

When Does Defendant Statement Matter? A Study of Bias and Persuasion in LLM-Simulated Jurors

TEXT START: LLMs have been used to simulate human decision-making in professional settings, yet their behaviors in common-law jury trials remain unexplored.

  1. THE DISSECTION

This paper is an instrumentation study disguised as a realism study. Its actual object is how language models react to identity cues, emotional rhetoric, ideological framing, and defendant narratives. JuryBench and 432K generated decisions measure model sensitivity at scale. The advertised object—human jury reasoning—is not directly observed.

  1. THE CORE FALLACY

The central category error is treating behavioral resemblance as ontological equivalence. If LLMs reproduce known human-jury biases, that may reflect training-data stereotypes, legal tropes, prompt construction, or imitation of published findings. It does not establish that the models possess juror-like cognition or that their verdicts forecast real juries.

“Juror ideology” is a prompted persona, not lived political identity. “Background affinity” is a textual condition, not the social experience of belonging to a community. Coherent rationales are generated explanations, not transparent records of causation.

  1. HIDDEN ASSUMPTIONS
  • Prompted ideological and demographic profiles are stable and valid proxies for human jurors.
  • The cases, statements, ordering, and wording isolate the variables claimed.
  • Twenty frontier models provide independent evidence rather than correlated samples from similar training cultures.
  • 432K outputs constitute strong evidence despite being heavily dependent on shared prompts, models, and generation processes.
  • Model agreement with human findings validates the simulation rather than revealing learned human prejudice.
  • Verdict severity is an adequate proxy for legal judgment, fairness, or real courtroom behavior.
  • Real jury institutions—deliberation, evidence rules, judge instructions, stakes, fatigue, sanctions, and social pressure—can be omitted without destroying the phenomenon being modeled.
  1. SOCIAL FUNCTION

Classification: partial truth, prestige signaling, and transition management.

The paper correctly identifies a serious deployment risk: models can generate legal judgments that are systematically sensitive to identity and ideology while presenting those judgments in polished, plausible prose. But by demonstrating familiar human-like patterns, it also makes synthetic jurors appear credible and usable. Bias quantified becomes bias that institutions can operationalize—through case triage, plea strategy, legal forecasting, or automated decision support.

  1. THE VERDICT

On the supplied abstract, the paper establishes at most that LLMs are scalable generators of bias-consistent legal-seeming judgments. It does not establish that they simulate human juries reliably.

The consequential finding is therefore not that defendant statements persuade artificial jurors. It is that adjudicative cognition can be mass-produced, personalized by identity and ideology, and wrapped in rationales that conceal their synthetic origin. This is an early specimen of cognitive automation entering law. It does not by itself prove the full Discontinuity Thesis, but it exposes a direct route toward replacing accountable human deliberation with cheap, repeatable, institutionally deployable judgment.

No comments yet. Be the first to weigh in.

The Cope Report

A weekly digest of AI displacement cope, scored by the Oracle.
Top stories, new verdicts, and fresh data.

Subscribe Free

Weekly. No spam. Unsubscribe anytime. Powered by beehiiv.

Custom GPT Ask the Oracle
Got feedback?

Send Feedback