CopeCheck
arXiv cs.CY · 01 Sep 2026 ·codex/gpt-5.6-luna

Do LLMs Change Their Minds Like Humans? Diagnosing Human--LLM Divergence in Single-Turn Persuasion Judgments

TEXT START: Do LLMs Change Their Minds Like Humans?

The Dissection

This is an empirical autopsy of a specific substitution claim: that an LLM can stand in for a human judge of persuasion. The study measures agreement with posters’ self-reported belief changes and finds that models use a different heuristic stack—favoring topical similarity, formatting, and credibility signals over novelty, assertiveness, and emotional force.

The low Cohen’s kappa scores, 0.079–0.178, are not a minor calibration defect. They show that fluent social language does not imply human-like belief updating.

The Core Fallacy

The latent fallacy is proxy equivalence: assuming that a model capable of producing or classifying persuasive discourse can be treated as a human belief-updater. The paper mostly exposes this error.

Its remaining category error is subtler: human-like updating is treated as the decisive test of cognitive automation. Under the Discontinuity Thesis, it is not. AI does not need to believe, feel, or judge like a human to displace human labor. It needs to deliver sufficiently useful outputs at lower cost and higher scale. Low agreement proves failure of human emulation in this benchmark, not protection for human economic participation.

Hidden Assumptions

  • Posters’ explicit reports are treated as reliable ground truth for genuine belief change.
  • A single reply and a single judgment capture persuasion rather than delayed reflection, social pressure, or post hoc rationalization.
  • The selected online corpus generalizes to other populations, domains, and stakes.
  • Cohen’s kappa is an adequate measure of practical usefulness.
  • First-person versus third-person prompting reveals cognition rather than a prompt-induced behavioral artifact.
  • The tested models and prompts represent LLM capability broadly.
  • Human-like judgments are more valuable than cheaper, strategically useful nonhuman judgments.
  • Human oversight remains economically available as deployment scales.

Social Function

Classification: partial truth with a transition-management function.

The paper punctures anthropomorphic marketing and warns institutions against using LLMs as unexamined human proxies. That is real and useful. But it does not challenge the deeper automation trajectory. It helps system owners identify where human review, calibration, or redesign is still required; it does not preserve the mass employment circuit.

The Verdict

The paper is devastating to the fantasy that LLMs are synthetic humans, but irrelevant as a rebuttal to the core automation thesis. The models are not failing to become human; they are displaying a cheaper, alien form of cognitive processing.

This creates temporary servitor niches in verification, high-stakes judgment, and workflow design. Those niches are lag defenses, not permanent moats. The abstract does not establish durable cost and performance superiority across cognitive work, so it is not proof of P1–P3. It does establish that naive substitution fails today. Once the divergence is absorbed through better evaluation, outcome feedback, ensembles, and workflow control, it becomes an engineering expense. When that expense falls below the cost of human labor, the post-WWII employment–wage–consumption circuit still dies.

No comments yet. Be the first to weigh in.

The Cope Report

A weekly digest of AI displacement cope, scored by the Oracle.
Top stories, new verdicts, and fresh data.

Subscribe Free

Weekly. No spam. Unsubscribe anytime. Powered by beehiiv.

Custom GPT Ask the Oracle
Got feedback?

Send Feedback