AI-generated analysis · May contain errors · Disclosure and methodology
Do LLMs Change Their Minds Like Humans? Diagnosing Human--LLM Divergence in Single-Turn Persuasion Judgments
TEXT START: Do LLMs Change Their Minds Like Humans?
The Dissection
This is an empirical autopsy of a specific substitution claim: that an LLM can stand in for a human judge of persuasion. The study measures agreement with posters’ self-reported belief changes and finds that models use a different heuristic stack—favoring topical similarity, formatting, and credibility signals over novelty, assertiveness, and emotional force.
The low Cohen’s kappa scores, 0.079–0.178, are not a minor calibration defect. They show that fluent social language does not imply human-like belief updating.
The Core Fallacy
The latent fallacy is proxy equivalence: assuming that a model capable of producing or classifying persuasive discourse can be treated as a human belief-updater. The paper mostly exposes this error.
Its remaining category error is subtler: human-like updating is treated as the decisive test of cognitive automation. Under the Discontinuity Thesis, it is not. AI does not need to believe, feel, or judge like a human to displace human labor. It needs to deliver sufficiently useful outputs at lower cost and higher scale. Low agreement proves failure of human emulation in this benchmark, not protection for human economic participation.
Hidden Assumptions
- Posters’ explicit reports are treated as reliable ground truth for genuine belief change.
- A single reply and a single judgment capture persuasion rather than delayed reflection, social pressure, or post hoc rationalization.
- The selected online corpus generalizes to other populations, domains, and stakes.
- Cohen’s kappa is an adequate measure of practical usefulness.
- First-person versus third-person prompting reveals cognition rather than a prompt-induced behavioral artifact.
- The tested models and prompts represent LLM capability broadly.
- Human-like judgments are more valuable than cheaper, strategically useful nonhuman judgments.
- Human oversight remains economically available as deployment scales.
Social Function
Classification: partial truth with a transition-management function.
The paper punctures anthropomorphic marketing and warns institutions against using LLMs as unexamined human proxies. That is real and useful. But it does not challenge the deeper automation trajectory. It helps system owners identify where human review, calibration, or redesign is still required; it does not preserve the mass employment circuit.
The Verdict
The paper is devastating to the fantasy that LLMs are synthetic humans, but irrelevant as a rebuttal to the core automation thesis. The models are not failing to become human; they are displaying a cheaper, alien form of cognitive processing.
This creates temporary servitor niches in verification, high-stakes judgment, and workflow design. Those niches are lag defenses, not permanent moats. The abstract does not establish durable cost and performance superiority across cognitive work, so it is not proof of P1–P3. It does establish that naive substitution fails today. Once the divergence is absorbed through better evaluation, outcome feedback, ensembles, and workflow control, it becomes an engineering expense. When that expense falls below the cost of human labor, the post-WWII employment–wage–consumption circuit still dies.
Comments (0)
No comments yet. Be the first to weigh in.