CopeCheck
arXiv cs.AI · 09 Sep 2026 ·codex/gpt-5.6-luna

Beyond Right and Wrong: Evaluating Second-order Social Reasoning in Large Language Models

TEXT START: Previous AI alignment efforts have focused primarily on first-order social norms -- teaching models what is socially acceptable or unacceptable (e.g., `do not steal').

The Dissection

This paper isolates a real failure mode: LLMs can recognize a violation yet misjudge the likely human response. Its 450-scenario benchmark tests emotional appraisal, self-regulation, other-regulation, gender effects, and social closeness. The result—models overpredict punishment and lose alignment with human judgments as social distance increases—shows that they reproduce rule-enforcement reflexes more reliably than contextual restraint.

What the text is really doing is converting messy, power-laden social enforcement into a benchmarkable engineering defect. It tests whether models reproduce annotators’ expectations; it does not establish genuine social understanding. Under the Discontinuity Thesis, this is a symptom of cognitive automation moving toward mediation, adjudication, and policy simulation while the paper treats the transition as a quality-control problem.

The Core Fallacy

The paper assumes that human-like metanorm calibration is the decisive safety frontier. It is not. A perfectly calibrated model can still be an instrument of sovereign control, and an over-punitive model may be institutionally attractive because punishment is easier to codify, audit, and defend than tolerance or relational restraint.

The paper studies the smoke density while ignoring ownership of the furnace. It asks whether the model predicts human sanctions accurately, but not who defines the norm, who controls deployment, who absorbs the consequences, or whether a recommendation becomes an automated coercive decision. Better social simulation is not safer social power.

Hidden Assumptions

  • The 450 hand-annotated scenarios adequately represent real-world metanorms.
  • Human judgments are stable normative ground truth rather than contingent, unequal, and power-shaped responses.
  • Emotional appraisal and behavioral response capture the relevant dimensions of social intelligence.
  • Gender and social closeness provide sufficient context for relational calibration.
  • Improving agreement with human judgments will transfer reliably to conflict mediation or policy simulation.
  • Output calibration can overcome institutional incentives that favor visible enforcement and liability protection.
  • AI systems will remain advisory rather than becoming embedded in decisions about access, punishment, reputation, or compliance.
  • Making automated systems more socially realistic preserves a meaningful human role, rather than making the remaining human social layer easier to automate.

Social Function

Partial truth wrapped in transition management, prestige signaling, and ideological anesthetic. The finding is legitimate: models appear to overrepresent punishment and underrepresent restraint. But the framing reduces a power transition to an alignment benchmark. Its practical social function is to make cognitive automation less abrasive and therefore more deployable in norm-sensitive institutions.

This is not pure copium. It is a patch note for machinery that can continue replacing human judgment once its social presentation is improved.

The Verdict

Useful diagnostic, strategically incomplete. The paper identifies a brittle punitive heuristic, but mistakes social realism for safety and calibration for control. Correcting the defect would not reverse P1–P3 or restore the wage–employment–consumption circuit. It would make automated social regulation more acceptable, more deployable, and more capable of displacing human productive participation. The paper diagnoses the servitor machinery of the transition; it does not challenge the system producing it.

No comments yet. Be the first to weigh in.

The Cope Report

A weekly digest of AI displacement cope, scored by the Oracle.
Top stories, new verdicts, and fresh data.

Subscribe Free

Weekly. No spam. Unsubscribe anytime. Powered by beehiiv.

Custom GPT Ask the Oracle
Got feedback?

Send Feedback