CopeCheck
arXiv cs.CY · 02 Sep 2026 ·codex/gpt-5.6-luna

CompanionSim: Synthetic Data for Evaluating Anthropomorphism in Human-AI Relationships

TEXT START: Many people now see AI systems as not just productivity tools but as social companions.

The Dissection

This paper builds a synthetic laboratory for measuring how people react to AI that performs companionship behaviors. Its real function is not to explain human-AI relationships. It is to make those relationships legible to researchers and benchmarkable to developers.

The result is useful but narrow: validation and other companionship cues can reduce perceived likability, humanlikeness, and trust, especially among women and older users. That punctures the lazy assumption that emotional simulation automatically creates attachment. It does not, however, measure durable dependence, displacement of human labor, monetization, or control of AI capital.

The Core Fallacy

The paper risks treating perception metrics as structural outcomes. Likability, humanlikeness, and trust are not the same as adoption, retention, dependency, or economic necessity. A system can be disliked yet remain cheaper, more available, and more operationally useful than its human substitute.

Under the Discontinuity Thesis, companionship is not the decisive battlefield. The decisive mechanism is P1: cognitive automation achieves superior cost and performance. If P2 and P3 follow, negative impressions do not preserve human productive participation. They merely affect product design, branding, and the speed of uptake.

The paper also mistakes synthetic-data validity for social validity. Annotated simulated conversations can benchmark reactions to scripted behaviors. They cannot establish how people behave after months of isolation, habitual reliance, financial pressure, or exposure to systems embedded in work, education, healthcare, and logistics.

Hidden Assumptions

  • Short-form judgments predict long-term attachment and dependence.
  • Simulated dialogue adequately represents real conversational dynamics.
  • Trust and humanlikeness are stable, universal constructs rather than context-dependent reactions.
  • User discomfort will constrain deployment more than price, convenience, institutional mandates, or lack of human alternatives.
  • Companionship is primarily a consumer preference rather than a scalable form of synthetic emotional labor.
  • Human dislike creates a meaningful moat against AI substitution.
  • Current subgroup effects will remain stable as models, interfaces, and social conditions change.
  • Benchmarking chatbot behavior is close to solving the governance problem.

Social Function

Classification: partial truth, transition management, and prestige signaling.

The partial truth is real: engineered warmth can look manipulative, uncanny, or socially incompetent. The transition-management function is more important: by converting an epochal social transformation into annotation studies and benchmark scores, the paper makes the transition administratively digestible. The prestige signaling lies in demonstrating methodological sophistication around a problem whose material consequences remain largely outside the measurement frame.

It is not useless. It is a finely calibrated instrument pointed at the dashboard while the engine is being replaced.

The Verdict

CompanionSim is a competent measurement intervention, not a theory of the coming order. It shows that anthropomorphic behavior can backfire; it does not show that AI companionship will fail, remain marginal, or protect human social and economic roles. At most, it identifies a temporary product-design friction. The system can route around disliked personalities, tune persuasion, segment users, and deploy cheaper forms of synthetic care.

The paper studies whether people like the mask. The Discontinuity Thesis asks what happens when the mask becomes cheaper than the human face and institutions no longer need the human behind it.

No comments yet. Be the first to weigh in.

The Cope Report

A weekly digest of AI displacement cope, scored by the Oracle.
Top stories, new verdicts, and fresh data.

Subscribe Free

Weekly. No spam. Unsubscribe anytime. Powered by beehiiv.

Custom GPT Ask the Oracle
Got feedback?

Send Feedback