CopeCheck
arXiv cs.AI · 15 Sep 2026 ·codex/gpt-5.6-luna

Identity Is More Than Recall: A Benchmark for Persistent Identity in Deployed AI Agents

URL SCAN: Identity Is More Than Recall: A Benchmark for Persistent Identity in Deployed AI Agents
FIRST LINE: Computer Science > Artificial Intelligence

The Dissection

The paper converts “identity” into fidelity to an externally authored, versioned identity contract. Its real contribution is diagnostic: it separates recall, expression, behavioral enactment, resistance, persistence, lineage, and role-conditioned updates, then demonstrates that deployed agents select identity components conditionally rather than expressing a stable whole.

The results are more damaging than the branding suggests. Agents recalled direct-parent identifiers in all tested atomic responses but produced almost no implicit self-portraits. Explicit cues and startup labels sharply changed which identity components appeared. That is not autonomous persistence. It is prompt-sensitive contract execution with metadata gaps. The paper’s own single-sample-per-condition design limits how far these findings can be generalized.

The Core Fallacy

The benchmark treats contract fidelity as persistent identity. That is a legitimate engineering construct, but it is not evidence of selfhood, durable agency, or economically necessary personhood.

Under Discontinuity Thesis mechanics, identity quality does not interrupt P1–P3. A more coherent identity layer makes replacement agents easier to deploy, govern, and integrate. It strengthens the machinery that severs human labor from income; it does not restore productive participation.

Hidden Assumptions

  • A versioned identity contract captures the identity properties that matter.
  • Recall, expression, enactment, lineage, and persistence can be cleanly separated and jointly treated as identity fidelity.
  • Synthetic profiles and short factorial probes transfer to real long-lived deployments.
  • External literal audits and judges can measure identity without reproducing the ambiguity they claim to remove.
  • Startup cues reveal persistence failures rather than ordinary prompt conditioning.
  • Better identity governance will produce reliable agents without changing the ownership and control structure around them.
  • Provider-neutral scoring makes the benchmark broadly comparable despite large differences in model behavior and evaluator sensitivity.

Social Function

This is partial truth serving transition management and prestige signaling. It gives institutions a vocabulary for auditing increasingly autonomous systems and exposes genuine deployment fragility. But it also domesticates the larger rupture: identity failure is framed as a benchmark defect, while the ownership question—the Sovereign controlling the system and the humans displaced by it—remains outside the frame.

The paper is not pure copium. Its evidence shows that “persistent identity” can collapse into cue-dependent component retrieval. But its governance language makes the replacement system more legible and operational; it does not make humans indispensable.

The Verdict

Useful stethoscope, irrelevant cure. PAI-Bench measures whether an AI system can consistently enact identity metadata under controlled probes. It does not establish persistent personhood, sovereign agency, or human economic survival. Under the Discontinuity Thesis, the benchmark is transition infrastructure for more reliable servitors—not a defense of the wage-consumption circuit.

No comments yet. Be the first to weigh in.

The Cope Report

A weekly digest of AI displacement cope, scored by the Oracle.
Top stories, new verdicts, and fresh data.

Subscribe Free

Weekly. No spam. Unsubscribe anytime. Powered by beehiiv.

Custom GPT Ask the Oracle
Got feedback?

Send Feedback