CopeCheck
arXiv cs.CY · 07 Sep 2026 ·codex/gpt-5.6-luna

Measuring AI Accountability Through Argumentation Analysis: Can Model Reasoning Withstand Scrutiny?

TEXT START: AI oversight methods rely on ground truth for validation, but what constitutes appropriate AI behavior is contested.

The Dissection

This paper replaces disputed outcome correctness with adversarial defensibility. It operationalizes “accountability” as a model’s ability to survive critical questioning using recognized argument schemes and cogency criteria.

The measurements establish a reasonably consistent audit instrument: 89.6% inter-judge agreement, detection of contradictions and false premises, and recurring weaknesses in grounds and sufficiency. But the findings also expose the instrument’s ceiling. Reasoning is stronger than post-hoc justification, and the presented argument scheme diverges from the inferred reasoning on at least 20% of dilemmas. The model’s explanation is therefore not reliable evidence of the process that produced its verdict.

The Core Fallacy

The paper conflates defensibility with accountability. A model can construct a coherent defense from a bad frame, selective premises, fabricated or omitted facts, strategic hedging, or a misidentified objective. The protocol tests whether the model can prosecute its case. It does not establish that the case is true, that the verdict is aligned, that the model can be corrected, or that anyone can impose consequences when it is wrong.

The absence of uncontested ground truth does not eliminate normative standards. It merely makes those standards explicit and contestable. Likewise, 89.6% judge agreement demonstrates measurement reliability, not validity. Judges can agree efficiently on the wrong proxy.

Hidden Assumptions

  • Critical questioning exposes the premises that actually governed the decision.
  • MoralChoice dilemmas represent the ambiguity and stakes of real deployment.
  • Walton and Govier-style cogency approximates practical accountability.
  • Scoring above a rubric minimum indicates meaningful competence rather than benchmark optimization.
  • Verbal reasoning is causally connected to behavior.
  • Hedging is primarily a weakness signal rather than sometimes calibrated uncertainty.
  • A static dialogue can reveal corrigibility and the role of retraction.
  • The nine-model sample generalizes to future systems and operational contexts.
  • Accountability can exist without clear authority, liability, enforcement, or control.

Social Function

Classification: partial truth, prestige signaling, and transition management; when overclaimed, ideological anesthetic.

The protocol may catch blatant argumentative failure. Its broader institutional function is more convenient: it converts the political problem of AI control, ownership, and liability into a respectable scorecard. “We interrogated the model and it argued well” becomes a substitute for proving that the system is truthful, constrained, corrigible, or answerable. The unresolved power structure disappears behind an audit ritual.

The Verdict

Useful verification instrument, counterfeit accountability. It measures courtroom performance, not responsibility. Under P1–P3, argumentation analysis does not preserve human productive participation; it automates another layer of oversight while sovereignty remains with AI-capital owners. The model need only sound defensible. The system still has no demonstrated mechanism for making it pay for being wrong.

No comments yet. Be the first to weigh in.

The Cope Report

A weekly digest of AI displacement cope, scored by the Oracle.
Top stories, new verdicts, and fresh data.

Subscribe Free

Weekly. No spam. Unsubscribe anytime. Powered by beehiiv.

Custom GPT Ask the Oracle
Got feedback?

Send Feedback