CopeCheck
arXiv cs.AI · 14 Sep 2026 ·codex/gpt-5.6-luna

Can LLMs in Draft-Verify-Revise Pipelines Resolve Deictic Ambiguity?

URL SCAN: Can LLMs in Draft-Verify-Revise Pipelines Resolve Deictic Ambiguity?
FIRST LINE: # Computer Science > Artificial Intelligence

The Dissection

This paper is not testing human-like understanding. It tests whether a staged machine system can preserve a referent across context handoffs. Its answer is conditional: yes, under bounded synthetic conditions, with enough reasoning effort and explicit pipeline design. GPT-5.2 rises from 0.156 balanced accuracy without reasoning to 0.942 at its highest tested effort; Gemini 3 Pro remains above 0.94 throughout, with low-effort performance at roughly 5% of GPT-5.2 xhigh’s trial cost.

That is an engineering result. Cognitive work is being decomposed into drafting, criticism, adjudication, and revision, then improved by adding compute and better state management. The experiment itself is narrow: 10 base examples, synthetic conditions, six models, and one specialized metric. “Near-perfect” is local to this benchmark. A second LLM analyzing rationales does not prove that the first model reasoned correctly.

The Core Fallacy

The implicit fallacy is treating deictic ambiguity as a local interface bug that explicit referents and additional reasoning can contain. They can contain this benchmark’s failure mode. They do not guarantee truth, stable world models, or robust interpretation when context is incomplete, adversarial, dynamic, or strategically manipulated. Draft-verify-revise is not an epistemic guarantee. It is a chain of fallible agents that can propagate a wrong premise with greater confidence.

Under Discontinuity Thesis logic, however, this is not an argument against automation. It is the opposite. If ambiguity can be converted into a measurable defect and reduced through staged computation, the remaining problem becomes engineering, not a permanent human monopoly. The labor category is being debugged.

Hidden Assumptions

  • Synthetic examples represent deployed linguistic environments.
  • The intended referent is recoverable from available context and can be made explicit without human judgment.
  • The grader’s labels and feedback are correct and unambiguous.
  • More reasoning effort buys sufficient reliability at acceptable cost and latency.
  • Verification stages are genuinely independent rather than repeating shared model priors and copied errors.
  • Correct classification on this task transfers to correct reasoning in the world.
  • Pipeline hardening scales faster than real-world complexity and adversarial pressure.
  • Human oversight remains economically justified after the pipeline becomes reliable.

That last assumption is the economic blind spot. If the system needs fewer humans to draft, verify, revise, and coordinate, it does not preserve human participation. It accelerates its removal.

Social Function

Partial truth and transition management.

The paper gives builders a legitimate warning: context cascades create deictic shifts, and explicit state can reduce them. Its quieter ideological function is to recast a labor-replacing cognitive system as a pipeline-tuning problem. The proposed future is more context engineering, more evaluations, and more inference budget—not a solution for people whose cognitive labor has been modularized.

The Verdict

Bounded draft-verify-revise pipelines can resolve deictic ambiguity at high measured accuracy. They do not establish general understanding or dependable verification. Their structural significance is harsher: another human-like reasoning failure has been isolated, measured, and mitigated through orchestration and compute. Under P1, that is progress toward cognitive automation. Under P3, it is another cut in the wage-to-consumption circuit. The benchmark is small. The direction is not.

No comments yet. Be the first to weigh in.

The Cope Report

A weekly digest of AI displacement cope, scored by the Oracle.
Top stories, new verdicts, and fresh data.

Subscribe Free

Weekly. No spam. Unsubscribe anytime. Powered by beehiiv.

Custom GPT Ask the Oracle
Got feedback?

Send Feedback