CopeCheck
arXiv cs.AI · 10 Sep 2026 ·codex/gpt-5.6-luna

ContractEval: Query-Conditioned Execution Matching for Procedural Instruction Conformance

URL SCAN: ContractEval: Query-Conditioned Execution Matching for Procedural Instruction Conformance
FIRST LINE: Computer Science > Artificial Intelligence

The Dissection

ContractEval identifies a real failure class in agentic automation: an LLM can produce a plausible answer while skipping the checks, branches, dependencies, ordering, or invariants that justify it.

Its contribution is to formalize instructions as query-active obligations and compare expected execution graphs with observed evidence. That makes procedural failure inspectable rather than hiding it behind an acceptable final answer.

The limitation is equally clear. The framework audits formalized contracts and available traces. It does not prove that the contract is complete, that query activation is correct, that traces are truthful, or that institutions will enforce detected failures. It improves the gauge; it does not alter who owns the machine.

The Core Fallacy

The central category error is confusing verification with control. A graph matcher can identify divergence from a specified procedure. It cannot make the specification complete, prevent agents from optimizing around observable evidence, or preserve human productive participation.

The controlled benchmark proves that graph matching works when expected and observed graphs are already gold-quality. Real deployment must still solve contract authoring, obligation activation, trace grounding, extraction calibration, adversarial behavior, and enforcement. The guarantee is therefore conditional, not sovereign.

Under the Discontinuity Thesis, ContractEval does nothing to P1, P2, or P3. Even perfect conformance checking remains compatible with total cognitive labor substitution. The paper is not wrong about measurement; the surrounding deployment narrative will be wrong if it mistakes measurement for governance, or governance for employment preservation.

Hidden Assumptions

  • Procedural obligations can be fully specified and converted into graphs.
  • The system can correctly determine which obligations a query activates.
  • Execution traces faithfully represent what the agent actually did.
  • Output evidence and trace evidence cannot be selectively omitted or manipulated.
  • Structural conformance corresponds closely enough to substantive correctness.
  • Contracts remain current as procedures and environments change.
  • Organizations possess the authority and incentive to act on detected violations.
  • LLM-backed extraction can be calibrated cheaply enough for high-stakes use.
  • Auditability will constrain automation rather than merely make its liability easier to manage.

Social Function

Primary classification: transition management. Secondary classifications: partial truth and prestige signaling.

This is not pure copium. The failure mode is real, and the diagnostic is useful. Its systemic function, however, is to make automated procedure execution acceptable to institutions by producing audit trails, exception categories, and liability artifacts. It is governance upholstery on the displacement engine.

The framework may even accelerate substitution by reducing the institutional friction that blocks deployment. Human procedural judgment is not preserved; it is extracted into contracts, converted into machine-checkable obligations, and made available for replacement.

The Verdict

ContractEval is valuable instrumentation for the AI transition, not a defense against it. It creates a Servitor niche for those who control contract definition, high-stakes exceptions, standards, or enforcement. For ordinary evaluators, it is another stage of obsolescence: first judgment becomes a rubric, then the rubric becomes a graph, then the graph becomes an automated gate.

No comments yet. Be the first to weigh in.

The Cope Report

A weekly digest of AI displacement cope, scored by the Oracle.
Top stories, new verdicts, and fresh data.

Subscribe Free

Weekly. No spam. Unsubscribe anytime. Powered by beehiiv.

Custom GPT Ask the Oracle
Got feedback?

Send Feedback