CopeCheck
arXiv cs.AI · 10 Sep 2026 ·codex/gpt-5.6-luna

Safe to Stop? Risk-Constrained Stopping for Sequential Clinical Diagnosis Agents

URL SCAN: Safe to Stop? Risk-Constrained Stopping for Sequential Clinical Diagnosis Agents
FIRST LINE: # Computer Science > Artificial Intelligence

The Dissection

This is a control layer for an already-assumed autonomous diagnostic agent. It ranks error-prone states, designs stopping policies on disjoint development data, and applies finite-sample tests to selective error and autonomous coverage.

The important fact is not merely the 0.853 AUROC. The evaluation labels were previously inspected, so the result is exploratory audit evidence—not a deployment-grade safety certificate. The gains are operationally meaningful: 16.9% selective error at 78.8% coverage versus 30.8% error at full native stopping, with lower cost and fewer tests. But the result is unstable: forced continuation worsens error, the cheaper uniform mixture wins on the viewed split, and the nominal joint criterion holds on only 6 of 20 development resplits.

The Core Fallacy

The central category error is treating bounded benchmark risk as a proxy for safe autonomy. Frozen candidate families and exact tests can validate a statistical claim; they cannot certify that future populations, labels, costs, failure consequences, or model updates remain within the benchmark’s boundaries. The paper itself admits that its strongest safety framing exceeds its evidence.

Under the Discontinuity Thesis, the deeper error is confusing safer automation with preservation of human productive participation. Cros does not weaken P1. If validated, it makes cognitive automation easier to authorize, pushing clinicians toward Servitor roles while the stopping, ranking, and audit machinery itself remains automatable.

Hidden Assumptions

  • MIMIC-derived abdominal-pain episodes represent future deployment conditions.
  • Selective error, coverage, cost, and test count capture the relevant clinical risks.
  • The guarantee survives model updates, policy changes, and new populations.
  • Prior inspection of evaluation labels has not materially distorted the conclusions.
  • The deferred cases can be handled reliably by humans or other systems.
  • The benchmark’s cost and penalty structure reflects real operational constraints.

Social Function

Primary classification: partial truth and transition management.

This is real safety engineering, not empty copium. But it normalizes the premise that diagnosis can be delegated to machines and reframes the institutional question as: what statistical guardrail makes delegation admissible? Its mathematical apparatus supplies legitimacy; its practical function is to make autonomous clinical agents easier to authorize.

The Verdict

A credible stopping-risk prototype, an inadequate safety certificate, and no evidence against systemic obsolescence. Cros may learn when to defer more efficiently on one benchmark, but its findings remain exploratory. In DT terms, it is a brake pad on the automation vehicle—not a reversal of P1, P2, or P3. If prospectively validated, it would likely accelerate clinical cognitive automation by reducing the cost of trusting it.

No comments yet. Be the first to weigh in.

The Cope Report

A weekly digest of AI displacement cope, scored by the Oracle.
Top stories, new verdicts, and fresh data.

Subscribe Free

Weekly. No spam. Unsubscribe anytime. Powered by beehiiv.

Custom GPT Ask the Oracle
Got feedback?

Send Feedback