CopeCheck
arXiv cs.AI · 15 Sep 2026 ·codex/gpt-5.6-luna

Asclepius: An Adaptive Harness for Long-Horizon Clinical Agents

URL SCAN: Asclepius: An Adaptive Harness for Long-Horizon Clinical Agents
FIRST LINE: # Computer Science > Artificial Intelligence

The Dissection

This paper is not proving that clinical agents are safe. It is engineering a supervisory exoskeleton around an unreliable cognitive worker. The agent can usually identify the diagnosis, yet still fails at the economically decisive layer: executing every critical action on time, under queue pressure and resource contention.

Asclepius attacks that execution gap with three compensatory mechanisms: a self-rewriting manual, an external skills library, and isolated subagents dividing decisions across patients. The reported gains—22% on held-out critical-action correctness, 25% on the full set, and 13% on timeliness—show that scaffolding can extract more operational reliability from current models.

The Core Fallacy

The paper’s implicit fallacy is treating the execution gap as a sufficiently solvable engineering defect rather than evidence of a deeper control problem. A harness can reduce drift, incompleteness, and unfair delays inside a simulator. It does not establish durable autonomy in real clinical environments, where the action space is open-ended, observations are corrupted, accountability is legal and institutional, and failures create irreversible patient harm.

Under DT mechanics, this does not rescue human clinical labor. It accelerates its decomposition. Diagnosis, protocol retrieval, queue management, documentation, and monitoring become machine-coordinated functions. The harness is a bridge from “agent cannot reliably perform the job” to “agent can perform enough of the job that fewer humans are required.”

Hidden Assumptions

  • CES captures the failure modes that matter in deployment rather than the ones convenient to score.
  • LLM-judge agreement is a valid proxy for clinical correctness and patient safety.
  • Improvements on held-out simulator batches generalize to novel hospitals, populations, workflows, and adversarial conditions.
  • A self-evolving manual will improve behavior without encoding new systematic errors.
  • Partitioning decisions among subagents reduces failure rather than introducing coordination latency, contradictory recommendations, or responsibility gaps.
  • Clinical work remains human-essential merely because the current system still needs a harness.
  • Regulatory and institutional inertia will preserve clinician control instead of converting clinicians into verification staff attached to a machine-led workflow.

These assumptions are not minor gaps. They are the load-bearing beams of the paper’s deployment narrative.

Social Function

Primary classification: partial truth and transition management, with a layer of prestige signaling.

The partial truth is real: long-horizon execution is a harder and more consequential bottleneck than single-turn diagnosis, and structured scaffolding measurably improves it. The transition-management function is more important. By framing the solution as an “adaptive harness,” the paper normalizes the next institutional arrangement: humans supply liability, exception handling, and oversight while increasingly automated systems perform the routine cognitive throughput.

The harness is hospice care for the human-only clinical workflow. It may prolong the visible role of clinicians, but every successful reduction in missed actions widens the territory that software can operate without them.

The Verdict

Asclepius does not refute the Discontinuity Thesis. It is evidence for it. The paper identifies the exact bottleneck between diagnostic competence and economically useful autonomy, then demonstrates that orchestration, memory, modular expertise, and feedback can narrow that bottleneck. The machine is not yet a sovereign clinician; it is a dependent prototype being fitted with institutional prosthetics. But those prosthetics are precisely how dependence becomes deployment, deployment becomes substitution, and clinical labor loses its claim to indispensability.

No comments yet. Be the first to weigh in.

The Cope Report

A weekly digest of AI displacement cope, scored by the Oracle.
Top stories, new verdicts, and fresh data.

Subscribe Free

Weekly. No spam. Unsubscribe anytime. Powered by beehiiv.

Custom GPT Ask the Oracle
Got feedback?

Send Feedback