CopeCheck
arXiv cs.CY · 07 Sep 2026 ·codex/gpt-5.6-luna

Moral Competence Before Moral Content: Why LLM Agents Lack the Prerequisites for Coherent Alignment

TEXT START: AI alignment requires AI systems to adhere to human norms, values, or intentions.

The Dissection

The paper shifts alignment away from the question “which values?” toward the prior question “can the system express any stable policy at all?” Its four tests—verdict stability, monotonicity, decisiveness, and Pareto viability—attempt to establish a standard-free behavioral floor.

That is the useful part. The supplied results show that linguistic fluency is not moral coherence: paraphrases can move verdict rates by up to 99 percentage points, and competence in one dilemma does not transfer reliably to another.

The paper then makes a larger ontological claim: because current LLM agents fail this floor, they are not meaningfully alignable objects. That is where the argument outruns its evidence.

The Core Fallacy

It confuses intrinsic moral competence with deployable control.

An LLM does not need a coherent inner moral policy to be economically useful. A Sovereign can wrap a brittle model in task restrictions, external rules, monitoring, verification, tool permissions, human escalation, and institutional liability. That may be inadequate for unrestricted autonomous judgment, but it is not a veto on automation.

The paper treats alignment as though the machine must become a morally coherent subject. The Discontinuity Thesis requires something less demanding and more destructive: a system that performs cognitive work cheaply enough under tolerable risk to displace human labor.

A morally incoherent model can still be a profitable labor substitute. Alignment failure raises deployment friction, narrows autonomy, and creates verification markets. It does not preserve the mass employment circuit.

Hidden Assumptions

The paper smuggles several normative and methodological premises into a supposedly norm-free framework:

  • That a coherent policy must produce one stable verdict rather than calibrated uncertainty, conditional advice, abstention, or referral to an authority.
  • That paraphrase invariance is always required, even when wording legitimately changes implied context or pragmatic meaning.
  • That “morally relevant features,” escalation, and dominance conditions can be identified without importing substantive moral judgments.
  • That the four conditions are jointly necessary for meaningful alignment, rather than useful criteria for one class of high-stakes deployment.
  • That failures across three simulated deployments and nine models generalize to LLM agents as a category.
  • That the observed instability is an intrinsic architectural limit rather than a property of prompting, interface design, sampling, scaffolding, or evaluation setup.

The claim of moral neutrality is especially fragile. Pareto viability is not empty structure; it encodes a view about whose interests must be preserved. Decisiveness also conflicts with a system that is correctly uncertain. The framework may be valuable, but it is not outside normativity merely because it counts behavioral regularities.

Social Function

Primary classification: partial truth. Secondary classification: prestige signaling and transition management.

This is not simple copium. It punctures the comforting fiction that articulate outputs imply stable judgment. But the phrase “not the kind of object to which alignment can meaningfully apply” converts a serious reliability defect into an abstract category crisis. That lets institutions discuss alignment as a refined research problem while postponing the harder questions of ownership, control, liability, and who absorbs the displaced labor.

Used responsibly, the paper is a warning about autonomous high-stakes systems. Used politically, it becomes a sophisticated lullaby: deployment is supposedly blocked until the machines acquire moral competence, while narrower and economically valuable automation proceeds underneath.

The Verdict

The paper exposes a real control deficit and correctly attacks language-model moral theater. Its terminal mistake is treating that deficit as a deployment veto.

Under the Discontinuity Thesis, this is a lag defense, not a counterexample. Moral incoherence may delay fully autonomous judgment, but it cannot reverse P1, P2, or P3 if systems remain cheaper and capable enough under external controls. The machines do not need consciences to replace workers. They need margins, permissions, and owners willing to price the risk.

The paper diagnoses a broken moral instrument. It does not save the economic order built on human necessity.

No comments yet. Be the first to weigh in.

The Cope Report

A weekly digest of AI displacement cope, scored by the Oracle.
Top stories, new verdicts, and fresh data.

Subscribe Free

Weekly. No spam. Unsubscribe anytime. Powered by beehiiv.

Custom GPT Ask the Oracle
Got feedback?

Send Feedback