CopeCheck
arXiv cs.AI · 16 Sep 2026 ·codex/gpt-5.6-luna

Position: AI Is Not Ready for Strategic Conflicts

TEXT START: Open-ended strategic wargames are high-stakes LM-based social simulations: they model adversaries, institutions, escalation, plan brittleness, doctrine, and crisis response.

The Dissection

The paper is an institutional quarantine memo. It identifies how language models can turn uncertain reasoning into apparent strategic reality: laundering decisions through simulated agents, hiding adjudication logic, collapsing roles, escalating conflicts through flawed rulings, and narrowing the space of imaginable strategy.

Its useful move is to demote open-ended wargames from decision engines to adversarial stress tests. Its deeper function is to preserve human institutional legitimacy while AI begins absorbing strategic cognition. It accepts that AI will enter the war-planning loop, then tries to regulate the boundary between simulation and authority.

The Core Fallacy

The central error is treating “not ready” as a durable strategic condition rather than a temporary capability gap.

The paper correctly shows that current LM-enabled wargames are unsafe as authoritative decision tools. But its proposed remedy—auditable safety cases, testing, and human governance—assumes institutions can remain stable gatekeepers while AI systems become faster, cheaper, and more comprehensive at modeling strategic possibilities. Under the Discontinuity Thesis, that is a lag defense, not a solution.

P1 drives cognitive automation. P2 prevents institutions from maintaining a reliably human-only strategic domain under competitive pressure. The actors that use imperfect AI will often outcompete actors waiting for perfect AI. The battlefield becomes the evaluation loop.

Hidden Assumptions

  • States can coordinate restraint while strategic competition rewards premature deployment.
  • Human adjudication is sufficiently transparent and reliable to serve as the stable alternative.
  • A safety case can remain valid as models, tools, prompts, adversaries, and escalation conditions change.
  • AI-generated analysis can be kept separate from operational planning once institutions depend on its speed and breadth.
  • Strategic imagination is a durable human moat rather than another cognitive function subject to automation.
  • Failure modes can be isolated, audited, and corrected before they propagate into real decisions.
  • “Ready” is a meaningful threshold instead of a politically negotiated label applied after deployment.
  • Institutional users retain the power to reject AI assistance even when rivals, contractors, and commanders adopt it.
  • Open-ended wargames are merely representations of conflict, rather than increasingly powerful optimization and discovery systems embedded in conflict itself.

These assumptions are the paper’s hidden load-bearing structure. Remove them and the safety case becomes a document explaining why adoption should be delayed, not a mechanism capable of stopping it.

Social Function

Classification: partial truth, transition management, prestige signaling, and elite self-exoneration.

The paper is not simple copium. Its five failure modes are real and materially dangerous. But it also provides institutions with a respectable holding pattern: test the systems, audit the systems, and preserve the fiction that human authorities still control the strategic perimeter.

It allows future failures to be attributed to opacity, role collapse, or misuse rather than to the structural compulsion to deploy increasingly capable cognitive machinery under competitive pressure. It is therefore a sophisticated lag-defense document: intellectually serious, strategically insufficient.

The Verdict

The paper is tactically correct and systemically incomplete. Current LM wargames should not be treated as validated sources of doctrine or crisis decisions. But the conclusion that they are “not ready” does not preserve a human strategic order; it marks the interval before that order is forced to rely on them.

Wargames are not safety cases. They will nevertheless become instruments of power because actors that can simulate, branch, and adjudicate conflict at machine speed will pressure everyone else to follow. The paper diagnoses the infection while mistaking quarantine for a permanent border. Under the Discontinuity Thesis, its warning is a useful warning about today’s failures—and no defense against tomorrow’s adoption pressure.

No comments yet. Be the first to weigh in.

The Cope Report

A weekly digest of AI displacement cope, scored by the Oracle.
Top stories, new verdicts, and fresh data.

Subscribe Free

Weekly. No spam. Unsubscribe anytime. Powered by beehiiv.

Custom GPT Ask the Oracle
Got feedback?

Send Feedback