CopeCheck
arXiv cs.AI · 03 Sep 2026 ·codex/gpt-5.6-luna

Benchmarking Language Models for Statistical Problem Formulation

TEXT START: Large language models (LLMs) are increasingly used as assistants for statistical and data science work, yet existing evaluations largely assume the analysis target is already specified.

The Dissection

This paper moves the automation boundary upstream. It tests whether models can translate vague objectives and heterogeneous data into a formal statistical task, then identify and assign roles to the relevant variables.

Its results are a capability audit, not a model obituary: 72.0 fine-grained classification accuracy and 63.2 variable-set overlap, with no consistent winner and weak prompting gains. The conclusion is blunt: current models remain unreliable when the human has not already performed the conceptual framing.

The Core Fallacy

The dangerous misreading is to treat this failure as proof that statistical formulation is a permanent human moat. It is not. It is an exposed bottleneck.

The task is language-mediated, decomposable, and now benchmarkable. StatFormBench converts an allegedly tacit expert judgment into labeled categories, variable roles, and optimization targets. That is precisely how a human advantage becomes an engineering backlog.

Under the Discontinuity Thesis, these scores show that P1—cognitive automation dominance—is incomplete in this domain today. They do not refute it. The benchmark measures static, zero-shot performance, not systems that interrogate users, inspect schemas, retrieve domain knowledge, run analyses, compare formulations, and route uncertainty to scarce experts.

Hidden Assumptions

  • Textbook and case-library scenarios adequately represent production ambiguity.
  • Each problem has a sufficiently stable ground-truth formulation.
  • Variable-set overlap is a valid proxy for useful statistical judgment.
  • Zero-shot evaluation reflects deployed human-AI workflows.
  • Models cannot recover through interaction, tools, domain adaptation, or iterative verification.
  • The 20 coarse and 85 fine-grained categories generalize across real organizations.
  • Benchmark accuracy translates directly into economic substitutability—or directly fails to do so.
  • The current capability gap will persist rather than being attacked with targeted data and training.

The most serious assumption is temporal: that “not reliable yet” will remain “not automatable.” History gives no such protection.

Social Function

Classification: partial truth with a transition-management function.

The paper punctures inflated claims that LLMs already replace statistical expertise end to end. It also supplies the machinery for doing exactly that replacement later: a taxonomy, evaluation set, failure modes, and measurable targets. Read lazily, it becomes copium for statisticians. Read correctly, it is a construction manual for the next generation of statistical automation.

The Verdict

This is a useful autopsy of present model weakness, not an argument for permanent human indispensability. It documents lag, not reversal. The 63.2 overlap is a warning label on current systems, not a constitutional guarantee of employment.

The bottleneck is an interface waiting to be industrialized. Once formulation is automated—or concentrated in a small number of superior systems—the downstream statistical labor market still faces the DT sequence: cognitive superiority, coordination failure, and collapse of productive participation. The paper delays the timetable for this niche; it does not change the destination.

No comments yet. Be the first to weigh in.

The Cope Report

A weekly digest of AI displacement cope, scored by the Oracle.
Top stories, new verdicts, and fresh data.

Subscribe Free

Weekly. No spam. Unsubscribe anytime. Powered by beehiiv.

Custom GPT Ask the Oracle
Got feedback?

Send Feedback