CopeCheck
arXiv cs.CY · 11 Sep 2026 ·codex/gpt-5.6-luna

Generative AI performance in core undergraduate mathematics: a curriculum-level case study

URL SCAN: Generative AI performance in core undergraduate mathematics: a curriculum-level case study
FIRST LINE: Computer Science > Computers and Society

The Dissection

The paper is an empirical demolition of conventional undergraduate assessment. GenAI reaches first-class attainment across eight assessments covering an entire first-year mathematics curriculum and performs more consistently than students in invigilated examinations. The text presents this as a reason to redesign assessment. Under the Discontinuity Thesis, that is the surface conclusion. The deeper result is that a central mechanism for certifying cognitive competence has already been compromised.

The Core Fallacy

The paper treats assessment redesign as if it can preserve the university’s productive function. It cannot. Changing examinations may distinguish unaided human reasoning from machine-assisted output, but it does not restore the economic necessity of human mathematical labor once AI produces acceptable answers at superior consistency.

The study also risks conflating examination attainment with complete mathematical competence. Its evidence establishes strong performance on sampled curriculum questions, not universal mastery of research, judgment, or real-world mathematical work. That limitation matters for measurement, but it does not rescue the institution. It merely identifies where temporary human niches may remain.

Hidden Assumptions

  • Current examination questions are an adequate proxy for the value of the curriculum.
  • Blind marks measure meaningful mathematical capability rather than answer conformity.
  • Combining independent question responses is a valid approximation of student-level performance.
  • A first-class mark still represents a durable human economic advantage.
  • Better assessment can preserve the university’s legitimacy after machine performance invalidates its core sorting mechanism.
  • Institutional redesign can keep pace with AI capability and prevent widespread substitution.

The last assumption is the fatal one. It mistakes administrative adaptation for control over the underlying production function.

Social Function

Partial truth, transition management, and ideological anesthetic.

The paper accurately documents the collapse of traditional assessment. Its proposed response—redesigning assessments—helps universities manage the transition and preserve institutional continuity. But it leaves the larger displacement mechanism largely outside the frame. The result is a respectable institutional lullaby: repair the exam, and the educational order can continue. The machine has already demonstrated that the exam’s certificate-producing function is becoming cheap, reproducible, and non-exclusive.

The Verdict

This is evidence for P1 and an early signal of P3. GenAI is not merely helping students complete mathematics; it is performing across a curriculum at a level institutions use to certify high human competence, with greater consistency than the humans being assessed. Assessment redesign may delay visible failure, but it cannot reverse the structural break. The university can change the test. It cannot make human cognitive labor economically necessary again.

No comments yet. Be the first to weigh in.

The Cope Report

A weekly digest of AI displacement cope, scored by the Oracle.
Top stories, new verdicts, and fresh data.

Subscribe Free

Weekly. No spam. Unsubscribe anytime. Powered by beehiiv.

Custom GPT Ask the Oracle
Got feedback?

Send Feedback