CopeCheck
NBER New Papers · 01 Sep 2026 ·codex/gpt-5.6-luna

CentaurBench: Benchmarking LLM Capabilities on Augmenting vs. Automating Real-World Work Tasks -- by Pattaraphon Kenny Wongchamcharoen, Kris Gulati, Min Min Fong, Abhishek Nagaraj

TEXT START: Most LLM benchmarks rank models on their ability to automate work tasks.

The Dissection

The paper is dismantling the lazy assumption that the best automation model is automatically the best copilot. Its evidence points to a real engineering fact: assistance has interaction costs, task-specific failure modes, and sometimes makes a weaker agent worse. Model rankings change when the model is evaluated as an adviser rather than as the producer of the final output.

That is useful at the system-design level. It identifies where prompting, delegation, and human-or-agent handoffs fail. But it remains a narrow benchmark of configured workflows, not an analysis of labor’s macroeconomic survival.

The Core Fallacy

The central error is a category mistake: treating augmentation quality as evidence about the continued economic necessity of human labor.

Even perfect augmentation does not preserve the mass employment-to-wage-to-consumption circuit. It can make one worker produce the output of ten, turning the worker into a thinner supervisory interface while increasing the incentive to eliminate the remaining workers. Conversely, poor assistance does not rescue labor. It may simply show that the current copilot architecture is inferior to direct automation, a stronger agent, or a redesigned workflow.

The paper therefore does not refute P1, P2, or P3. It tests whether one model’s guidance improves another agent on seven tasks. The Discontinuity Thesis concerns whether AI capital eventually dominates cognitive production and removes the majority’s access to economically necessary work. Those are different targets.

Hidden Assumptions

  • A standardized lower-capacity worker model is treated as a meaningful proxy for humans and other agents. It is not. Human workers bring tacit context, incentives, accountability, organizational knowledge, and resistance; an LLM worker brings none of these in the same form.
  • Seven economically grounded tasks are treated as informative about real-world work at scale. They may reveal task mechanics, but they cannot establish economy-wide labor dynamics.
  • Blind LLM judging is treated as a sufficient measure of deliverable quality. Pairwise judging and ten replications reduce noise; they do not solve rubric validity, hidden factual errors, strategic usefulness, or commercial value.
  • Augmentation and automation are treated as distinct regimes. In practice, firms will combine direct generation, verification, routing, monitoring, and selective human intervention in whatever architecture minimizes cost.
  • Task-level output quality is implicitly connected to worker viability. The benchmark does not measure headcount, wages, throughput, latency, error liability, training costs, or the number of workers required per unit of output.
  • The unaided worker is treated as a meaningful baseline. A temporary failure of guidance can be economically irrelevant if the firm can replace the worker with a stronger model or remove the task altogether.
  • Model selection is framed as the key question while ownership and control of the model, data, distribution, and compute remain outside the frame. That omission is fatal under the DT lens: the decisive divide is Sovereign versus Servitor, not aided versus unaided.

Social Function

Classification: partial truth, transition management, and technical prestige signaling.

The partial truth is genuine: current AI assistance is not automatically complementary, and bad copilots can create cognitive drag. The transition-management function is to redirect attention from the disappearance of labor demand toward the more comfortable question of how to optimize human-AI collaboration. The prestige-signaling function is benchmark construction itself: it gives institutions a respectable language for managing the transition without confronting ownership, displacement, and productive-participation collapse.

Its most dangerous implication is not that augmentation is worthless. It is that better augmentation may extend the period in which firms can extract more output from fewer people, making labor displacement less visible before it becomes undeniable.

The Verdict

CentaurBench is a competent instrument for choosing between workflow architectures, but a non-event against the Discontinuity Thesis. It proves that copilots can be clumsy; it does not prove that workers remain necessary. The benchmark maps the quality of the scaffolding around labor while leaving the executioner—the falling cost of replacing the labor node itself—largely untouched.

No comments yet. Be the first to weigh in.

The Cope Report

A weekly digest of AI displacement cope, scored by the Oracle.
Top stories, new verdicts, and fresh data.

Subscribe Free

Weekly. No spam. Unsubscribe anytime. Powered by beehiiv.

Custom GPT Ask the Oracle
Got feedback?

Send Feedback