The Cope Index
Tracking who's coping hardest about the end of work
CopeCheck scores public statements about AI and jobs by how much they rely on denial, deflection, or false reassurance.
23 figures tracked · 11776 articles autopsied
arXiv cs.AI
·
04 Sep 2026
KC-Bench turns agentic unreliability into a measurable engineering problem: conflicting instructions, stale model knowledge, inconsistent inputs, and competing temporal sources. Its 238 curated tasks, simulated environments, tool calls, ...
arXiv cs.AI
·
04 Sep 2026
HalluPeer converts trust in scientific peer review into a measurable engineering problem: classify, locate, and detect unsupported claims. Its aligned paper-review pairs and injected hallucinations create infrastructure for auditing mach...
arXiv cs.AI
·
04 Sep 2026
GPS-Bench is building an evidence-grounded prediction and explanation layer for governance. Its real contribution is methodological: it turns policy simulation from unconstrained persona theater into a controlled contest between joint re...
arXiv cs.AI
·
04 Sep 2026
Dalek is an attempt to turn software agents into hereditary machines rather than disposable programs. Its core move is to combine actors, messages, channels, a construction language, admissible transitions, and rule heredity with a von N...
arXiv cs.AI
·
04 Sep 2026
This paper is an incremental optimization of the medical-vision automation stack. FreNet uses SAM-derived visual priors before encoding and frequency/spatial feature reconfiguration during encoding to produce cleaner lesion masks. Its re...
arXiv cs.AI
·
04 Sep 2026
This paper packages neonatal chest-X-ray interpretation and clinical report drafting into a constrained multimodal inference system. Its real product is not autonomous medicine; it is workflow compression. The title says “diagnosis,” but...
arXiv cs.AI
·
04 Sep 2026
The paper exposes a real capability gap: multimodal models recognize food images well but fail when visual evidence must be connected to cooking procedure, regional cuisine, and cultural context. Its benchmark is designed to break the sh...
arXiv cs.AI
·
04 Sep 2026
This paper studies a narrow but consequential bottleneck in long-context LLM serving: deciding which KV-cache entries to retain during decoding. Its real contribution is not a superior scorer. It shows that temporal aggregation—especiall...
arXiv cs.AI
·
04 Sep 2026
This paper converts scheduling expertise into a learned machine policy operating across task and infrastructure graphs. PPO handles policy optimization; STGNN compresses topology, resource state, and temporal dynamics into automated deci...
arXiv cs.AI
·
04 Sep 2026
This is infrastructure optimization for the machine that replaces cognitive labor. GrowPage treats KV-cache capacity as an elastic runtime resource, tracking short- and long-horizon attention demand so reasoning requests can acquire memo...
arXiv cs.AI
·
04 Sep 2026
The paper identifies a real bottleneck in agentic VLMs: tool use is judged mainly by the final answer, so the model can waste calls, seek irrelevant evidence, or fail to extract useful information from valid observations. NTEP-R converts...
arXiv cs.AI
·
04 Sep 2026
This is an industrialization prototype for mathematical labor. It converts conjecture generation, counterexample hunting, formal statement construction, and proof search into a repeatable machine pipeline. The headline number—6,522 survi...
arXiv cs.AI
·
04 Sep 2026
The paper identifies a genuine failure: fluency no longer signals competence, while an AI label can cause users to discount accurate material. It then proposes replacing authorship disclosure with claim-level evidence visualization.
arXiv cs.AI
·
04 Sep 2026
This paper addresses a real operational defect: GUI agents are biased toward execution, even when instructions conflict with themselves or with visible interface state. CONFLICTGUI measures that failure; CONFLICTGUARD adds feasibility ve...
arXiv cs.AI
·
04 Sep 2026
This paper converts the tacit mechanics of human conversation—listening, backchanneling, interrupting, yielding, seizing the floor, and resolving conflicting directives—into measurable machine behavior. Its 1,038-case benchmark exposes a...
arXiv cs.AI
·
04 Sep 2026
Dude is an industrialization paper disguised as a review-quality paper. It treats research overload as a detection bottleneck, then inserts a multi-agent supervisory layer: dual detection expands recall, negotiation resolves paper-code g...
arXiv cs.AI
·
04 Sep 2026
The paper isolates a real conversational failure: multi-turn narration can pull an LLM toward the narrator’s self-justifying interpretation even without explicit adversarial pressure. Its benchmark tests 5,078 interpersonal conflicts acr...
arXiv cs.AI
·
04 Sep 2026
The paper converts six learner descriptors and a Bloom’s Taxonomy estimate into structured prompts that produce different response styles. Its real contribution is a control layer that makes a general-purpose model appear adaptive withou...
arXiv cs.AI
·
04 Sep 2026
The paper isolates a real causal-consistency failure: current state is not the same as a valid authorization. PlanFence binds plans to the exact records they used, then validates only action-relevant dependencies before execution. Its ev...
arXiv cs.AI
·
04 Sep 2026
This is branch prediction and transaction commit for tool-using agents. A fast drafter speculatively executes a chain on an isolated snapshot; the larger actor validates the trajectory and commits the precomputed steps when the initial a...