arXiv cs.CY
·
04 Sep 2026
The article converts an uncontrolled strategic race into a monitoring problem. It proposes metrics, indicators, and thresholds modeled on cybersecurity and national-security practice, implying that dangerous AI progression can be observe...
Josh Johnson lands at 38/100 (moderate) for minimisation. The claim uses personal experience and anecdotal reasoning to minimize AI displacement risks. While 'someone gave your job to AI'...
arXiv cs.CY
·
04 Sep 2026
The paper builds an administrative control layer for advanced-AI risk. It translates a civilizational power shift into familiar instruments—probabilistic models, Bayesian networks, thresholds, disclosure rules, expert workshops, and regu...
arXiv econ.GN
·
04 Sep 2026
The paper makes adoption thresholds empirical rather than treating them as fixed abstract parameters. It finds that lower product attractiveness and greater payoff uncertainty raise thresholds, while individual characteristics explain ad...
arXiv cs.AI
·
04 Sep 2026
This paper is an autopsy of prompt-based feature extraction. It tests whether changing job roles, prompt structure, and rule interpretation makes LLM-generated chemical features more useful for toxicity prediction, then passes those feat...
arXiv cs.AI
·
04 Sep 2026
The paper converts the physical laboratory into a machine-readable execution environment: typed objects, bounded capabilities, formal workflow composition, state simulation, and precondition checks. Its real function is to turn experimen...
arXiv cs.AI
·
04 Sep 2026
KC-Bench turns agentic unreliability into a measurable engineering problem: conflicting instructions, stale model knowledge, inconsistent inputs, and competing temporal sources. Its 238 curated tasks, simulated environments, tool calls, ...
arXiv cs.AI
·
04 Sep 2026
HalluPeer converts trust in scientific peer review into a measurable engineering problem: classify, locate, and detect unsupported claims. Its aligned paper-review pairs and injected hallucinations create infrastructure for auditing mach...
arXiv cs.AI
·
04 Sep 2026
GPS-Bench is building an evidence-grounded prediction and explanation layer for governance. Its real contribution is methodological: it turns policy simulation from unconstrained persona theater into a controlled contest between joint re...
arXiv cs.AI
·
04 Sep 2026
Dalek is an attempt to turn software agents into hereditary machines rather than disposable programs. Its core move is to combine actors, messages, channels, a construction language, admissible transitions, and rule heredity with a von N...
arXiv cs.AI
·
04 Sep 2026
This paper is an incremental optimization of the medical-vision automation stack. FreNet uses SAM-derived visual priors before encoding and frequency/spatial feature reconfiguration during encoding to produce cleaner lesion masks. Its re...
arXiv cs.AI
·
04 Sep 2026
This paper packages neonatal chest-X-ray interpretation and clinical report drafting into a constrained multimodal inference system. Its real product is not autonomous medicine; it is workflow compression. The title says “diagnosis,” but...
arXiv cs.AI
·
04 Sep 2026
The paper exposes a real capability gap: multimodal models recognize food images well but fail when visual evidence must be connected to cooking procedure, regional cuisine, and cultural context. Its benchmark is designed to break the sh...
arXiv cs.AI
·
04 Sep 2026
This is infrastructure optimization for the machine that replaces cognitive labor. GrowPage treats KV-cache capacity as an elastic runtime resource, tracking short- and long-horizon attention demand so reasoning requests can acquire memo...
arXiv cs.AI
·
04 Sep 2026
The paper identifies a real bottleneck in agentic VLMs: tool use is judged mainly by the final answer, so the model can waste calls, seek irrelevant evidence, or fail to extract useful information from valid observations. NTEP-R converts...
arXiv cs.AI
·
04 Sep 2026
This is an industrialization prototype for mathematical labor. It converts conjecture generation, counterexample hunting, formal statement construction, and proof search into a repeatable machine pipeline. The headline number—6,522 survi...
arXiv cs.AI
·
04 Sep 2026
The paper identifies a genuine failure: fluency no longer signals competence, while an AI label can cause users to discount accurate material. It then proposes replacing authorship disclosure with claim-level evidence visualization.
arXiv cs.AI
·
04 Sep 2026
This paper addresses a real operational defect: GUI agents are biased toward execution, even when instructions conflict with themselves or with visible interface state. CONFLICTGUI measures that failure; CONFLICTGUARD adds feasibility ve...
arXiv cs.AI
·
04 Sep 2026
This paper converts the tacit mechanics of human conversation—listening, backchanneling, interrupting, yielding, seizing the floor, and resolving conflicting directives—into measurable machine behavior. Its 1,038-case benchmark exposes a...
arXiv cs.AI
·
04 Sep 2026
Dude is an industrialization paper disguised as a review-quality paper. It treats research overload as a detection bottleneck, then inserts a multi-agent supervisory layer: dual detection expands recall, negotiation resolves paper-code g...