arXiv cs.CY
·
04 Sep 2026
OBER+ is an instrumentation layer for educational bureaucracy. It links five previously disconnected operations: measuring attainment, detecting persistent shortfalls, selecting a corrective practice, recording the intervention, and meas...
arXiv cs.CY
·
04 Sep 2026
This paper builds an audit-and-control layer around LLM decision-making. It moves clinical-trial matching from opaque neural judgment into explicit policies encoded as SMT/MaxSMT constraints, then derives decisions and rationales from th...
arXiv cs.CY
·
04 Sep 2026
The paper models LLM adoption as contagion: transmission, coupling, persistence, recovery, reinforcement, tipping points, lock-in, and competence loss. Its useful insight is nonlinear adoption. Its strategic weakness is that it treats AI...
arXiv cs.CY
·
04 Sep 2026
CHARM is a classifier presented as a psychologically grounded measurement system. Its reported benchmark gains establish improved task performance, while the COVID-19 Twitter application extends that performance into a broader claim: tha...
arXiv cs.CY
·
04 Sep 2026
This paper is an instrument for ranking artists inside the existing music economy. It compares Polish, Italian, Danish, and merged collaboration networks, then tests whether graph structure improves popularity prediction. Its own result ...
arXiv cs.CY
·
04 Sep 2026
This paper dismantles the fiction that an “LLM judge” is a single objective instrument. It treats models, providers, and versions as raters with measurable severity, halo, reliability, and drift. The results are damaging: severity spread...
arXiv cs.CY
·
04 Sep 2026
This survey’s real function is to consolidate web security and LLM security into one operational control problem. It maps the expanded attack surface—client, server, pipeline, prompt, output, and autonomous-agent layers—and proposes vali...
arXiv cs.CY
·
04 Sep 2026
The paper converts a structural inequality problem into an implementation problem. TechMate packages more than 25 recommended actions, case studies, and resources, then validates the package primarily through educator perceptions: 18 par...
arXiv cs.CY
·
04 Sep 2026
This is a legitimacy framework for algorithmic rule. It correctly identifies that formal fairness metrics do not guarantee that affected people will experience a system as just, trustworthy, or acceptable. Its proposed bridge—literature ...
arXiv cs.CY
·
04 Sep 2026
The paper is not really analyzing the future of the Metaverse. It is constructing a governance and procurement rubric for choosing a delivery stack. Its central move is to recast WebXR and commercial engines as complementary points on a ...
arXiv cs.CY
·
04 Sep 2026
The paper builds a procedural containment device around an uncontained capability. Its 5P sequence—Purpose, Process, Product, Pitfalls, and Plan—makes AI-assisted student work more auditable, but it does not make that work scarce, unique...
arXiv cs.CY
·
04 Sep 2026
The article converts an uncontrolled strategic race into a monitoring problem. It proposes metrics, indicators, and thresholds modeled on cybersecurity and national-security practice, implying that dangerous AI progression can be observe...
Josh Johnson lands at 38/100 (moderate) for minimisation. The claim uses personal experience and anecdotal reasoning to minimize AI displacement risks. While 'someone gave your job to AI'...
arXiv cs.CY
·
04 Sep 2026
The paper builds an administrative control layer for advanced-AI risk. It translates a civilizational power shift into familiar instruments—probabilistic models, Bayesian networks, thresholds, disclosure rules, expert workshops, and regu...
arXiv econ.GN
·
04 Sep 2026
The paper makes adoption thresholds empirical rather than treating them as fixed abstract parameters. It finds that lower product attractiveness and greater payoff uncertainty raise thresholds, while individual characteristics explain ad...
arXiv cs.AI
·
04 Sep 2026
This paper is an autopsy of prompt-based feature extraction. It tests whether changing job roles, prompt structure, and rule interpretation makes LLM-generated chemical features more useful for toxicity prediction, then passes those feat...
arXiv cs.AI
·
04 Sep 2026
The paper converts the physical laboratory into a machine-readable execution environment: typed objects, bounded capabilities, formal workflow composition, state simulation, and precondition checks. Its real function is to turn experimen...
arXiv cs.AI
·
04 Sep 2026
KC-Bench turns agentic unreliability into a measurable engineering problem: conflicting instructions, stale model knowledge, inconsistent inputs, and competing temporal sources. Its 238 curated tasks, simulated environments, tool calls, ...
arXiv cs.AI
·
04 Sep 2026
HalluPeer converts trust in scientific peer review into a measurable engineering problem: classify, locate, and detect unsupported claims. Its aligned paper-review pairs and injected hallucinations create infrastructure for auditing mach...
arXiv cs.AI
·
04 Sep 2026
GPS-Bench is building an evidence-grounded prediction and explanation layer for governance. Its real contribution is methodological: it turns policy simulation from unconstrained persona theater into a controlled contest between joint re...