arXiv econ.GN
·
10 Sep 2026
This is a literature inventory framed as a policy-manageable product review. It traces BNPL from online fashion checkout to spending “anything ranging from a pizza to your rent,” then narrows the battlefield to consumer effects and credi...
arXiv cs.AI
·
10 Sep 2026
LexAgentHallu is not merely measuring hallucination. It is converting legal-agent failure into an engineering observability problem: classify the error, localize it in the trajectory, quantify it, then improve the system. Its decisive mo...
arXiv cs.AI
·
10 Sep 2026
The paper packages a narrow content-classification system as healthcare and mental-health protection. CareGuard combines zero-shot labeling, fine-tuned BERT-family models, emotion filtering, and cosine-similarity screening to identify al...
arXiv cs.AI
·
10 Sep 2026
The paper reallocates SFT’s gradient pressure. TrimSFT suppresses supervision for tokens that are already mastered or barely supported, concentrating it on an intermediate confidence band. Its real function is training-efficiency enginee...
arXiv cs.AI
·
10 Sep 2026
This is a control layer for an already-assumed autonomous diagnostic agent. It ranks error-prone states, designs stopping policies on disjoint development data, and applies finite-sample tests to selective error and autonomous coverage.
arXiv cs.AI
·
10 Sep 2026
PRAGMA turns lifelong personal history into an engineering pipeline: retrieve the right evidence, align it with memory, and convert it into recommendations, planning, and decision support. It correctly identifies that factual recall is n...
arXiv cs.AI
·
10 Sep 2026
This paper isolates a genuine weakness: LLMs can read local emotion but lose the evolving structure of relationships among multiple people. It then converts that weakness into a benchmark—191 samples, 7,079 annotated turns, 1,064.8 minut...
arXiv cs.AI
·
10 Sep 2026
This paper builds an inspection layer for systems already being forced into production. It converts agentic risk into a seven-domain taxonomy, automates 120 adversarial scenarios per domain through SAGE-RT, and uses LLM judges to classif...
arXiv cs.AI
·
10 Sep 2026
RobustSGPO turns agent-harness design into an executable search loop: propose an edit, validate the patch, test it, then continue from incumbent or retained variants. That is not merely prompt tuning. It is machine-assisted optimization ...
arXiv cs.AI
·
10 Sep 2026
The paper constructs a provenance taxonomy for Physical AI. Its seven sources—Recorded-Experience, Predictive-Modeling, Evaluative-Interaction, Surrogate-Environment, Mechanism-Grounded, Embodied-Coupling, and Evolution-Driven formation—...
arXiv cs.AI
·
10 Sep 2026
This paper is an engineering blueprint for moving AI from an advisory tool into the control layer of physical operations. Its decisive move is the transition from state synchronization to self-driven cognition: systems do not merely repr...
arXiv cs.AI
·
10 Sep 2026
This paper turns urban planning into an executable loop: inspect files, construct a plan, run an evaluator, and optimize against feedback. Its real function is not merely to improve planning software. It converts a professional cognitive...
arXiv cs.AI
·
10 Sep 2026
This is an engineering upgrade for automated graph reasoning. It decomposes heterogeneous graphs into community-specific agents, compresses structure into fixed-size permutation-invariant signatures, filters semantic evidence, and invoke...
arXiv cs.AI
·
10 Sep 2026
ContractEval identifies a real failure class in agentic automation: an LLM can produce a plausible answer while skipping the checks, branches, dependencies, ordering, or invariants that justify it.
arXiv cs.AI
·
10 Sep 2026
This paper builds instrumentation for an increasingly autonomous machine economy. LTD and ARP convert internal representations into forecasts of eventual task success, making agents easier to trust, gate, escalate, and deploy. It address...
GoogleAlerts/AI automation workers
·
16 Sep 2026
Wall Street Journal lands at 5/100 (lucid) for lucid. This claim is lucid and structurally aware. It identifies the pipeline problem—junior roles being automated across multiple professions—which is...
arXiv cs.AI
·
10 Sep 2026
This paper converts subjective XAI assessment into an LLM-mediated scoring pipeline. The advertised product is not merely better evaluation; it is scalable, reproducible automation of evaluative cognition. Human judgment becomes calibrat...
arXiv cs.AI
·
10 Sep 2026
Valerant targets a specific production bottleneck: manually discovering, reconstructing, and making 3D game spaces navigable. It couples a pretrained action-conditioned world model with SLAM and exploration-driven action selection, conve...
arXiv cs.AI
·
10 Sep 2026
This paper is not merely optimizing experiment selection. It is converting materials-recovery research into a closed decision loop: models choose the next wells, experiments update beliefs, and the system selects the process with the low...
The Cool Down
·
16 Sep 2026
Vince Gilligan lands at 28/100 (moderate) for lucid. Gilligan's quote explicitly names billionaires as the beneficiaries of AI at workers' expense and acknowledges AI displacement of human labor...