cs.AI · 2026-07-28 · No. 67

Artificial Intelligence, 2026-07-28.

50 new papers in cs.AI. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

50 entries
  1. 01

    ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams

    Ali Ansari, Yasmin Mohammadi, Farnoush Nili, Parsa Esmaeilkhani, Longin Jan Latecki, Eduard Dragut

    cs.AI · cs.CV · cs.DB

    Entity-Relationship Diagrams (ERDs) are central to conceptual database design, yet they are typically available only as rendered images rather than machine-readable schemas, limiting AI-assisted database engineering. We introduce ERUnderstand, the first large-scale benchmark for structured understanding of ER diagrams, comprising 2,960 diagrams collected from curated educational sources, real-world schemas, and synthetically generated...

    arxiv.org/abs/2607.24707 · PDF

  2. 02

    Eviction as Estimation: A Fixed-Lag Smoothing View of Test-Time Memory, and When Measuring Beats Accumulating

    Maruthi Vemula, Neeraj Praneeth Gajula

    cs.AI

    A language model with a bounded working memory must repeatedly decide which stored items to keep. Every deployed method decides the moment an item arrives, from the past (StreamingLLM, H2O) or from a guess about the future (SnapKV). We recast the choice as an estimation problem on a hidden signal, whether an item will be reused, placing existing methods on one axis, the commit lag $H$: online filters and learned predictors commit at $H=0$,...

    arxiv.org/abs/2607.24667 · PDF

  3. 03

    Reason-Mediated Behavioral Models for Auditing LLM Social Simulators

    Atharva Pandey, Gautam Jajoo

    cs.AI

    Large language models are increasingly used as social simulators, including as synthetic survey respondents. Most evaluations ask whether simulated outcomes resemble human outcomes. We argue that this is necessary but too weak: a simulator can match the final answer while using the wrong rationale-derived reason pattern. We study this problem through a 94-person sunscreen concept test in which each respondent evaluated three product concepts...

    arxiv.org/abs/2607.24649 · PDF

  4. 04

    Efficiency Matters in Autonomous Research

    Haiqian Yang, Yuan Cao

    cs.AI · cs.LG

    AI-driven autonomous research (AR) systems are becoming increasingly effective across a broad range of tasks. Their performance, however, is still evaluated primarily by the quality of the final outcome. In this paper, we argue that the efficiency of the solution-search process is an equally important but often overlooked dimension of performance. A strong AR system should not only produce high-quality results, but also reach them with as...

    arxiv.org/abs/2607.24647 · PDF

  5. 05

    Artificial Intelligence and Innovation Ecosystem: Evolutionary Developments, Challenges, and Future Directions

    Zhimin Zhang, Chengzhen Ma, Jia Chai, Rongxin Zhan, Huansheng Ning, Lingfeng Mao, Dan Zhang, Suiping Jiang

    cs.AI

    The development of the Innovative Ecosystem (IE) presents a new paradigm for economic integration, collaborative advancement, and shared achievements. The rise of Artificial Intelligence (AI) has significantly accelerated the global processes of digitization, informatization, and intelligence. Exploring how AI can leverage inherent characteristics to influence the development trajectory of IE is a topic that warrants further investigation....

    arxiv.org/abs/2607.24589 · PDF

  6. 06

    SIREN: Towards End-to-End Extreme-Weather Early Warning with Experience-Grounded LLM Agents

    Hang Ni, Weijia Zhang, Fan Liu, Mengqian Lu, Hao Liu

    cs.AI

    Early warning of extreme weather is essential for mitigating the societal, economic, and environmental risks posed by hazardous weather events. However, expert-centered warning workflows are costly, labor-intensive, and difficult to scale throughout the warning-to-action process. Although recent advances in Large Language Model (LLM) agents have enabled the automation of weather-related tasks, existing studies remain centered on isolated...

    arxiv.org/abs/2607.24588 · PDF

  7. 07

    LLM-SoccerArena: Benchmarking LLMs on Real-World Predictions in Sports

    Jonas Schröder, Jonas Schweisthal, Oliver Müller, Markus Weinmann, Stefan Feuerriegel

    cs.AI

    Large language models (LLMs) increasingly support decisions about uncertain future events, yet evaluating their ability to forecast real-world outcomes remains difficult. In particular, existing benchmarks are typically static and retrospective, and therefore cannot test how information is synthesized by LLMs to predict future events under uncertainty. We introduce LLM-SoccerArena (https://llm-soccerarena.com), a prospective live benchmark...

    arxiv.org/abs/2607.24573 · PDF

  8. 08

    DSCH-Loss: A Dynamic Semantic Channel Objective for Deep Semantic Hashing

    Tobias J. Bauer, Christian Riess, Daniel Loebenberger, Christian Bergler

    cs.AI · cs.CV · cs.IR

    Semantic hashing methods for generating short binary hash codes that allow efficient approximate nearest neighbor search in high-dimensional data spaces have gained extensive consideration in recent years. Deep learning-based methods offer better semantic capturing capabilities than traditional approaches relying on manual feature engineering. Moreover, they enable a data-driven approach to semantic hashing across diverse data modalities,...

    arxiv.org/abs/2607.24567 · PDF

  9. 09

    TRACE-CTI: Auditable Post-Extraction Governance of TTP Claims with Knowledge Graphs

    Federico Valletta, Giacomo Longo, Enrico Russo, Alessio Merlo

    cs.AI · cs.CR

    Security Operations Centers increasingly rely on automated mapping of Cyber Threat Intelligence reports to MITRE ATT&CK, yet extractor outputs remain fallible and are often stored without the evidence, provenance, and validation history needed to decide whether an individual mapping should be trusted. We present TRACE- CTI, a post-extraction claim-governance framework that preserves run-level Predictions, aggregates them into...

    arxiv.org/abs/2607.24563 · PDF

  10. 10

    Hierarchical Group-Conditional Conformal Risk Control for Selective Prediction in Language Models

    Murilo Salem, Luísa Böhm, Daniel Pontes, Anderson Ferrugem

    cs.AI

    Large language models serve heterogeneous populations structured by domain, topic difficulty, and linguistic style. Conformal risk control (CRC) gives rigorous marginal risk guarantees for selective prediction with abstention, but marginal guarantees do not imply per-group ones: a model can meet the population budget while systematically over-exposing subgroups to errors. Under mild shift in group composition, standard CRC violates the budget...

    arxiv.org/abs/2607.24562 · PDF

  11. 11

    LLM-Assisted Ontology Engineering and Construction of a French Legal Knowledge Graph

    G{é}nesis Montenegro, Mokhtar Boumedyen Billami, Catherine Faron, Fabien Gandon, Pierre Monnin

    cs.AI

    Maintenance regulations are complex legal texts that are difficult to exploit when addressing a specific case and challenging to integrate into operational systems. This paper presents a two-stage LLM-assisted workflow for French maintenance regulations: ontology engineering from a SEMLEG-based core ontology, followed by construction of an ontology-grounded French legal knowledge graph. The first stage consists in the open extraction of typed...

    arxiv.org/abs/2607.24551 · PDF

  12. 12

    Task-Conditional Faithfulness Auditing of Multimodal LLMs for Grid Diagnosis

    Tianqiao Zhao, Meng Yue, Jianhui Wang

    cs.AI · eess.SY

    Multimodal large language models (LLMs) can combine topology, measurements, and incident text for grid diagnosis, yet answer accuracy does not establish that task-appropriate evidence was used. This letter proposes a general framework in order to conduct task-conditional faithfulness audit. It compares self-reported reliance, intervention-derived behavioral reliance, and preregistered engineering importance. The framework first registers...

    arxiv.org/abs/2607.24539 · PDF

  13. 13

    Making Mathematical Knowledge Explainable, Accessible and Interoperable Through Large Language Model Integration

    Jan Range, Björn Schembera, Dominik Göddeke

    cs.AI · cs.DL

    Mathematical models are central to formalizing research problems, yet their documentation often falls short of FAIR principles. Knowledge bases such as the Mathematical Model Database (MathModDB) address this gap by providing curated, semantically rich representations of mathematical models. Built on Wikibase, the same open-source infrastructure underlying Wikidata, MathModDB utilizes Semantic Web technologies to support Linked Open Data,...

    arxiv.org/abs/2607.24512 · PDF

  14. 14

    From Execution to Capability: Scientific Experience Consolidation via Procedural Knowledge Synthesis

    Liwei Dong, Jiahao Zhao, Nan Xu

    cs.AI

    Large language models increasingly solve scientific-computing tasks, but executable feedback from one problem rarely becomes durable capability on subsequent problems. We study scientific-computing experience consolidation: converting verified runtime experience into transferable procedural knowledge and persistent model improvement. This setting presents two challenges: trajectory-derived artifacts may encode source-specific repairs rather...

    arxiv.org/abs/2607.24459 · PDF

  15. 15

    Failures Reveal What Metrics Miss: An Evidence-Driven Agent for Recursive Refinement of ECG Classifiers

    Jinliang Deng, Yiming Niu, Yibo Pan, Zhiqi Shao, Qin Luo, Yongxin Tong

    cs.AI

    Deep models have substantially advanced 12-lead ECG classification, yet their refinement still relies heavily on human experts to inspect failures and iteratively revise classifier designs. Recent LLM-based agents have demonstrated the potential for automated model design, but when guided only by aggregate performance metrics, they lack insight into why individual cases fail and how the classifier should be revised. We present RecursiveECG,...

    arxiv.org/abs/2607.24419 · PDF

  16. 16

    Are Prompt Optimizers Blind? Cross-Modal Visual Feedback for Automatic Prompt Optimization

    Haoyue Liu, Xiaoyu Ma, Ye Chen, Yuexian Zou, Xiaoying Tang

    cs.AI

    Automatic prompt optimization (APO) has been widely adopted to adapt vision-language models (VLMs) to downstream tasks without weight updates, yielding promising results. However, on multimodal tasks, the effectiveness of APO is fundamentally bottlenecked by a blind feedback channel: the optimizer reads the question, the prediction, and the gold answer, but never the input image on which the model failed, and therefore cannot diagnose...

    arxiv.org/abs/2607.24354 · PDF

  17. 17

    Simulating Tenant Responses to Energy Policy Interventions with Transaction-Cost-Aware LLM Age

    Weijie Xia, Stefanie Horian, Hanyue Huang, Queena K. Qian, Jie Yang, Pedro P. Vergara Barrios

    cs.AI

    Recent studies use Large language models (LLMs) to simulate human opinions and decisions by prompting models with demographic, attitudinal, or persona-based descriptions. Yet such simulations rarely model the practical, cognitive, or social frictions that shape how people respond to policy interventions. Perceived transaction cost (PTC) provides a useful lens for modeling the practical frictions that shape policy responses, such as...

    arxiv.org/abs/2607.24341 · PDF

  18. 18

    Gubernaut: A Deterministic Homeostatic Controller for Affect-Regulated LLM Agents, Validated Across Independent Model Families

    Dushyant Sharma

    cs.AI · cs.CL

    Large language model (LLM) agents inherit reactive failure modes: escalation under provocation, sycophantic drift under flattery, perseveration when stuck. These are failures of propensity, not capability; they concern what a model does under sustained pressure, which training-time alignment reduces but does not eliminate at runtime. This research led to the Gubernaut Cognitive Controller (GCC), a model-agnostic runtime control layer in a...

    arxiv.org/abs/2607.24339 · PDF

  19. 19

    Unequal Trips, Unequal Places: Diagnosing and Mitigating Delay Inequity in Autonomous Vehicle Fleet Coordination

    Nicole Hu, Mingtao Zhang, Haoyang LI, Chen Jason Zhang, Li Qing

    cs.AI

    City-scale autonomous vehicle fleet coordinators are typically optimized for aggregate travel time, yet fleet averages conceal how delay is distributed across trips and regions. We conduct a distributional audit on three real-city road-network and taxi-demand datasets from Manhattan, Chicago, and San Francisco. The audit reveals pervasive trip-length inequity whose direction depends on the city and coordinator. After accounting for trip...

    arxiv.org/abs/2607.24336 · PDF

  20. 20

    From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search

    Junlin Liu, Jiangwang Chen, Zixin Song, Shuaiyu Zhou, Chunji Lv, Hank Wu, Kailin Jiang, Jinyang Wu, Bohan Yu, Chenxi Zhou

    cs.AI

    Agentic search enables large language models to solve knowledge-intensive tasks by interleaving multi-step reasoning with retrieval, yet optimizing this with outcome-based reinforcement learning (RL) provides only sparse supervision. Knowledge distillation can supply denser guidance, and advanced proprietary models with their strong reasoning capabilities are promising teachers. While distilling from proprietary models can densify this...

    arxiv.org/abs/2607.24280 · PDF

  21. 21

    Generative Artificial Intelligence (GenAI) to convert images of queuing networks into verifiable simulation models: an open-weight LLM workflow approach

    Thomas Monks, Alison Harper, Amy Heather, Navonil Mustafee

    cs.AI

    Recent work has explored the use of Large Language Models (LLMs) to automate simulation model building, typically by generating executable code directly from natural language descriptions. However, this raises challenges for verification and reproducibility particularly for users without programming expertise. We propose Sketch2DES, a sketch-to-simulation workflow that converts diagrammatic representations of queuing networks into verifiable...

    arxiv.org/abs/2607.24259 · PDF

  22. 22

    Epistemic Norms for AI Safety and Alignment Research

    Keivan Navaie

    cs.AI

    Mainstream AI research emphasises capability growth and tolerates low failure rates when average-case performance is high. AI safety and alignment research has a different mission: to ensure that catastrophic failures never occur, under sparse evidence, adversarial dynamics, and fat-tailed risk. We argue that the two domains differ along two analytically independent axes---{\it capability profile}, demonstrating the absence of hazardous...

    arxiv.org/abs/2607.24243 · PDF

  23. 23

    Integrating Factual and Normative Industrial Knowledge via Constraint-Aware Graph Attention for Process Plan Recommendation

    Yuntong Chen, Yingqi Li, Yingying Xiao, Ziang Wang, Zewei Liu, Jiahao Liu, Xitian Tian, Lijiang Huang

    cs.AI · cs.IR

    Integrating heterogeneous industrial knowledge, including factual relations and decision constraints, remains a core challenge in industrial information systems. Machining process planning exemplifies this problem because engineers must select operations by combining material properties, feature characteristics, and quality requirements. Existing methods rely mainly on similarity retrieval or classification, without a unified ranking...

    arxiv.org/abs/2607.24213 · PDF

  24. 24

    Myopia Prevention and Control 3.0: Artificial Intelligence--Driven Risk Stratification, Proactive Monitoring, and Personalized Intervention

    Tieniu Wang, Cangzhu Huang, Qianhui Li

    cs.AI

    The convergence of artificial intelligence (AI), digital sensing, and ubiquitous computing has created an unprecedented opportunity to transform myopia prevention from a reactive, population-based model into a proactive, precision-driven one. Despite evidence that half the world's population will be myopic by 2050, conventional approaches---school-based vision screening (Phase 1.0) and evidence-based risk factor management (Phase 2.0)---have...

    arxiv.org/abs/2607.24187 · PDF

  25. 25

    Falsifiable Commitment Planning for Self-Correcting Web Agents

    Guangyi Liu, Huan Zhao, Quanming Yao

    cs.AI

    Long-horizon web agents often go off track before final failure: a trajectory can remain locally plausible even after the current state, reused skill, or plan assumption no longer supports the user instruction. Existing agents can plan, reflect, or reuse experience, but their plans rarely specify the evidence under which an active step should still be trusted. We propose FCPAgent, a falsifiable commitment planning framework for robust...

    arxiv.org/abs/2607.24167 · PDF

  26. 26

    Agent-UCT: Upper Confidence Bounds Applied to Trees for Agentic Workflow Optimization with Cost-Awareness

    Yang Li, Hai Liu, Dian Shao, Yu Wang, Xiyu Chen, Sergey Volkov, Bozhi Wang, Ziyu Sun, Sihang Liu, Ye Luo, Xiaowei Zhang

    cs.AI · cs.LG

    Optimizing agentic workflows, such as retrieval-augmented generation (RAG) pipelines, requires navigating a combinatorial space of discrete component choices under tight evaluation budgets. Existing approaches - heuristic search, black-box optimization, and standard tree search methods - do not explicitly exploit the compositional structure of these workflows, leading to redundant computation and inefficient budget allocation. We introduce...

    arxiv.org/abs/2607.24162 · PDF

  27. 27

    A Motion-Aware Vector Quantization Framework with Centroid Reuse for Efficient VLA Inference

    Zhuoran Song, Haozhe Jiang, Chunyu Qi, Minnan Pei, Gang Li, Xiaoyao Liang, Haibing Guan

    cs.AI

    Vision-Language-Action (VLA) models have demonstrated strong potential for embodied AI, yet their high inference latency on GPUs limits real-time deployment. Existing accelerators, such as Dadu-Corki, improve efficiency but treat VLA models as full-precision workloads, leaving substantial redundancy in both memory and computation underexploited. In this paper, we propose VQVLA, an algorithm-hardware co-design framework that accelerates VLA...

    arxiv.org/abs/2607.24148 · PDF

  28. 28

    Grading the Narrators: An Isnad-Rijal Framework for Claim-Level Provenance in Multi-Agent Knowledge Systems

    Ali Zahid Raja

    cs.AI · cs.MA

    Modern multi-agent knowledge systems increasingly accumulate knowledge through chains of autonomous transformations rather than direct retrieval. Existing provenance work records what happened - execution traces, tool calls, evidence links - and source-reliability estimation is long established (truth discovery, reputation systems). What is missing is an operational framework that attaches graded, per-domain transmitter reliability to...

    arxiv.org/abs/2607.24117 · PDF

  29. 29

    Scaling GUI Agents with Visual State Transitions

    Xiangyan Liu, Kaixin Li, Haonan Wang, Biao Wu, Meng Fang, Longxu Dou, Chao Du, Michael Qizhe Shieh, Tianyu Pang

    cs.AI

    We introduce State Transition Pretraining (STP) as a new scaling axis for GUI agents. During the STP stage, we continually pretrain a unified multimodal model on visual state transitions by jointly optimizing inverse dynamics (predicting actions from state changes) and forward dynamics (predicting next states from current states and actions). This optimization equips the model with better action-grounded visual representations and an internal...

    arxiv.org/abs/2607.24112 · PDF

  30. 30

    MemChain: Learning Interpretable Memory Traces for Memory-Augmented LLM Agents

    Yiwen Ma, Songjun Tu, Qichao Zhang, Dong Li, Linjing Li, Dongbin Zhao

    cs.AI

    Memory-augmented LLM agents typically answer queries by retrieving relevant memories and feeding them directly to an answer model. This retrieval-as-evidence paradigm assumes retrieved memories are already suitable for reasoning, leaving the answer model to resolve redundancy, conflicts, and weak relevance while incurring substantial context overhead in long-term memory tasks. We propose MemChain, a trainable post-retrieval memory policy that...

    arxiv.org/abs/2607.24097 · PDF

  31. 31

    Towards High-Level Semantic Intelligence

    Xiujie Song, Gefei Yang, Yining You, Jiahui Gan, Qi Jia, Shota Watanabe, Tianxi Wan, Mengyue Wu, Kai Yu

    cs.AI

    Recent advances in AI have substantially expanded its cognitive and reasoning capabilities. From the perspective of semantic complexity, the development of AI reveals a clear trajectory from simple to complex semantic processing. While early AI systems mainly addressed tasks involving direct and literal semantic perception or expression, contemporary systems are increasingly expected to perform more sophisticated cognitive reasoning, enabling...

    arxiv.org/abs/2607.24082 · PDF

  32. 32

    MiSS: A Logic-Driven Explanation of Minimal Sufficient Coalitions for Point Cloud Classifiers

    Mengda Xing, Jean-Marie Lagniez

    cs.AI

    We present MiSS, a black-box, query-based framework for explaining 3D point cloud classifiers through perturbation-relative sufficiency reasoning. MiSS treats a superpoint partition as an interpretable abstraction layer and asks whether the original prediction can be certified from a minimal coalition of geometric regions under a specified perturbation distribution. Unlike abductive explainers that require Boolean feature spaces or white-box...

    arxiv.org/abs/2607.24074 · PDF

  33. 33

    The Cost of Knowing: A Resource-Aware Protocol for Benchmarking Hallucination Beyond Static Leaderboards

    Keyu Li, Jin Gao, Dequan Wang

    cs.AI

    On standard factuality tasks, frontier models now cluster near the top of the scale. The question is therefore shifting from how factual a system is toward how much compute that factuality costs. Static leaderboards score factuality in isolation and treat compute as free, so they cannot tell a genuinely better system apart from one that simply spends more. Consider a ranking reversal. A brute-force Best-of-4 agent posts the higher raw...

    arxiv.org/abs/2607.24063 · PDF

  34. 34

    Success Is Not Self-Explanatory: Auditing Success Provenance in Agent Evaluation

    Jingkun Luo, Da-Tian Peng

    cs.AI

    A correct answer can conceal why an agent succeeded. Once agents change their information state during evaluation, correctness no longer distinguishes intended reasoning from answer acquisition. Outcome evidence and exposure detection do not establish whether success depended on an acquired target; we call this missing evaluation object success provenance. AcquaBench audits it through matched CLEAN, GOLD, and SHAM value substitution on four...

    arxiv.org/abs/2607.24054 · PDF

  35. 35

    Quantum-Inspired Evolutionary Neighborhood Search for Arrival-Departure Track Utilization Adjustment under Short-Term Disturbances

    Xiaobin Li, Wuming Lei, Yanbin Gao, Weiguang Wang

    cs.AI

    Short-term disturbances at major passenger railway stations alter train arrival and departure times as well as the release sequence of station resources. Effective recovery therefore requires coordinated adjustment of arrival-departure track allocation, station resource occupation, and train retiming. This study represents the station resources involved in train arrival, track occupancy, and departure operations as zone-level...

    arxiv.org/abs/2607.24049 · PDF

  36. 36

    The Half-Lives of Generative-AI Evidence: A 40-Record Audit, a Claim-Currency Framework, and a Reflexive Case of Frontier-Model-Assisted Research

    Carlo Iacono

    cs.AI

    Generative-AI evaluations can become historical before publication, yet calendar age does not affect every conclusion equally. This paper has two linked purposes. First, it audits a maximum-variation purposive corpus of 40 empirical records appearing between 18 July 2025 and 17 July 2026. The audit coded publication route, execution timing, model identity, age of the newest named generation or immutable snapshot, same-family supersession and...

    arxiv.org/abs/2607.24032 · PDF

  37. 37

    A Cyclic Adaptation-Generalization Framework with Uncertainty-Guided Self-Paced Learning for Long-Term Brain-Machine Interfaces

    Jiyu Wei, Di Hong, Zhanjie Zhang, Dazhong Rong, Qinming He, Yueming Wang

    cs.AI · cs.HC · cs.RO · eess.SP

    Brain-Machine Interfaces (BMIs), which link the brain to external devices, hold great potential in rehabilitation, human performance augmentation, and human-centered robotics. However, invasive BMIs face a critical challenge for long-term deployment due to neural drift, which degrades decoding performance over time and necessitates frequent recalibration. Existing methods designed to mitigate neural drift typically rely on either domain...

    arxiv.org/abs/2607.24031 · PDF

  38. 38

    Self-Supervised Consistency Enhanced Disentangled Learning for Neural Decoding Generalization in Brain-Machine Interface

    Jiyu Wei, Di Hong, Zhanjie Zhang, Dazhong Rong, Qinming He, Yueming Wang

    cs.AI · cs.HC · cs.RO · eess.SP

    Brain-Machine Interfaces (BMIs) provide a direct communication pathway between the brain and external devices, enabling humans to control assistive and robotic technologies, with potential applications in rehabilitation, human motor augmentation, and human-centered robotics. However, due to neural drift, the performance of BMIs decreases over time, posing challenges for long-term viability, particularly for invasive BMIs (iBMIs). Existing...

    arxiv.org/abs/2607.24023 · PDF

  39. 39

    Exploring Budgeted Image Classification with Content-Sensitive Resource Allocation

    Athanasios G. Papadopoulos

    cs.AI

    The ever-growing adoption of Artificial Intelligence (AI) creates the need to deploy Deep Neural Networks in a variety of computational environments. We consider dynamic environments, where computational requirements are subject to change, and we pose the following question: How do we adjust the complexity of an AI classification system, in order to maximize its accuracy, while meeting changing computational constraints? We call this problem...

    arxiv.org/abs/2607.23997 · PDF

  40. 40

    Plato-Bio: verification-first biological novelty screening with temporal rediscovery and structural benchmarks

    Stefan G. Creadore

    cs.AI · q-bio.QM

    Large language model research agents can connect literature retrieval, analysis code, and manuscript preparation, but coherent output does not establish scientific validity. We developed Plato-Bio, a biology-routed extension of the open Plato/Denario architecture that couples explicit workflow states with provenance records, citation checks, claim-to-evidence links, scoped file writes, and publication gates. A source audit identified and...

    arxiv.org/abs/2607.23975 · PDF

  41. 41

    Grokking on the Weight-Decay Clock: A Rate Hierarchy from Softly Broken Symmetries

    Taeyoung Kim

    cs.AI · cs.LG

    Delayed generalization, or grokking, remains poorly understood despite extensive empirical study. We identify an exactly solvable late-time relaxation mechanism for grokking in linear models trained with full-batch heavy-ball optimization and weight decay, together with a locally quadratic extension to nonlinear neural networks. Our analysis reveals a distinguished population-active component of the empirical null space, which we call the...

    arxiv.org/abs/2607.23967 · PDF

  42. 42

    EviBack: Search-Agent Reinforcement Learning via Evidence-Constrained Teacher Backoff

    Xiao Ma, Zhiquan Hu, Yi Wei, Chenchen Zhao, Yijun Chen, Jicheng Zhao, Yuming Li Chuang Dai

    cs.AI

    Reinforcement learning enables Agentic RAG systems to learn multi-turn search from verifiable outcome rewards, but all- zero rollout groups provide no comparative signal and may hide useful search behavior. We present EviBack, an evidence- constrained Teacher backoff that supplies auxiliary super- vision to such groups while preserving verifiable Actor re- wards. It separates evidence assessment from answer refine- ment, preventing reference...

    arxiv.org/abs/2607.23955 · PDF

  43. 43

    DICA: Dual-Indicator Guided Contrastive Alignment in Multimodal Large Language Models

    Hao Yang, Jin Wang, Xuejie Zhang

    cs.AI

    Human visual reasoning typically follows a coarse-to-fine attention process, starting from global scene understanding and gradually focusing on question-relevant regions. However, multimodal large language models may deviate from this pattern due to attention drift and the underutilization of visual evidence, which can lead to hallucinations. To mitigate these issues, this study proposes a Dual-Indicator Guided Contrastive Alignment (DICA),...

    arxiv.org/abs/2607.23944 · PDF

  44. 44

    From Cognitive Architectures to Language Agents: A Mechanism-Level Review of Lineage, Convergence, and Migration Gaps

    Haodi Fan, Zucong Lan

    cs.AI

    Memory, planning, reflection, and tool use are often compared as feature labels, obscuring the control semantics that determine how an agent actually runs. This review connects ten historical cognitive architectures, eight language-agent runtime families, and forty-two mechanism-focused modern systems. We reconstruct each mechanism through state, control, transition, persistence, failure, learning, and resource governance, then code evidence...

    arxiv.org/abs/2607.23942 · PDF

  45. 45

    MemTX: Transactional Belief Commit for Stateful Agent Memory

    Xiaoyang Li, Yiqi Wang, Haohui Lu, Zhi Chen, Mo Li, Pingan Song, Taotao Cai

    cs.AI

    LLM agents increasingly coordinate through persistent shared memory: one agent's write becomes another agent's premise, and eventually a tool call with real side effects. Current agent memory systems treat every accepted write as immediately actionable truth, so a polluted tool result, a stale update, or a teammate's half-finished note can silently drive an irreversible action. We argue that a memory write is not a belief commit. We present...

    arxiv.org/abs/2607.23929 · PDF

  46. 46

    Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory

    Saurabh Ranjan, Konstantina Sokratous, Brian Odegaard

    cs.AI · cs.CL · cs.CY · q-bio.NC

    A conversational AI that cannot tell its own output from what a user said will treat its own mistakes as user-provided facts. In humans, this capacity is called reality monitoring, and its failures are linked to hallucinations, delusions, and confabulation, yet whether LLMs possess it remains untested. Here we show, across two experiments and six LLMs, that source attribution depends on how conversational memory is structured: ceiling...

    arxiv.org/abs/2607.23927 · PDF

  47. 47

    GOTS: Greedy Orthogonal Token Selection for High-Resolution Vision-Language Models

    Jun Ling, Tao Huang, Junzhuo Liu, Bowen Tang, Peng Wang

    cs.AI

    Modern vision-language models (VLMs) increasingly rely on dynamic or high-resolution visual encoding, producing thousands of visual tokens that substantially increase downstream language-model inference cost. Existing token-reduction methods assess token utility through token-wise importance, query relevance, coverage, pairwise diversity, or subset-level objectives. Our key insight is to view visual token reduction through selected-span...

    arxiv.org/abs/2607.23913 · PDF

  48. 48

    Cost-Aware Recovery-Pathway Identification and Bayesian Optimization for Autonomous Materials Discovery

    Debajyoti Ray, Niranjan Srinivas

    cs.AI · cs.RO

    Autonomous laboratories automate experimental execution, but a campaign must also decide which recovery pathway merits optimization. We formulate this as a sequential decision problem with a discrete pathway-identification stage and a continuous within-pathway optimization stage under heterogeneous experimental costs. Our implementation, Coactive learning, combines a cost-sensitive Bayesian hypothesis-discrimination policy motivated by EC2...

    arxiv.org/abs/2607.23896 · PDF

  49. 49

    Understanding Human-like Solutions in Combinatorial Optimization via Learning and Search

    Haijiang Yan, Jian-Qiao Zhu, Liqiang Huang, Ming Meng

    cs.AI

    Humans often find good solutions to combinatorial optimization problems that are computationally hard even for advanced computer algorithms. In the Euclidean traveling salesman problems (TSP), people rapidly produce tours that are near-optimal, despite severe limits on time and computation. What makes a tour human-like, and how might such solutions be learned? Here we address these questions through a large-scale behavioral and computational...

    arxiv.org/abs/2607.23854 · PDF

  50. 50

    Do Visual Features Improve Other-Initiated Repair Detection? A Dyadic Multimodal Approach

    Anh Ngo, Nicolas Rollet, Catherine Pelachaud, Chloé Clavel

    cs.AI

    Other-initiated Self-repair, or in short Other-initiated Repair (OIR), is an essential mechanism in conversational interaction, whereby a recipient signals a problem in speaking, hearing, or understanding, prompting the previous speaker to resolve it. In the case of conversational agents, it is essential to accurately identify these repair initiation strategies to address communication breakdowns efficiently. While conversational analysis...

    arxiv.org/abs/2607.23845 · PDF

This edition is part of The Daily Abstract — cs.AI archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.