cs.AI · 2026-06-14 · No. 23

Artificial Intelligence, 2026-06-14.

54 new papers in cs.AI. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

54 entries
  1. 01

    Automated reproducibility assessments in the social and behavioral sciences using large language models

    Tobias Holtdirk, Pietro Marcolongo, Anna Steinberg Schulten, Felix Henninger, Stefan Rose, Sarah Ball, Bolei Ma,...

    cs.AI

    Reproducibility in the social and behavioral sciences is typically evaluated by independent researchers who reanalyze the original data to assess whether the published findings can be recovered. However, such approaches are resource-intensive and difficult to scale. Here, we show that large language models (LLMs) can automate reproducibility assessments. Using N=76 published studies with predefined claims from the behavioral and social...

    arxiv.org/abs/2606.13670 · PDF

  2. 02

    Agents-K1: Towards Agent-native Knowledge Orchestration

    Zongsheng Cao, Bihao Zhan, Jinxin Shi, Jiong Wang, Fangchen Yu, Zhijie Zhong, Zijie Guo, Tianshuo Peng, Zhuo Liu, Yi...

    cs.AI

    Current LLM-based research agents have advanced through agent orchestration, yet largely overlook scientific knowledge orchestration. Existing works often reduce papers to abstracts, surface mentions, and flat \texttt{cites} edges, omitting key entities, claims, evidence, mechanisms, and method lineages essential for scientific reasoning. To this end, we introduce \textbf{Agents-K1}, an end-to-end knowledge orchestration pipeline that...

    arxiv.org/abs/2606.13669 · PDF

  3. 03

    EurekAgent: Agent Environment Engineering is All You Need For Autonomous Scientific Discovery

    Amy Xin, Jiening Siow, Junjie Wang, Zijun Yao, Fanjin Zhang, Jian Song, Lei Hou, Juanzi Li

    cs.AI · cs.CL

    LLM-based agents have shown increasing potential in automating scientific discovery. Given an optimizable metric and an execution environment, they can propose, validate, and iterate scientific solutions, and have produced results that outperform human-designed approaches. As model capabilities continue to improve, we argue that the bottleneck for autonomous scientific discovery is shifting from prescribing agent workflows to designing agent...

    arxiv.org/abs/2606.13662 · PDF

  4. 04

    Before You Think: System 0, AI-Mediated Cognition and Cognitive Colonization

    Marianna Bergamaschi Ganapini, Massimo Chiriatti, Enrico Panai, Giuseppe Riva

    cs.AI

    This paper examines three recent frameworks for understanding the cognitive and epistemic consequences of artificial intelligence: Tri-System Theory, Thinkframes, and System 0. It argues that while the first two capture important dimensions of AI's influence on individual reasoning and collective epistemic practices, System 0 occupies a theoretically distinctive position that neither can fully replicate. The paper introduces the concept of...

    arxiv.org/abs/2606.13658 · PDF

  5. 05

    Beyond Runtime Enforcement: Shield Synthesis as Defensibility Analysis for Adversarial Networks

    Achraf Hsain, Sultan Almuhammadi

    cs.AI · cs.CR · cs.GT · cs.LG · cs.MA

    Shielded reinforcement learning is typically presented as a runtime safety mechanism that compiles temporal-logic specifications into automata restricting an agent's actions. We argue this is the wrong product. The same automata-theoretic machinery -- specification compilation, product game construction, attractor computation, and winning-region extraction -- is better read as a design-time analytical instrument whose outputs are structural...

    arxiv.org/abs/2606.13621 · PDF

  6. 06

    AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility

    Xiaoyuan Liu, Jianhong Tu, Yuqi Chen, Siyuan Xie, Sihan Ren, Tianneng Shi, Gal Gantar, Evan Sandoval, Donghyun Lee,...

    cs.AI · cs.LG

    Agent systems are advancing quickly across domains, but their evaluation remains fragmented. Most benchmarks rely on fixed, LLM-centric harnesses that require heavy integration, create test-production mismatch, and limit fair comparison across diverse agent designs. The root problem is the lack of an open, agent-agnostic assessment interface. We advocate Agentified Agent Assessment (AAA), where evaluation is performed by judge agents and all...

    arxiv.org/abs/2606.13608 · PDF

  7. 07

    Reasoning as Pattern Matching: Shared Mechanisms in Human and LLM Everyday Reasoning

    Zach Studdiford, Gary Lupyan

    cs.AI

    When large language models (LLMs) fail to generalize or make haphazard errors in reasoning, it is often taken as evidence that LLMs are not truly reasoning, but rather performing a kind of pattern matching. The implication is that people's behavior does not exhibit the same types of failures because human reasoning uses principled and abstract world models. We evaluate human participants and 25 LLMs on their ability to engage in common-sense...

    arxiv.org/abs/2606.13607 · PDF

  8. 08

    Multi-Agent Reinforcement Learning from Delayed Marketplace Feedback for Objective-Weight Adaptation in Three-Sided Dispatch

    Haochen Wu, Yi Hou, Shiguang Xie

    cs.AI · cs.LG · cs.MA

    Dispatch in three-sided marketplaces provides a natural setting for reinforcement learning from world feedback: decisions are evaluated by delayed operational outcomes such as delivery speed, courier utilization, and merchant congestion. We present a deployed reinforcement learning system at DoorDash that adapts dispatch objective weights in a large-scale food-delivery marketplace using delayed signals. Rather than replacing the combinatorial...

    arxiv.org/abs/2606.13604 · PDF

  9. 09

    EpiBench: Verifiable Evaluation of AI Agents on Epigenomics Analysis

    Harihara Muralidharan, Reema Baskar, Soo Hee Lee, Tim Proctor, Kenny Workman

    cs.AI

    We introduce EpiBench, a verifiable benchmark for short-horizon epigenomics analysis. EpiBench evaluates whether agents can make well-defined analysis decisions from realistic workflow states and return deterministically gradable answers. The benchmark includes 106 evaluations across CUT\&Tag/CUT\&RUN, ATAC-seq, ChIP-seq, and DNA methylation workflows. Across 5,088 valid trajectories from 16 model-harness pairs, no system passed a majority of...

    arxiv.org/abs/2606.13602 · PDF

  10. 10

    Reward Modeling for Multi-Agent Orchestration

    King Yeung Tsang, Zihao Zhao, Vishal Venkataramani, Haizhou Shi, Zixuan Ke, Semih Yavuz, Shafiq Joty, Hao Wang

    cs.AI · cs.CL · cs.LG · cs.MA

    Multi-Agent Systems (MAS) built on Large Language Models (LLMs) require effective orchestration to coordinate specialized agents, yet training such orchestrators is hindered by limited supervision and high computational cost. We propose Orchestration Reward Modeling (OrchRM), a self-supervised framework for evaluating orchestration quality without human annotations. OrchRM leverages intermediate artifacts from multi-agent executions to...

    arxiv.org/abs/2606.13598 · PDF

  11. 11

    Multiagent Protocols with Aggregated Confidence Signals

    Ali Elahi, Barbara Di Eugenio

    cs.AI · cs.LG · cs.MA

    Confidence is used for reliability, oversight, and a range of downstream decision tasks in Natural Language Processing (NLP), yet no existing method produces or evaluates a confidence for the output of a multiagent system. Prior work uses confidence within multiagent debate (MAD) to weight messages, trigger debate, or calibrate individual agents, but it never aggregates these into a single confidence for the system itself. We introduce three...

    arxiv.org/abs/2606.13591 · PDF

  12. 12

    A Three-Layer Framework for AI in Scientific Discovery

    Guojun Liao

    cs.AI

    Current discussions of AI in scientific discovery are often dominated by two visible capabilities: search over existing knowledge and execution through optimization, simulation, and automation. Both are important, but neither fully captures the central act of discovery: the formation and evolution of models. This paper proposes a three-layer view of AI in discovery. Layer 1 is search and retrieval by large language models. Layer 2, as the...

    arxiv.org/abs/2606.13566 · PDF

  13. 13

    Is It You or Your Environment? A Bayesian Inference Framework for Genomically-Anchored Personalized Physiological Interpretation

    Aruna Dey, Suraj Biswas

    cs.AI · cs.HC · q-bio.BM · q-bio.GN · q-bio.MN

    Personalized health AI systems face a fundamental cold-start problem: machine learning models for physiological interpretation require weeks of individual behavioral data before they can distinguish constitutional variation from environmentally driven deviation. We propose a solution grounded in causal inference and Bayesian prior design. An individual's genomic profile serves as an exogenous genetic anchor -- a domain-informed, personalized...

    arxiv.org/abs/2606.13556 · PDF

  14. 14

    Uncertainty-Aware Hybrid Retrieval for Long-Document RAG

    Hoin Jung, Xiaoqian Wang

    cs.AI · cs.CL

    Retrieval augmented generation (RAG) depends critically on the quality and granularity of retrieved evidence. Large retrieval units preserve context but often introduce irrelevant content, which can dilute answer bearing evidence and worsen long context utilization. Fine-grained units are more compact, but they may be difficult to retrieve reliably because short chunks can lack semantic, lexical, or bridging cues needed to match the query. We...

    arxiv.org/abs/2606.13550 · PDF

  15. 15

    CloudCons: A Comprehensive End-to-End Benchmark for Cloud Resource Consolidation

    Xiaobin Zhang, Lefei Shen, Mouxiang Chen, Zhuo Li, Hongkai Li, Han Fu, Jianling Sun, Xiaoxue Ren, Chenghao Liu

    cs.AI

    Driven by conservative over-provisioning to guarantee service reliability, resource utilization in cloud data centers remains at low levels. To mitigate this, the forecast-then-optimize paradigm has emerged to optimize consolidation by anticipating future demands. While emerging time series foundation models promise to enhance this paradigm through zero-shot generalization, existing benchmarks focus solely on prediction error metrics. The...

    arxiv.org/abs/2606.13513 · PDF

  16. 16

    Why Sampling Is Not Choosing: Intentionality, Agency, and Moral Responsibility in Large Language Models

    Joseph Keshet

    cs.AI · cs.CL

    Recent advances in large language models (LLMs) have prompted claims that such systems exhibit agency or qualify as moral agents. This paper argues that these attributions are misguided. We maintain that moral responsibility requires commitment-bearing agency grounded in intrinsic intentionality and self-attributed action, and that such agency constitutes the form of free will relevant to responsibility. Although LLMs generate coherent and...

    arxiv.org/abs/2606.13441 · PDF

  17. 17

    Evaluation Sovereignty in Metadata-Driven Classification: A Multi-Track Framework for Weakly Supervised Information Systems

    Raymond Vasquez

    cs.AI

    Evaluation in machine learning is typically treated as a neutral measurement process. However, in operational information systems, evaluation outcomes are often conditioned by the processes used to generate labels. This paper does not seek to improve classification performance. Instead, it examines the validity of performance measurement under differing label-authority regimes. This issue is particularly relevant in large-scale...

    arxiv.org/abs/2606.13436 · PDF

  18. 18

    Optimizing Appliance Scheduling for Solar Energy Management Using Metaheuristic Algorithms

    Hiba Ahmed, Alexander E. I. Brownlee, Jason Adair, Simon T. Powers

    cs.AI

    Renewable energy is essential for meeting future energy demands; however, solar energy generation, which occurs only during daylight hours often does not align with household consumption patterns. Appliances such as cookers, washing machines, and dryers are typically operated according to user preferred schedules rather than solar energy availability, creating a scheduling optimization problem. The objective is to determine optimal appliance...

    arxiv.org/abs/2606.13407 · PDF

  19. 19

    Neuro-Symbolic Agents for Regulated Process Automation: Challenges and Research Agenda

    Alexander Rombach, Chantale Lauer, Nijat Mehdiyev

    cs.AI · cs.MA

    LLM-based agents are entering regulated industries where they automate judgment intensive quality management processes. We argue that symbolic structures already embedded in these domains, including regulations, typed process models, and compliance constraints, should be treated not merely as external monitoring mechanisms but as core architectural components that shape the agent's decision-making and behavior. We propose...

    arxiv.org/abs/2606.13405 · PDF

  20. 20

    MiniMax Sparse Attention

    Xunhao Lai, Weiqi Xu, Yufeng Yang, Qiaorui Chen, Yang Xu, Lunbin Zeng, Xiaolong Li, Haohai Sun, Haichao Zhu, Vito...

    cs.AI

    Ultra-long-context capability is becoming indispensable for frontier LLMs: agentic workflows, repository-scale code reasoning, and persistent memory all require the model to jointly attend over hundreds of thousands to millions of tokens, yet the quadratic cost of softmax attention makes this untenable at deployment scale. We introduce MiniMax Sparse Attention (MSA), a blockwise sparse attention built upon Grouped Query Attention (GQA). A...

    arxiv.org/abs/2606.13392 · PDF

  21. 21

    A Quantitative Experimental Repeated Measures Study of Training Dynamics in a Small Llama Style Language Model Under a Compute-Aware Token Budget

    Joe Dwyer

    cs.AI

    This study examines training dynamics in a small Llama-style language model trained under a fixed, compute-constrained token budget. Rather than evaluating efficiency solely through endpoint performance, the study uses a quantitative experimental repeated measures design to analyze how validation loss, validation perplexity, rolling volatility, backslide behavior, spike behavior, and between-seed variability change across token-based training...

    arxiv.org/abs/2606.13370 · PDF

  22. 22

    IterCAD: An Iterative Multimodal Agent for Visually-Grounded CAD Generation and Editing

    Tao Hu, Jiaxin Ai, Licheng Wen, Xueheng Li, Shu Zou, Siqi Li, Nianchen Deng, Xinyu Cai, Hongbin Zhou, Pinlong Cai,...

    cs.AI · cs.CV

    Computer-Aided Design is pivotal in modern manufacturing, yet existing automated methods predominantly rely on open-loop, one-shot generation, creating a mismatch with iterative real-world practices. In this paper, we present IterCAD, a unified multimodal agent framework for closed-loop, interactive CAD generation and editing. We formulate the task as a multi-turn interaction between a multimodal agent and an executable CAD sandbox, covering...

    arxiv.org/abs/2606.13368 · PDF

  23. 23

    Can I Buy Your KV Cache?

    Luoyuan Zhang

    cs.AI · cs.CE · cs.MA

    Right now, across the world, AI agents are repeating the same absurd act: to read one document, they each recompute it from scratch. Every agent re-runs prefill, the most compute-intensive step a large model takes, over identical text, only to rebuild a key-value (KV) cache identical to the one the agent before it just built. The same answer, computed a million times. We make a proposal that is almost offensively simple: compute it once. Let...

    arxiv.org/abs/2606.13361 · PDF

  24. 24

    ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning

    Xucong Wang, Ziyu Ma, Yong Wang, Shidong Yang, Hailang Huang, Renda Li, Pengkun Wang, Xiangxiang Chu

    cs.AI

    Reinforcement Learning with Verifiable Rewards (RLVR) is a central technique for improving long-horizon reasoning in Large Language Models (LLMs). However, existing RLVR methods often encourage unnecessarily long reasoning rollouts, which can degrade reasoning coherence and exhaust the available context budget. Existing approaches to long-context organization often depend on external mechanisms to organize rollouts, rather than enabling the...

    arxiv.org/abs/2606.13316 · PDF

  25. 25

    Physics-Guided Spatiotemporal Learning for Coastal Wave Peak Period Estimation from Video

    Abubakar Hamisu Kamagata, Dharm Singh Jat, Attlee Munyaradzi Gamundani, Abhishek Srivastava, Paramasivam Saravanakumar

    cs.AI · cs.LG

    Wave parameters in the nearshore are crucial for coastal engineering, shoreline protection, marine hazard assessment, and coastal management for climate resilience. Traditional monitoring systems like buoys and radar platforms offer accurate monitoring but can have high installation and maintenance expenses and limited spatial coverage. Passive ocean monitoring using video has been achieved by leveraging deep learning, however, many methods...

    arxiv.org/abs/2606.13302 · PDF

  26. 26

    ERTS: Adversarial Robustness Testing of Ethical AI via Semantic Perturbation in a Bounded Consequence Space

    Pratyush Chaudhari

    cs.AI

    As AI systems are deployed in high-stakes ethical contexts such as healthcare triage, autonomous vehicle control, and employment screening, formal methods for evaluating their robustness against adversarial manipulation of ethical reasoning remain underdeveloped. This paper introduces the Ethical Robustness Testing System (ERTS), a closed-pipeline framework that: (1) encodes ethical dilemmas into a 22-dimensional Ethical Consequence Space...

    arxiv.org/abs/2606.13282 · PDF

  27. 27

    From Verdict to Process: Agentic Reinforcement Learning for Multi-Stage Fact Verification

    Rongxin Yang, Shenghong He, Siyuan Zhu, Chao Yu

    cs.AI

    Recent approaches combining Large Language Models (LLMs) with retrieval-augmented reasoning have shown promise for automated fact verification. To process complex claims, these verification pipelines typically execute multi-stage workflows that coordinate tightly coupled modules, including claim decomposition, evidence gathering, and verdict prediction. However, existing methods optimize individual stages in isolation or rely on fixed...

    arxiv.org/abs/2606.13262 · PDF

  28. 28

    MOSAIC: Modality-Specific Adaptation for Incremental Continual Learning in Parkinson's Disease Gait Assessment

    Minlin Zeng, Zhipeng Zhou, Yang Qiu, Martin J. McKeown, Zhiqi Shen

    cs.AI

    Gait-based Parkinson's disease assessment increasingly relies on heterogeneous sensors, but clinical systems rarely collect all modalities simultaneously. New sensors may arrive through device upgrades, protocol changes, or multi-center deployment, while historical patient data are often unavailable because of privacy and storage constraints. This modality-incremental setting faces three challenges: unreliable cross-modal distillation,...

    arxiv.org/abs/2606.13258 · PDF

  29. 29

    Multi-Field Hybrid Retrieval-Augmented Generation for Maritime Accident Root Cause Analysis

    Seongjin Kim, Sungil Kim

    cs.AI

    Maritime accident adjudication reports contain critical tribunal findings for root cause analysis (RCA), yet retrieving relevant precedents and drafting consistent reports from decades of records remains labor-intensive. This paper proposes a multi-field hybrid retrieval-augmented generation (RAG) framework for automated maritime RCA, utilizing a comprehensive dataset of 13,329 Korea Maritime Safety Tribunal (KMST) reports (1971-2025). We...

    arxiv.org/abs/2606.13249 · PDF

  30. 30

    EPIG: Emotion-Based Prompting for Personalised Image Generation

    Emna Othmen, Mohamed Yassine Landolsi, Lotfi Ben Romdhane

    cs.AI

    Text-to-image diffusion models have achieved impressive results in synthesizing high-quality images from natural language prompts. However, commonly used prompting strategies remain relatively generic, limiting the model's ability to accurately express emotional intent and nuanced affective attributes. This work proposes EPIG, a method that enhances emotional expressiveness at the prompt level prior to image generation. Grounded in...

    arxiv.org/abs/2606.13247 · PDF

  31. 31

    Brick: Spatial Capability Routing for the Mixture-of-Models (MoM) Paradigm

    Francesco Massa, Marco Cristofanilli

    cs.AI

    Defining query difficulty is one of the hardest problems in deployment engineering. Existing LLM routers rely on surface features such as domain labels, keywords, and token count, ignoring the within-domain variance that actually determines model success. Frontier models cost ten to one hundred times more than local open-weight models, so at production scale even small per-request savings become a direct cloud-bill lever. We present Brick, a...

    arxiv.org/abs/2606.13241 · PDF

  32. 32

    LLM-as-an-Investigator: Evidence-First Reasoning for Robust Interactive Problem Diagnosis

    Fabrizio Marozzo, Pietro Liò

    cs.AI · cs.CE · cs.ET · cs.LG · cs.MA

    Large language models (LLMs) are increasingly used as interactive assistants for technical problem solving. However, when users provide incomplete descriptions or plausible but unverified explanations, LLMs may prematurely align with these assumptions and propose solutions before collecting sufficient evidence. We refer to this behavior as user-driven sycophancy: the tendency of an LLM to reinforce a user-provided hypothesis instead of...

    arxiv.org/abs/2606.13220 · PDF

  33. 33

    Hallucination in Medical Imaging AI: A Cross-Modality Analytical Framework for Taxonomy, Detection, and Mitigation under Regulatory Constraints

    Omar Alshahrani, Muzammil Behzad

    cs.AI

    AI systems are being deployed across medical imaging faster than their failure modes are understood. At this point in time, the failure of greatest clinical concern is hallucination: clinically plausible but factually incorrect outputs, including fabricated anatomical structures, missed findings, incorrect laterality, and invented measurements in generated reports, with direct consequences, for example, for biopsy decisions, staging, and...

    arxiv.org/abs/2606.13211 · PDF

  34. 34

    A Minimal Model of Bounded Trade-Off Screening in Multi-Attribute Choice

    Manisha Dubey, Anirban Sarkar, Subramanian Ramamoorthy

    cs.AI

    Human decision-making often involves choosing between multi-attribute alternatives, yet classical models assume fully compensatory utility aggregation despite evidence that people reject options with poor performance on critical attributes. We propose a bounded trade-off reasoning framework in which decisions are governed by a screening process that evaluates the balance between gains and losses across attributes. The model introduces a...

    arxiv.org/abs/2606.13201 · PDF

  35. 35

    ARMOR-MAD: Adaptive Routing for Heterogeneous Multi-Agent Debate in Large Language Model Reasoning

    Fuqiang Niu, Bowen Zhang

    cs.AI

    Multi-agent debate (MAD) can improve large language model reasoning, but fixed debate pipelines often waste computation and can amplify correlated errors among similar agents. We propose ARMOR-MAD, a training-free heterogeneous MAD framework that treats debate as conditional computation. ARMOR-MAD combines three components: Pre-debate Agreement Routing (PAR) decides whether independently generated Round-0 answers require debate; Early...

    arxiv.org/abs/2606.13197 · PDF

  36. 36

    Under What Conditions Can a Machine Become Genuinely Creative?

    Yong Zeng

    cs.AI · cs.CY

    Recent AI systems can generate texts, software architectures, hypotheses, designs, and scientific workflows that appear creative. This paper asks under what conditions a machine can become genuinely creative, and how human agency can be preserved within shared cognitive and creative environments. It develops a requirement framework derived from Designics, the science of meaning-bearing intentional change. The paper argues that genuine machine...

    arxiv.org/abs/2606.13196 · PDF

  37. 37

    Reasoning for Mobile User Experience with Multimodal LLMs: Task, Benchmark, and Approach

    Ruichao Mao, Zhou Fang, Teng Guo, Hao Yang, Yaping Li, Shaohua Peng, Maji Huang, Xiaoyu Lin, Shuoyang Liu, Xuepeng...

    cs.AI

    User experience (UX) centered on usability, perceived consistency, and functional clarity is fundamental to real-world user interfaces (UI). The application of multimodal large language models (MLLMs) in the field of user interfaces is evolving rapidly, such as visual element grounding, graphical user interface (GUI) agents, and design-to-code generation. However, research efforts on evaluating UX based on UI screenshots are still immature....

    arxiv.org/abs/2606.13192 · PDF

  38. 38

    Mental-R1: Aligning LLM Reasoning for Mental Health Assessment

    Xin Wang, Boyan Gao, Yibo Yang, David A. Clifton

    cs.AI

    Mental health problems such as anxiety, depression, and suicide remain urgent global challenges, where timely and accurate assessment is critical for effective intervention. Recently, large language models have been explored for mental health assessment. However, existing general-purpose post-training methods do not align with the cognitive processes of human assessment, which may lead to unreliable reasoning outcomes. To bridge this gap, we...

    arxiv.org/abs/2606.13176 · PDF

  39. 39

    TerraBench: Can Agents Reason Over Heterogeneous Earth-System Data?

    Dat Tien Nguyen, Thao Nguyen, Fadillah Adamsyah Maani, Huy M. Le, Muhammad Umer Sheikh, Numan Saeed, Muhammad Haris...

    cs.AI

    Climate and environmental decision-making increasingly requires reasoning across heterogeneous inputs, including gridded physical data, satellite imagery, geospatial context, and simulator outputs. Weather and climate foundation models can forecast well, but do not reason interactively in language, while large language models (LLMs) reason in language but cannot operate directly on high-dimensional Earth-system data. As a result, real...

    arxiv.org/abs/2606.13148 · PDF

  40. 40

    Rethinking RAG in Long Videos: What to Retrieve and How to Use It?

    Yuho Lee, Jisu Shin, Nicole Hee-Yeon Kim, Jihwan Bang, Juntae Lee, Kyuwoong Hwang, Fatih Porikli, Hwanjun Song

    cs.AI

    Retrieval-augmented generation is moving beyond text into long, egocentric video, where systems must select query-relevant chunks across multiple modalities and temporal granularities. Yet progress in VideoRAG is limited by two gaps: existing benchmarks allow queries to be answered without the video, obscuring retrieval errors, and prior methods apply a single modality-granularity configuration per query, ignoring chunk-level variability. We...

    arxiv.org/abs/2606.13141 · PDF

  41. 41

    AAbAAC: An Annotated Corpus for Autoimmunity Information Extraction

    Fabien Maury, Solène Grosdidier, Maud de Dieuleveult, Adrien Coulet

    cs.AI

    Despite advances in information extraction driven by deep learning and large language models, performance gaps remain in highly specialized biomedical fields, where domainspecific complexity poses challenges for generalist models. In this work, we focus on the domain of autoimmunity, where the main entities of interest are autoimmune diseases, autoantibodies (i.e., molecules that may mark or cause these diseases), their molecular targets,...

    arxiv.org/abs/2606.13051 · PDF

  42. 42

    Augmentation techniques for video surveillance in the visible and thermal spectral range

    Vanessa Buhrmester, Ann-Kristin Grosselfinger, David Munch, Michael Arens

    cs.AI · cs.CV

    In intelligent video surveillance, cameras record image sequences during day and night. Commonly, this demands different sensors. To achieve a better performance it is not unusual to combine them. We focus on the case that a long-wave infrared camera records continuously and in addition to this, another camera records in the visible spectral range during daytime and an intelligent algorithm supervises the picked up imagery. More accurate, our...

    arxiv.org/abs/2606.13042 · PDF

  43. 43

    Nous: An Attempt to Extract and Inject the Cognition Behind Prediction-Market Behavior

    Haowei Qian

    cs.AI

    As LLM agents proliferate in prediction markets and collective decision-making, they risk a cognitive monoculture: agents built on shared foundation models produce correlated forecasts, and recent measurement finds frontier-model errors correlated at r ~ 0.77. We ask whether human cognitive diversity can be recovered from behavior and transferred to LLM agents. Nous extracts a structured eight-dimension behavioral profile from real Polymarket...

    arxiv.org/abs/2606.13038 · PDF

  44. 44

    SciR: A Controllable Benchmark for Scientific Reasoning in LLMs

    Pierre Beckmann, Marco Valentino, Andre Freitas

    cs.AI

    Three paradigmatic forms of inference recur across scientific reasoning: deduction, induction, and causal abduction. Reliably evaluating LLMs on these in scientific settings is currently out of reach: scientific benchmarks built on human annotations are costly and lack mechanistic ground truth, while synthetic logical-reasoning benchmarks do not resemble real scientific documents. We introduce SciR, a benchmark that combines multi-paradigm...

    arxiv.org/abs/2606.13020 · PDF

  45. 45

    Otters++: A Time-to-first-spike Based Energy Efficient Optical Spiking Transformer

    Zhanglu Yan, Jiayi Mao, Kaiwen Tang, Fanfan Li, Gang Pan, Tao Luo, Bowen Zhu, Qianhui Liu, Weng-Fai Wong

    cs.AI

    Spiking neural networks (SNNs) are promising for energy-efficient inference, and time-to-first-spike (TTFS) coding is especially attractive because each neuron fires at most once. In practice, however, this benefit is often reduced by the cost of computing a temporal decay term and multiplying it by the synaptic weight. We address this issue by turning a physical hardware "bug," the natural signal decay in optoelectronic devices, into the...

    arxiv.org/abs/2606.13016 · PDF

  46. 46

    The Illusion of Multi-Agent Advantage

    Prathyusha Jwalapuram, Hehai Lin, Chuyuan Li, Fangkai Jiao, Sudong Wang, Yifei Ming, Zixuan Ke, Chengwei Qin,...

    cs.AI · cs.CL · cs.MA

    Prevailing wisdom posits that Multi-Agent Systems (MAS) are superior to Single-Agent Systems (SAS), citing advantages like context protection, parallel processing and distributed decision-making. However, empirical support for this claim relies primarily on comparisons with SAS baselines using benchmarks that prioritize isolated reasoning tasks, which do not adequately assess these advantages. Focusing on automatically generated MAS that are...

    arxiv.org/abs/2606.13003 · PDF

  47. 47

    APCyc: Property-Informed Design of Cyclic Peptides via Automated Cyclization

    Yifan Zhao, Lang Qin, Jintai Chen

    cs.AI

    Cyclic peptides represent a promising class of therapeutic compounds in modern drug discovery, often offering improved stability and binding affinity. However, the de novo design of cyclic peptides remains challenging because methods must identify pocket-adaptive cyclization patterns and linkage sites while simultaneously controlling drug-relevant properties. This challenge is particularly pronounced for recent generative models trained...

    arxiv.org/abs/2606.12991 · PDF

  48. 48

    Structured Testbench Generation for LLM-Driven HDL Design and Verification-Oriented Data Curation

    En-Ming Huang, Yu-Hung Kao, Ren-Hao Deng, Wei-Po Hsin, Yao-Ting Hsieh, Cheng Liang, Hsiang-Yu Tsou, Mu-Chi Chen,...

    cs.AI

    Automated testbench generation has become a critical bottleneck in large language model (LLM)-driven Register Transfer Level (RTL) workflows, where large numbers of candidate designs must be verified rapidly and reliably. Existing prompt-based approaches treat testbench generation as unconstrained code synthesis, yielding stochastic outputs with high token cost, low reproducibility, and insufficient coverage. To address this gap, we present...

    arxiv.org/abs/2606.12983 · PDF

  49. 49

    A Mathematical Forum Platform for Collaborative Problem Solving and Dataset Generation for AI Reasoning

    Akbar Erkinov, Nurmukhammad Abdurasulov

    cs.AI

    Sharing mathematical content in online forums remains a significant friction point for students and educators: writing raw LATEX is error-prone, standalone optical character recognition tools require platform switching, and current forum software offers no integrated path from a photograph of a formula to a rendered post. We present a unified system that eliminates this friction by embedding an image to LATEX conversion pipeline directly...

    arxiv.org/abs/2606.12976 · PDF

  50. 50

    Multi-Modal Agents for Power Distribution Defect Detection: An Evaluation of Foundation Models

    Quan Quan

    cs.AI

    The power distribution network is critical to reliable electricity delivery, yet traditional inspection methods face limitations in semantic understanding, generalization, and closed-loop automation. To address these challenges, this paper proposes a Multi-Modal Agent framework specifically for power distribution defect detection. Central to this study is the systematic evaluation of multimodal foundation models as unified cognitive engines....

    arxiv.org/abs/2606.12969 · PDF

  51. 51

    OpenMedQ: Broad Open Pretraining for Medical Vision-Language Models

    Ibrahim Gulluk, Max Van Puyvelde, Olivier Gevaert

    cs.AI · cs.CV · cs.LG · eess.IV

    We present OpenMedQ, a medical vision-language model pretrained on the broadest fully-open medical mix to date: 14 datasets totaling ~3.35M pretraining samples spanning pathology, radiology, microscopy, and text-only clinical QA. OpenMedQ reaches state-of-the-art BLEU-1 on PathVQA (75.9), beating Med-PaLM M variants up to 562B parameters (~80x larger), and matches the best reported VQA-MED BLEU-1 (64.5). Its vision encoder, transferred to 8...

    arxiv.org/abs/2606.12953 · PDF

  52. 52

    Learning What to Remember: A Cognitively Grounded Multi-Factor Value Model for Agentic Memory

    Zhibao Chen, Qian Cheng

    cs.AI

    Long-running LLM agents accumulate interaction histories far larger than any context window, forcing a standing decision: what to encode deeply, what to forget, and what to retrieve under a fixed memory budget. Production systems answer with semantic similarity or recency -- both mis-specified for the forgetting decision, which is made at consolidation time before the future query is known. We propose a multi-factor memory value function...

    arxiv.org/abs/2606.12945 · PDF

  53. 53

    PRISMR: Overcoming Parse Collapse in Multimodal Listwise Ranking via Parameterized Representation Internalization

    Hao Jiang, Xin Li, Annan Wang, Zhi Yang, Haoxiang Zhang, Yichi Zhang, Weisi Lin

    cs.AI

    Generative listwise ranking with Large Multimodal Models (LMMs) aims to capture global list context in a single forward pass, but its effectiveness degrades in long-context multimodal scenarios. We identify a recurring failure mode, parse collapse, where the autoregressive decoder produces fluent yet incomplete rankings by silently omitting candidates and terminating early. This failure stems from limited context utilization rather than...

    arxiv.org/abs/2606.12942 · PDF

  54. 54

    MARS: Margin-Adversarial Risk-controlled Stopping for Parallel LLM Test-time Scaling

    Wenbo Chen, Puheng Li, Mengyang Liu, Weijie Su, Tianpei Xie

    cs.AI

    Parallel test-time scaling samples many reasoning traces and majority-votes their answers, improving LLM accuracy but requiring traces to run to completion, incurring substantial computational overhead. We observe that probing partial traces at intermediate checkpoints can extract current answers without disrupting generation, revealing an evolving aggregate vote. Based on this observation, we introduce MARS, a margin-adversarial stopping...

    arxiv.org/abs/2606.12935 · PDF

This edition is part of The Daily Abstract — cs.AI archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.