cs.AI · 2026-09-29 · No. 128
Artificial Intelligence, 2026-09-29.
50 new papers in cs.AI. Titles, authors,
abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →
01 — The papers
50 entries-
01
FinAutoRubric: Expert-Guided Automatic Rubric Generation for Evaluating Financial Research Agents
Hoyoung Lee, Suyeol Yun, Jack Haverty, Yunju Cho, Meesong Kim, Daekyung Park, Sumin Kim, Jihoon Kwon, Jasmine Jia...
cs.AI · q-fin.CP
Evaluating finance research agents requires rubrics that reflect expert standards and fix the values correct as of an information cutoff. Expert-reviewed finance benchmarks rely on fixed, per-item rubrics, which are costly to extend and cannot encode each institution's own standard. In FinAutoRubric, experts specify reusable evaluation guidance, while agents and code carry out query-specific rubric generation, review, and validation. This...
-
02
Shockingly Simple Self-retrospection Improves Agentic Models Without RL
Jonathan Light, Christopher Zhang Cui, Jeonghye Kim, Roger Creus Castanyer, Emiliano Penaloza, Zhengyan Shi,...
cs.AI · cs.CL
People learn not only by repeating successful actions, but also by recounting and explaining their experiences, revising their understanding to guide future behavior. Can a language-model agent improve its future actions by training only on explanations of its own experience? We investigate this question by studying Retrospection-Only Fine-Tuning (ROFT), a minimal online procedure designed to isolate the effect of explanation-only training on...
-
03
Failure-Transparent Agents: Benchmarking Post-Failure Reporting in Tool-Using Language Models
Junru Zhu, Shiming Xie, Aime Lu Fan Chen, Xiaoqing Ding, Chunxin Tang, Ruoyu Qi, Yulang Fei
cs.AI
Tool-using agents can fail twice: a required tool can fail, and the agent can then report success without the evidence needed to justify it. Existing benchmarks often entangle this reporting failure with tool selection, recovery, and environment dynamics. We introduce Failure-Transparent Agents (FTA), a controlled benchmark that fixes the failed observation and required evidence state before generation, making post-failure claims directly...
-
04
Reinforcing Agentic Creativity in Scientific Ideation with Night Science
Priyanka Kargupta, Silviu Cucerzan, Shweti Mahajan, Allen Herring, Jiawei Han, Ryen W. White, Sujay Kumar Jauhar
cs.AI · cs.CL
Large language models (LLMs) excel at structured, verifiable tasks, but their low-entropy bias can produce homogeneous and predictable outputs, limiting their utility for open-ended scientific ideation. Effective discovery, however, spans a broader creative spectrum: from structured day science to loosely structured, serendipitous night science that reaches ideas beyond those typically considered. We introduce AI Night-Scientist, an agentic...
-
05
Reasoning with Continuous Latent Diffusion
Xiang Cheng
cs.AI
Continuous diffusion generates complete reasoning solutions through iterative refinement in latent space. We introduce Latent Flow Reasoning Models (LFRMs), an ELF-based training and inference recipe. Our experiments show that accurate decoding alone does not ensure strong reasoning performance. We therefore learn compact representations from multiple layers of a strong autoregressive teacher. Their decomposition also enables asynchronous...
-
06
Report: Progressive Disclosure of Agent Skills
Guilin Zhang, Kai Zhao, Priyanka Mudgal, Waleed Ammar, Xiquan Cui, Xu Chu, Alet Blanken
cs.AI
Users of Workday's deployed LLM-based agents often request features which can be addressed by defining named procedures, also known as skills, in the LLM context, effectively augmenting agents' capabilities. However, as an agent's skills library grows in size, so does the agent's operational cost. Progressive disclosure (lazy-loading) of skills as needed may reduce operational costs, but its impact on overall latency and skill-retrieval...
-
07
Verifier Errors in RLVR: Reward Hacking, Limits of Feedback, and Selective Control
Christian Moya, Elliott Thornley, Guang Lin
cs.AI
In reinforcement learning with verifiable rewards (RLVR), imperfect verifiers can reward incorrect responses, creating opportunities for reward hacking. Using gradient flow with a fixed verifier, we characterize the conditions under which reward rises while correctness falls. We then show that the observations available during RLVR are, in general, insufficient to detect or identify accepted errors, or to guarantee their reduction without...
-
08
PhoneCLI: From App Interfaces to Callable Commands for Mobile Agents
Yangqin Jiang, Lingrui Xu, Chao Huang
cs.AI
Mobile GUI agents operate through a perception--action loop: at each step they screenshot the device, invoke a vision--language model (VLM), and emit an action. It is slow, costly, and brittle, yet most of what it does is navigation---and everyday navigation is static, ordered, and endlessly repeated. We present PhoneCLI, which compiles an app's GUI navigation into callable commands, without any app-internal API, runtime instrumentation, or...
-
09
Not All Thinking is Created Equal: Latent Reasoning Discovers a Recurrent Search Algorithm for Depth Generalization
Huzi Cheng, Zhewei Zhang
cs.AI
Large Language Models can perform multi-step reasoning and improve task performance through different forms of intermediate computation, from token-based traces to computation carried out in latent space. However, a question remains open: do these different forms of thinking rely on the same underlying mechanism? To address this, we train and compare five variants of the same GPTNeoX backbone from scratch on an extended multi-hop reasoning...
-
10
Verifiable Visual Rewards Transfer from Synthetic Scenes to Natural Prompts
Shuyue Stella Li, Xiaochuang Han, Yulia Tsvetkov, Luke Zettlemoyer
cs.AI · cs.CV
Precise instruction following in image generation, such as satisfying object counts and spatial relations, remains an open challenge at least in part because it is learned using unreliable reward models such as object detectors and vision-language models. We introduce Verifiable Visual Rewards (VVR), the first framework for programmatically verifiable image rewards, and show that training on it generalizes to natural prompts. Each VVR task is...
-
11
RIDE: Reference-Anchored Inference-Time Diffusion Editing for Scaffold Hopping
Ruoxi Gao, Frazier N. Baker, Trieu Nguyen, Xia Ning
cs.AI · cs.LG
Scaffold hopping is a critical task in drug discovery, which seeks to discover new, structurally distinct molecules that share key functional groups and similar 3D shape with a reference binding ligand. Existing diffusion-based scaffold hopping methods formulate the problem as conditional generation of scaffolds given the functional groups. However, they lack a principled mechanism to jointly enforce 2D structural novelty and preserve the 3D...
-
12
From cacophony to hierarchy: a principled framework for assessing AI consciousness
Shamil Chandaria, Arvo Muñoz Morán, Fernando Rosas, Anil Seth, Henry Shevlin, Marcus Hutter, Thore Graepel, Adam...
cs.AI · cs.CY
The question of AI consciousness is one of the most urgent pre-emptive problems in philosophy and computer science, yet progress is hampered by a cacophony of competing theories that often talk past each other. Separating the hard problem from the mapping problem allows the deepest metaphysical disagreements to be set aside: granting that experience supervenes on a system's organisation, the tractable question becomes at which grain of...
-
13
TCSAlgBench: Benchmarking Automated Proving for Research-Level Theoretical Computer Science
Chutong Yang, Xiyuan Zhang, Yu Huang, Boran Han, Soonho Kong, Shuai Zhang, Vihang Prakash Patil, Zhen Han, Michael...
cs.AI · cs.CL
Large language models perform strongly on competition mathematics, but their research-level reasoning remains difficult to evaluate systematically. Theoretical computer science (TCS) connects algorithm design to explicit guarantees and fundamental limits, providing a setting for evaluating whether models can justify computational improvements with arguments humans can inspect. We introduce TCSAlgBench, a benchmark and reusable pipeline for...
-
14
Signatures of semantic search in the activations of large language models
Luke Leckie, Peter M. Todd, Jacob G. Foster
cs.AI
When recalling lists of concepts (e.g., animals) during the semantic fluency task (SFT), both humans and large language models (LLMs) organise their output into clusters of related items (e.g., sea animals) that are punctuated by strategic switches between clusters. In humans, this pattern can be explained by a semantic foraging process, whereby distinct neural and behavioural signatures accompany within-cluster production ("exploit") and...
-
15
Source-preserving alignment for robust evidence localization in scientific PDFS
Zihao Liu, Wei Yang, Zixiao Dong, Chenshu Li, Longzhang Liu, Tao Tan, Hong Xie
cs.AI
Scientific information-extraction systems often return a claim with an evidence string, which users must locate in the original PDF. This is challenging because the extracted evidence and PDF text layer are different representations: line wrapping, Unicode variants, superscripts, citation markers, and fragmented items alter text sequences and geometry. We present a source-preserving alignment framework: normalize text for robust matching...
-
16
IMC-CLINIC: Coupled Loss-Informed Newton Iterations for Clipping in Analog In-Memory Computing
Yung-Chin Chen, Chia-Yu Chen, Naveen Verma
cs.AI
Analog in-memory computing (IMC) offers a promising path toward energy-efficient large language model (LLM) inference by executing matrix multiplications (MatMul) directly within memory arrays in the analog domain. Its efficiency, however, comes with an additional source of error: limited-precision analog-to-digital converters (ADCs) quantize accumulated analog partial sums, introducing output-side error distinct from conventional activation...
-
17
Share-Borne AI Virus: Memory-Hopping Attacks Across LLM Agents
Sidharth Pulipaka, Ansh Sharma, Stanislau Hlebik, Leonidas Raghav, Vyas Raina, Ivaxi Sheth, Mario Fritz
cs.AI · cs.CL · cs.CR · cs.LG
Large language models are increasingly deployed as stateful assistants that retain information across interactions and use tools to read, modify, and create persistent artifacts. As these artifacts are shared between users, they form an indirect communication channel between otherwise independent assistants. We study a failure mode in which this channel enables self-propagating attacks. We introduce artifact-mediated propagation, where...
-
18
Representation Alignment as a Bottleneck in LLM-Based Retrosynthesis Planning
Hyunwoo Yoo, Cassie Huang, Haebin Shin, Li Zhang, Gail L. Rosen
cs.AI · cs.CL
While LLMs show promise in general reasoning, symbolic planning in chemistry remains a bottleneck. Direct ''SMILES-to-PDDL'' attempts fail because they force models to juggle chemical analysis and planning-language structuring simultaneously. We hypothesize that this failure stems from a lack of intermediate abstractions rather than insufficient model capacity. By decomposing retrosynthesis into molecule mapping, reaction mapping, and PDDL...
-
19
RSI-Master: Structuring Experiments to Guide Autonomous Model Improvement
Yaxin Du, Xiyuan Yang, Zhifan Zhou, Yujie Ge, Cheng Wang, Jiajun Wang, Sijie Chen, Zehui Liu, Yuxin Zhang, Weicheng...
cs.AI
Recursive self-improvement (RSI) seeks to enable AI systems to participate in improving their own capabilities. A concrete pathway is autonomous model development, where agents iteratively explore post-training strategies to improve a base model. This setting faces two challenges: agents may exploit open-ended experimental actions through hacking, and repeated experimentation may lead to strategy lock-in, where an early direction is refined...
-
20
From Search to Research: Exploring Search Scaling in Autonomous Quantitative Factor Mining
Kangcheng Deng, Hui Cai, Jiacheng Lu, Chester Zhongshu Qian, Rui Sun, Beidi Luan, Jing Li, Daxin Jiang, Zuo Bai
cs.AI
Inference scaling has been shown to improve large language model (LLM) performance, and this principle naturally extends to autonomous LLM agents through increased search budgets, which we refer to as *search scaling*. Although prior work has characterized the mechanisms, scaling behavior, and performance limits of LLM inference scaling, much less is known about these questions in autonomous research. Therefore, we investigate how search...
-
21
BaRe-Mem: Bayesian Reliability Memory for Robust and Adaptive Agent Consultation
Peilin Feng, Zhengyang Huang, Soujanya Poria
cs.AI
In multi-agent systems, reliable consultation is challenging because advisor capabilities vary across tasks, and misleading information can make consultation worse than autonomous reasoning. We introduce BaRe-Mem, an online Bayesian reliability memory for multi-agent consultation. It estimates advisor reliability based on the central model's internal belief representations and updates these estimates from historical interactions. These...
-
22
RareDx: Controlled Knowledge Integration and Graph-Grounded Policy Optimization for Rare-Disease Diagnosis
Bo Zhang, Yuchen Wang, Dongbai Li, Matthew Yu Heng Wong, Qingkai Zeng, Lijun Wang, Tien-Yin Wong, Peng Cui, Tianyu Liu
cs.AI
Rare-disease diagnosis is a long-tail reasoning problem: phenotypes are incomplete, individual disorders are sparsely documented, and relevant evidence is distributed across ontologies, gene annotations, and biomedical text. Language models consequently favor common conditions, miss rare candidates, or produce plausible but invalid names. We introduce RareDx, which couples controlled evidence use with knowledge-graph-grounded policy...
-
23
Continuous Context Management
William Hoy, Jingxuan Fan, Nurcin Celik, Xu Pan
cs.AI
Long-horizon large language model (LLM) agents commonly retain their complete interaction history until compaction is triggered at a predefined threshold. We study Continuous Context Management (CCM), which performs compaction at every turn to prevent interaction history from accumulating in the active prompt. At each turn, a CCM agent emits an updated memory together with an environment action; its next prompt contains the original task,...
-
24
ARISE: Adapting to Evolving Capability Gaps in Agentic Reinforcement Learning
Kun Feng, Yuchen Fang, Yiyang Tan, Shuqi Gu, Yongxiang Zhao, Yu Liu, Xingyu Lu, Lintao Ma, Kan Ren
cs.AI
As a long-horizon agent improves through experience, previously observed weaknesses may recede while new limitations emerge, continually changing what it still needs to learn. Yet the learning process often remains tied to a static view of these needs: fixed behavioral criteria and training priorities can become misaligned with evolving agent capabilities, while sparse task-level feedback makes such misalignment more difficult to detect. Even...
-
25
MechBench: Can AI Scientific Agents Discover Mechanisms Beyond Phenomenal Laws?
Zihan Yu, Jiadong Zhang, Jialin Cheng, Jingtao Ding, Yong Li
cs.AI · cs.LG · cs.SC
Scientific discovery requires not only recovering mathematical laws that describe observable behavior, but also identifying the mechanisms that generate them. Existing benchmarks for symbolic regression and scientific agents primarily evaluate phenomenal-law recovery, leaving mechanism discovery largely untested. We introduce MechBench, a benchmark that explicitly separates these two capabilities. Each task is defined by a mechanistic model,...
-
26
SRHarness: A Harness for Agentic Symbolic Regression
Zihan Yu, Shixuan Zhou, Hao Huang, Jingtao Ding, Yong Li
cs.AI · cs.LG · cs.SC
Recent agentic symbolic regression approaches increasingly rely on large language models to analyze data, select scientific operations, and refine hypotheses over long search trajectories. In such systems, performance depends not only on the underlying model and search strategy, but also on the runtime infrastructure that supports scientific search. We introduce SRHarness, a domain-specific harness for agentic symbolic regression built around...
-
27
Why Deterministic PRM Guidance Underperforms in Discrete Diffusion Reasoning
Yan Zhan, Shaobo Liu, Zhijun Gao
cs.AI
Discrete diffusion language models (dLLMs) expose a denoised solution at every step, which makes process reward model (PRM) guidance look like a way to spend compute at test time. We show that once denoising, PRM scoring, and outcome reward model (ORM) scoring are charged in the same budget of forward passes, its deterministic form loses to a much simpler baseline. Our PRMs score intermediate denoising states and are trained on the...
-
28
A.D.A.M.O. (Agent for language-Driven Actions with Multimodal Observations): A Visual-Symbolic Framework for Virtual Humans
Alessandro Emmanuel Pecora, Stefano Calzolari, Francesco Strada, Andrea Bottino
cs.AI · cs.GR
Creating believable vh requires the coherent integration of perception, reasoning, and action mediated by language. A central challenge is to combine these components into a control loop grounded in interactive 3D environments. To this end, we present A.D.A.M.O. (Agent for language-Driven Actions with Multimodal Observations), a visual-symbolic framework for language-driven vh that leverages a pretrained vlm with tool calling to unify...
-
29
AutoBCI: Forecast-Guided Agentic Neural Architecture Discovery for EEG-Based Brain--Computer Interfaces
Muyun Jiang, Yi Ding, Wei Zhang, Jinbo Chen, Chenyu Liu, Zhenjie Yang, Yuxin Li, Jingyuan Chen, Yuhao Lu, Yong Li,...
cs.AI
EEG-based brain-computer interfaces support a broad range of applications, yet designing decoding architectures that perform well across diverse tasks remains challenging. We introduce AutoBCI, an agentic framework in which a Designer Agent and a Forecaster Agent support the discovery and selection of EEG decoding architectures across tasks. The Designer Agent performs Pool-Guided Architecture Discovery (PGAD), generating and refining...
-
30
Just Initialize: A Training-Free Initialization Component for Large-Scale Routing Optimization
Jiale Zhao, Sirui Mao, Zimu Chen, Wentao Yang, Zihan Wang, Xuefeng Huang, Junji Cheng, Liyuanjun Lai
cs.AI
Large-scale routing problems are difficult to solve efficiently as their search spaces grow rapidly with problem size. Existing approaches primarily improve the optimization procedure itself, often at increasing computational cost. We instead shift the focus to a useful initialization that can be refined into a high-quality solution with limited downstream refinement. We propose Just Initialize, a training-free and solver-agnostic...
-
31
Building Transformation Layers for Riemannian Neural Networks
Ziheng Chen
cs.AI · cs.LG
Recently, deep neural networks on manifold-valued representations have garnered significant attention across various machine learning applications. One recent focus is the generalization of Euclidean fully connected (FC) and convolutional layers to non-Euclidean geometries. However, previous approaches typically focus on a few selected manifolds and rely on specific properties of the target manifold. In contrast, this work proposes a...
-
32
Self-Adapting Group of Experts for Multi-Agent Reasoning
Mohammad Atif Quamar, Nurbek Tastan, Karthik Nandakumar, Junpei Komiyama
cs.AI · cs.CL · cs.MA
Multi-agent systems bring together language model agents with different roles to propose, review, and refine solutions. Each agent's response depends on its model's capabilities, the reasoning strategy defined by its system prompt, and the information in its input context. Existing frameworks often adapt communication by changing this context while leaving individual prompts fixed, even when a problem calls for different skills. We study...
-
33
Structural Alignment for Reliable Industrial AI: Bridging Physical Reality, Data, Models, and Human Intent
Lizhi Xiao, Sihong Wu, Victoria Xiao, Yiqiao Song, Chen Gu, Jianwei Ma, Xinming Wu, Aimé Fournier
cs.AI · physics.geo-ph
Artificial intelligence is increasingly deployed in critical industrial domains, including healthcare, energy grids, subsurface exploration, where failures can have severe consequences for human safety, system stability, and economic outcomes. Yet AI is still evaluated primarily through benchmark accuracy, a model-centric metric that fails to capture the structural complexity and risks of real-world deployment. We propose a framework that...
-
34
A decision-support system applied to Law: Reasoning and explainability of the decision
Jeremy Bouche-Pillon, Pascale Zarat{é}, Yannick Chevalier, Nathalie Aussenac-Gilles
cs.AI
The emergence of the digital transition brought an increasing need to control the processing of digital information, including in Law Enforcement Agencies (LEAs). At the EU level, in recent years, many regulations have emerged to control data processing and exchange. Texts other than the GDPR, such as the ''Law Enforcement Directive (LED)'', appeared to regulate specifically how Law Enforcement Agencies (LEAs) could process data. A formal...
-
35
Don't Inoculate Everything: Stratified Inoculation Prompting Narrows Backdoor Triggers and Preserves Desired Traits
Kajetan Dymkiewicz, Tim Farrelly, Adam Práda, Ishaan Panigrahi, Srishti Gureja, Helen Yannakoudakis, Robert Mullins,...
cs.AI
Supervised fine-tuning can teach language models undesired behaviours alongside desired ones. Inoculation prompting (IP) aims to limit unwanted generalisation by requesting the undesired behaviour during training and removing the request at inference. However, undesired behaviour can still appear under unrelated prompts. IP can also hinder learning of the desired behaviour. We address these limitations in settings where both behaviours...
-
36
Jailbreaks for Black-Box Uncertainty Quantification in Large Reasoning Models
Lucas Biechy, Cédric Eichler, Adrien Boiret, Nicolas Anciaux
cs.AI · cs.CL · cs.LG
While Large Reasoning Models (LRMs) excel at complex reasoning, alignment through reinforcement learning often induces systemic overconfidence. In production environments, where logits may be unavailable, robust black-box uncertainty quantification (UQ) is essential for trustworthiness and safety. Focusing on question-answering for LRMs, we show that existing black-box methods, such as paraphrase-based self-consistency and confidence...
-
37
Jev thinks "I don't know'', but doesn't say it: Introducing Sys1Cal-v1 Dataset for Probability Calibration
Riccardo Porcedda
cs.AI · cs.LO · eess.SY
The appearance of Jev marked the era of System One Models, foundation models that return structured decisions with probability distributions rather than text. Aside from low cost and great speed, Jev's central promise is that these probabilities are calibrated: such claim is not backed by any public test and available external benchmarks evaluate confidence calibration, not whether every returned option probability has the right numerical...
-
38
TMCS: Tool-Grounded Multi-Agent Reasoning for Compositional Chemical Problem Solving
Shengqin Wang, Jie Jin, Yu Cheng, Yihang Chen, Weilin Luo, Yuan Xie, Zhizhong Zhang
cs.AI · cs.CV
Despite the promise of Large Language Models (LLMs) in computational chemistry, rigorous combinatorial chemistry problems remain difficult because they require quantitatively constrained molecular modification, candidate validation, and systematic revision after failed attempts. Existing tool-augmented chemical agents demonstrate useful planning and tool use, but they rarely provide a unified loop for property-driven molecular optimization...
-
39
Hyper Algorithm Design Agent: Evolving Learnable Optimizer from Zero
Zipei Yu, Yue-Jiao Gong, Zeyuan Ma, Yuncheng Jiang, Zhiguang Cao
cs.AI
Meta-Black-Box Optimization (MetaBBO) is one of the highlights in the recent AI for Optimization trend. This paradigm's bi-level workflow leverages the learnable algorithm design policy at meta level to ensure the performance and generalization improvement on the low-level optimization task. While MetaBBO helps advance the performance lower bound of the resulted optimization system, it is currently handcrafted and customized case by case to...
-
40
Reliability Engineering for AI Systems: Challenges, Methods, and Directions
Rong Pan, Yili Hong, Min Xie
cs.AI · stat.AP
AI reliability concerns whether an AI system performs its intended function dependably over a stated period and under stated operating conditions, with stated evidence. As these systems become more autonomous, that function includes more than a correct output. Retrieval, memory, tool use, permissions, human oversight, and interactions among systems must operate consistently and safely, and, for generative systems, so must the reasoning...
-
41
Narrowing the Horizon: Quantifying Topic Saliency Shifts in Generative Monoculture
Oriane Peter, Elena Simperl, Kate Devlin
cs.AI
As Large Language Models (LLMs) become central to how we access and share information, they play an increasingly powerful role in shaping global knowledge. However, as these models evolve, their outputs risk converging into a \textit{generative monoculture}, where the diversity of perspectives they represent narrows over time. Studies at the model level often fail to pinpoint which specific topics or viewpoints are being marginalised or...
-
42
Training-Free Clinical Reasoning through Medical Ontologies and Cognitive Mapping: A Symbolic-Probabilistic Knowledge Graph Framework
Surajit Das
cs.AI
Most clinical prediction systems learn patient-variable-outcome associations; we investigate a training-free diagnostic paradigm mapping patient observations to explicit medical knowledge. CKG Reasoner integrates candidate-specific Evidence Feature Nodes, patient-reference matching, a bounded Information Gate, knowledge-weighted evidence accumulation, disease similarity, and decisive clinical rules. Missing-aware normalization and coverage...
-
43
AbGaze: Attentive Geometric Representation Learning for End-to-End Antibody Design
Jiashuo Wang, Siqi Fan, Yizhen Luo, Zaiqing Nie
cs.AI
Computational antibody design requires representations that capture the geometric patterns underlying antigen--antibody interactions, yet existing approaches often rely on scalar distances or surface-intrinsic features, leaving cross-molecular geometry largely implicit. We present AbGaze, an end-to-end antibody design framework based on attentive geometric representation learning, which encodes distance, spatial direction, and surface-normal...
-
44
EvoIn: Bridging Evolution and Internalization for Agent Fine-Tuning
Shihan Dou, Shaofan Liu, Zhonghang Lu, Jiahang Lin, Shichun Liu, Binghai Wang, Jiajie Jin, Guanting Dong, Tao Gui,...
cs.AI · cs.CL
Recent work has explored improving agents by jointly evolving their harnesses and models, but often takes a ''potpourri'' approach that bundles together new tools, new decision-making procedures, and model adaptation to the evolved harness under a single notion of agent improvement. In this paper, we instead investigate how agents can improve their decision-making procedures. In particular, we propose EvoIn, an agent fine-tuning framework...
-
45
The Argument and the Letterhead: Source-Position Coherence in AI Evaluation
Michele Loi
cs.AI
An argument can be surprising coming from a particular speaker without being a bad argument. Do AI evaluators keep these judgments apart? Two preregistered descriptive studies and a later Jev supplement collected 2,976 usable evaluations of six fixed texts about US AI policy, Germany's debt brake and Swiss nuclear energy. Each text was presented under several source attributions. The key comparison asks whether the gap between two sources...
-
46
Textual User Taste: Natural-Language User Context for Foundation-Model Recommender System at Scale
Ghazal Fazelnia, Paul Gigioli, Eliza Klyce, Sharon Zheng, Katie Zelvin, Ye Myat Thein, Anurag Deshpande, Seda...
cs.AI
Foundation model recommender systems require user context that can be consumed by large language models, reasoned over, and refined through natural-language interaction. Traditional behavioral embedding vectors remain highly effective for retrieval and ranking, but they are opaque to users and not natively expressed for language model workflows. We present Textual User Taste, a system that generates structured natural-language taste profiles...
-
47
Imprint Reader: From Weight-Update Readout to Behavioral Intervention
Guanxu Chen, Qihao Lin, Jing Shao
cs.AI
As language models take a growing role in AI development, a natural aspiration is for them to reflect on their own learning process, as humans do, and use that reflection to improve themselves. At the same time, these models have an advantage that human learners lack, since training leaves parameter-level traces that can, in principle, be inspected directly. However, current models cannot decode these traces into an explicit account of what...
-
48
Towards Reliable AI Data Scientists: Data Agents with Workflow Harnesses
Huachi Zhou, Yujing Zhang, Jiahe Du, Jiacheng Cai, Zijin Hong, Chuang Zhou, Zheng Yuan, Qinggang Zhang, Qing Li, Xiao Huang
cs.AI
Large language model agents are increasingly deployed for data-intensive work, yet reliable data analysis requires more than general-purpose reasoning and ad hoc tool augmentation. Data Agents, equipped with workflow harnesses, offer a promising paradigm for automating the end-to-end data science lifecycle. This paper examines Data Agents from a harness-centric perspective. First, we introduce a taxonomy of Data Agents and associated data...
-
49
EP-Mem: Elastic Privacy Memory for Social Relationship-Aware LLM Agents
Fengzhou Sun, Yuan Zhang, Xintong Yu, Jinyao Yan
cs.AI
Large language model (LLM) agents face critical privacy risks when acting as delegates in human-agent-human communication. To prevent such breaches, agents must understand users' social relationships and adhere to context-dependent social information disclosure boundaries. Current studies on agent memory privacy focus on instantaneous interactions, leaving the long-term relational disclosure problem unexplored. In this paper, we propose...
-
50
ASCT: Attentive Search over Counterfactual Trees for Credit Assignment in Agentic Reinforcement Learning
Yang Li, Jinhan Yang, hai liu, Di Wan, Xiyu Chen, Zongsi Xu, Tuo Zhou, Sheng Zhong, Sergey Volkov, Ye Luo, Hao Sun
cs.AI
Terminal utility evaluates a complete agentic workflow, but learning requires credit for the decisions within it. We introduce Attentive Search over Counterfactual Trees (ASCT), a framework that turns training-time multi-step search into local action credit. At actor-visited states, an auxiliary tree evaluates alternative legal actions from the same recoverable prefix. Its action-value table is centered by the frozen actor's probabilities and...
This edition is part of The Daily Abstract — cs.AI archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.
#D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.