cs.AI · 2026-09-16 · No. 115

Artificial Intelligence, 2026-09-16.

51 new papers in cs.AI. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

51 entries
  1. 01

    ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents

    Shuhan Xue, Jianyuan Zhong, Ziyuan Nan, Wenbin Li, Zhaochen Yu, Jinchao Ding, Qiang Gao, Pengyu Zhan, Yuntong Zhang,...

    cs.AI · cs.CL

    We introduce and release ScienceBuddy, an interactive scientific research workspace that brings continually improving scientific agents into researchers' everyday workflows. ScienceBuddy supports researchers in carrying out scientific tasks while transforming their requests, feedback, and execution evidence into tasks and evaluation rubrics for continual learning. At its core is recursive-in-recursive self-improvement, a paradigm that couples...

    arxiv.org/abs/2609.17523 · PDF

  2. 02

    Verifiable Social Reasoning for LLM Assistants

    Amir Taubenfeld, Zorik Gekhman, Avigail Grinstein-Dabush, Itay Laish, Ariel Goldstein, Marian Croak, Avinatan...

    cs.AI · cs.CL

    LLM assistants are widely used for daily social advice, yet evaluating their social reasoning in such consultation settings remains challenging since (i) it requires setups where the assistant learns about social situations from subjective user narratives, and (ii) social properties, such as others' intentions, typically lack verifiable ground truth. To address these challenges, we introduce Fuse, a multi-agent simulation framework for...

    arxiv.org/abs/2609.17496 · PDF

  3. 03

    LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence

    Xingxuan Zhang, Gang Ren, Hao Yuan, Hao Zou, Hongze Tan, Hui Wang, Jianhao Song, Jiansheng Li, Jiayao Zhang, Jinghan...

    cs.AI

    We introduce LimiX-2, a new model in the LimiX family, developed through model and data scaling guided by our previously established scaling laws. LimiX-2 adopts the Contextual Mechanism Networks (CMNs) paradigm and is pretrained with Context-Conditional Masked Modeling (CCMM). CMNs shifts the organizing principle of in-context learning from target-centric prediction to mechanism-oriented joint modeling. Rather than centering the network on...

    arxiv.org/abs/2609.17488 · PDF

  4. 04

    JustFit: 200K-Token LLM Serving on a 24 GiB Laptop with Just-in-Time State Management

    Yuhua Chen

    cs.AI · cs.PF

    Capable open-weight models make local coding and reasoning attractive, but their context and execution state strain laptop memory. We present JustFit, an MLX-based inference runtime that combines KVExec for compressed KV execution, PhaseSwap for component residency, and StateTrans for state-preserving serving transitions. These mechanisms fuse reconstruction and coordinate just-in-time materialization and release, independently of...

    arxiv.org/abs/2609.17475 · PDF

  5. 05

    FlashVector: Agent for Hierarchical Model Serving Stack Optimization

    Qi Wu, Lohan Lemire, Kai Meng, Zhongmou Cai, Raphael Bargues, Petr Zhitnikov, Zeyuan Cao, Yao Wang, Shujun Bian, Wei...

    cs.AI · cs.PF

    Model serving is one of the largest cost drivers in production recommender systems. Maximizing its throughput requires navigating a deeply layered hierarchy: GPU kernels, the ML framework computation graph, the model server, and on-demand feature processing -- each demanding specialized domain expertise. Such cross-layer expertise is inherently difficult to acquire, and does not scale with a workload that continuously grows and evolves,...

    arxiv.org/abs/2609.17391 · PDF

  6. 06

    Self-Emergence Agent Architecture:Behavior-Inertia HMM, Reflexive Metacognition,and Social-Contrastive Self-Modeling

    Xiaoyang Liu

    cs.AI

    Large language model (LLM) agents exhibit strong language-generation and problem-solving capabilities, yet suffer from three structural limitations: personality drift, non-evolutionary reflection, and the absence of a self-other boundary. Existing generative-agent simulations rely on static memory and fixed prompts, maintaining neither behavioral inertia nor endogenous self-evolution. We propose the Self-Emergence Agent Architecture (SEAA),...

    arxiv.org/abs/2609.17331 · PDF

  7. 07

    From Transient Prompts to Persistent Control: Scientific Poster Generation via Recursive Semantic-Geometric Contracts

    Runze Li, Yukun Zhao, Can Xu, Yucheng Shen, Shuaiqiang Wang, Jianmin Wu, Lingyong Yan, Dawei Yin

    cs.AI

    Scientific poster generation distills a multimodal paper into a single-page visual artifact, forcing strict trade-offs between informational coverage and readability under a fixed spatial budget. Existing methods pass plans as transient prompts and validate individual stages in isolation. This strategy causes requirements to drift across content and layout modules, and previous checks to be silently invalidated. We introduce PosterVisor, a...

    arxiv.org/abs/2609.17326 · PDF

  8. 08

    Intrinsic Motivation in Reinforcement Learning: A Research Agenda for Adaptive Self-Organisation

    Anatoly Belikov

    cs.AI

    Biological cells can be viewed as individual, interacting agents whose collective dynamics give rise to adaptive behaviour at multiple levels of organisation, from individual cells through tissues to whole multicellular organisms. In this perspective and tutorial article we discuss whether intrinsic rewards in artificial neural systems can support adaptation, functional specialisation and higher-level self-organisation without a shared...

    arxiv.org/abs/2609.17325 · PDF

  9. 09

    Extracting ontology-compliant knowledge from scientific text describing irradiated materials using large language models

    Marco Luca Sbodio, Marcos Martínez Galindo, Vanessa Lopez, Blanca Biel, Pablo Canca, Pedro Delgado, Jesús I....

    cs.AI

    The quest for new materials increasingly relies on predictive models and comprehensive simulations that span scales from atomic to macroscopic levels. However, essential data necessary for these models and simulations are often embedded in scientific literature as unstructured text, limiting reusability and posing challenges for researchers seeking to leverage existing knowledge effectively. While extracting structured data from unstructured...

    arxiv.org/abs/2609.17291 · PDF

  10. 10

    End-to-End Latency-Minimizing and Load-Balanced Request Scheduling for Edge LLM Inference in Agentic AI Services

    Zhen Li, Jun Cai, Haoran Gao, An Li, Tan Li

    cs.AI

    Large language model (LLM)-powered agentic AI services increasingly demand low-latency inference, motivating the deployment of LLMs across distributed edge servers. However, heterogeneous communication and computing capabilities, together with dynamically evolving inference states, make the edge server selection for each incoming request time-varying and tightly coupled across slots. In this paper, we investigate an online request scheduling...

    arxiv.org/abs/2609.17193 · PDF

  11. 11

    MOCC-R1: Reinforcing Reasoning-Response Consistency for Multimodal Counselor Response Generation

    Wenjie Zheng, Qiming Xie, Jianfei Yu, Rui Xia

    cs.AI

    Multimodal counselor response generation (MCRG) aims to generate an appropriate counselor response from multimodal dialogue histories. Progress is limited by two gaps: first, existing datasets rarely capture sustained, human-recorded counseling interactions conducted by qualified counselors; Second, existing methods do not explicitly optimize consistency between counseling reasoning and the generated response, potentially undermining the...

    arxiv.org/abs/2609.17180 · PDF

  12. 12

    FirmCORe: A Benchmark for Structured Reasoning about Inter-Firm Collaboration Opportunities

    Tian Du, Tiantong Wu, Yafei Wang, Mengyu Liu, Xingyan Chen, Mu Wang

    cs.AI

    Comprehensive structured data on inter-firm relationships is often scarce or inaccessible because many relationships are privately negotiated, selectively disclosed, and fragmented across proprietary databases. This scarcity hinders the discovery of collaboration opportunities, particularly for startups and small and medium-sized enterprises. Firm profiles are readily available, but collaboration potential cannot be inferred from business...

    arxiv.org/abs/2609.17128 · PDF

  13. 13

    Shared-Prefix KV Reuse Across Standard LoRA Adapters: Quality and Serving Tradeoffs

    Dushyant Rajput

    cs.AI · cs.CL

    A common small-model deployment runs one shared backbone with several LoRA specialists that answer over the same context. Serving them naively re-prefills that shared context once per specialist. We study a narrow, practical question: for already-trained standard LoRA adapters -- not adapters retrained for cache compatibility -- how much task quality is preserved if the backbone's prefill KV cache is computed once and reused across...

    arxiv.org/abs/2609.17109 · PDF

  14. 14

    Symbolic Separation: Grounding Deep Agents in Knowledge Graphs for Trustworthy Operational Data Analytics

    Baibek Davletiyarov, Junaid Ahmed Khan, Andrea Bartolini

    cs.AI

    Generative AI promises natural language access to the massive numerical telemetry of data centers and Industry 4.0 installations, yet text-to-query and tool-using agents stay unreliable: even frontier models answer little more than half of real-world database questions, and far fewer of the multi-step, operational ones, because the LLM must compose how heterogeneous sources relate and hallucinates the relations, not just the fields. We...

    arxiv.org/abs/2609.17107 · PDF

  15. 15

    Semi-Supervised Learning-Based Genetic Biomarkers Dataset for Multiple-Stage Hepatocellular Carcinoma Prediction

    Ahmed Ammar Kubba, Manar Abu Talib, Jibran Sualeh Muhammad, Ali Bou Nassif, Abdalla Sayed Mohamed, Darko Castven,...

    cs.AI

    Liver cancer is a complex disease responsible for a high number of deaths across the globe each year, making automated solutions for liver cancer classification urgent. The most common form of liver cancer is hepatocellular carcinoma (HCC), accounting for over 90% of liver cancer cases. There is a distinct lack of publicly available HCC datasets utilizing genomic data, which is necessary for training artificial intelligence (AI) models for...

    arxiv.org/abs/2609.17100 · PDF

  16. 16

    Scaling-Score Conformal Prediction for Multi-Target Regression

    Sylvain Rousseau, Soundouss Messoudi

    cs.AI

    Multi-target regression requires a model to simultaneously predict several related outputs. Conformal prediction provides distribution-free, finite-sample marginal coverage guarantees, but extending these to joint multi-dimensional regions in a model-agnostic, sample-efficient manner remains challenging: max-aggregation ignores scale differences, copula-based methods are only asymptotically valid, rectangular methods typically split the...

    arxiv.org/abs/2609.17091 · PDF

  17. 17

    Interactive Memory Learning for Long-Term Conversations

    Cai Ke, Jiangyue Yan, Han Zhang, Xin Liu, Zike Yuan, Yue Yu, Hui Wang, Ruifeng Xu

    cs.AI · cs.CL

    Recent advancements in large language models have significantly enhanced the capabilities of agents in modeling long-term conversations. Despite these successes, existing approaches typically adopt a static heuristic paradigm, where information is passively archived without adaptive memory valuation. Consequently, these methods fail to self-evolve or align their memory management with evolving user needs. To address this, we propose ICML...

    arxiv.org/abs/2609.17088 · PDF

  18. 18

    Sample-Conditioned Representation Selection for Audio Few-Shot Learning

    Fengrui Liu, Ningxin Shen, Yi Li, Yiwei Fu, Feng Liu, Jiangmeng Li

    cs.AI · cs.SD

    Few-shot audio classifiers may rely on foreground-background co-occurrences and fail when those correlations shift. On SpurAudio, the resulting representation shift is concentrated and class dependent: for ResNet12, the top 10 percent of channels explain 82.80 percent of the null-corrected shift contribution. We propose SAMPLESELECT, which predicts a fixed-budget feature mask independently for each input while keeping the encoder and source...

    arxiv.org/abs/2609.17076 · PDF

  19. 19

    Neuro-Symbolic Hierarchical Intention Anticipation in Human Behavior

    Farnaz Soleimani, Abdelghani Chibani, Yacine Amirat, Ghazaleh Khodabandelou

    cs.AI · cs.CV · cs.HC · cs.LG · cs.NE

    Assistive autonomous systems must anticipate human goals before an observed behavior is complete. This article formulates anticipation as goal inference from a partially observed multimodal episode together with structured prediction of the remaining behavior, rather than exact motor forecasting. A compact Hierarchical Planning Decoder (HPD) is attached to a frozen neuro-symbolic recognition encoder and predicts, at four ontological levels,...

    arxiv.org/abs/2609.17064 · PDF

  20. 20

    Sparse MLLM Anchors, Dense Adaptation: Breaking the Self-Referential Loop in Wild Test-Time Adaptation

    Zhenbin Wang, Lei Zhang, Lituan Wang, Yan Wang, Zhao Zhang, Wei Huang

    cs.AI

    Wild test-time adaptation (WTTA) updates a source model online under small test batches, concurrent distribution shifts, and time-varying class imbalance. Most WTTA methods derive their adaptation signals, including predictive uncertainty, sample reliability, and local feature geometry, from the model being adapted. When the source model is unreliable under shift, these signals can reinforce its own errors, forming a self-referential loop. We...

    arxiv.org/abs/2609.17040 · PDF

  21. 21

    SKIP: a Self-knowledge-guided Step-wise Preference Learning Framework for Concise Reasoning

    Qinhong Lin, Yuhao Zhang, Yinglun Feng, Zhongliang Yang, Linna Zhou

    cs.AI

    While Chain-of-Thought (CoT) reasoning has been proven to be effective, it often leads to overthinking, resulting in computational overhead, inference latency, and even degraded performance in large language models (LLMs). Existing concise reasoning frameworks significantly compromise accuracy while compressing the length of output. In this paper, we propose SKIP, a self-knowledge-guided step-wise preference learning framework. Starting with...

    arxiv.org/abs/2609.17019 · PDF

  22. 22

    ORDER: Task-Conditioned Routing for Retrieval-Augmented Generation

    Aurélien Pellet, Julien Perez, Marie Puren

    cs.AI

    Retrieval-Augmented Generation (RAG) pipelines typically rely on a fixed indexing and retrieval configuration determined at preprocessing time. This one-size-fits-all design is ill-suited to domain-expert settings, where heterogeneous queries require different chunking granularities, metadata constraints, and source-selection strategies. As a result, configurations that are effective for one family of queries often perform poorly for others....

    arxiv.org/abs/2609.17012 · PDF

  23. 23

    ThinkFlow: Self-Evolving Probabilistic Latent Memory for Lifelong Conversational Agents

    Cai Ke, Xin Liu, Han Zhang, Jiangyue Yan, Zike Yuan, Ling Deng, Yue Yu, Hui Wang, Ruifeng Xu

    cs.AI · cs.CL

    Lifelong conversational agents rely on memory systems to maintain deep, context-aware interactions with users. However, existing explicit textual memory pipelines suffer from a severe information bottleneck, often losing subtle behavioral patterns and emotional shifts. Furthermore, being typically static post-deployment, they cannot autonomously adapt to personal habits and preferences without manual feedback. Cognitive science, however,...

    arxiv.org/abs/2609.17010 · PDF

  24. 24

    FlexEE: Self-Speculative and KV-Compatible Early Exiting for Offloading-Aware LLM Inference

    Qihu Xie, Ziwei Li, Yi Kang

    cs.AI

    Large language model (LLM) inference is often constrained by both computation and memory, especially in offloading-based deployments where model weights are transferred across memory hierarchies during autoregressive decoding. In this setting, reducing the number of executed layers can lower per-token latency while also avoiding costly weight movement. Motivated by this observation, we present FlexEE, an early exiting framework for...

    arxiv.org/abs/2609.17008 · PDF

  25. 25

    Affect-Prototype Guided Fusion for Open-Vocabulary Incomplete Multi-modal Emotion Recognition

    Yichi Zhang, Shenyue Wang, Jing Luo, Chunyang Yu, Xinyu Yang

    cs.AI

    Open-vocabulary multimodal emotion recognition (OV-MER) aims to generate open natural-language emotion labels from multimodal affective cues. In real-world scenarios, however, complete and synchronized modal data are difficult to obtain due to limitations of acquisition devices and user privacy constraints. Existing OV-MER methods are largely designed for full-modal inputs, and fail to perform effective feature fusion under modal missing...

    arxiv.org/abs/2609.16962 · PDF

  26. 26

    AntennaFlow: A Generative Flow Model for Offset Correction in Phaseless Antenna Testing

    Yongzhi Li, Chongting Shen, Menglin Chen, Xun Jiang, Zhengpeng Wang

    cs.AI · cs.IT

    Near-field to far-field transformation is central to large-aperture antenna testing, yet two coupled challenges remain: costly phase acquisition at millimeter-wave bands and violations of the centering assumption under offset mounting. Existing methods address these issues separately, requiring either dense full-field data or offset vectors. We tackle both jointly by exploiting a key observation: amplitude fields under different offsets are...

    arxiv.org/abs/2609.16948 · PDF

  27. 27

    QART: A Quantum-Classical Hybrid Architecture for Long-Horizon Reasoning -- Exploring a Conditional Path toward Quantum Scaling

    Lehao Lin, Yuheng Cheng, Guolong Liu, Yao Li, Xuning Tan, Xiyuan Zhou, Ruixi Zou, Shi Wang, Huan Zhao, Wenxuan Liu,...

    cs.AI

    Long-horizon reasoning is vulnerable to early errors that compromise later decisions. We present QART, the Quantum-Augmented Reasoning Transformer, a quantum--classical hybrid architecture combining a backbone language model with quantum encoding, CIM-based QUBO optimization, and quantum decoding. Semantic information can come from hidden representations or model-generated text; detailed encoding and optimization procedures remain...

    arxiv.org/abs/2609.16887 · PDF

  28. 28

    Bridging Learned Visual Perception and Symbolic Belief-Space Planning

    Guy Azran, Michael Navat, Sarah Keren

    cs.AI · cs.RO

    In partially observable settings, agents must act without full knowledge of the world state and rely on uncertain state-estimation pipelines. Obtaining grounded and verifiable symbolic plans under such uncertainty remains a key challenge. Recent work has integrated Vision-Language Models (VLMs) to bridge perception and symbolic reasoning, following two main paradigms. The first, VLM-as-planner, maps images directly to action sequences, and...

    arxiv.org/abs/2609.16884 · PDF

  29. 29

    CoAdapt: An LLM-based Framework for Adaptive Collaborative Perception in IIoT Robotic Swarms

    Houssam Hajj Hassan, Antonia Maria Masucci, Lynda Zitoune, Salah-Eddine Elayoubi

    cs.AI · cs.RO

    Industrial IoT environments increasingly deploy autonomous mobile robots for tasks such as material handling, product assembly, or infrastructure inspection. In such deployments, collaborative perception enables robots to share LiDAR observations and collectively construct a richer model of their environment than an individual agent could produce alone. However, industrial environments are dynamic spaces where robot positions shift...

    arxiv.org/abs/2609.16852 · PDF

  30. 30

    Execution Flexibility in Automated Planning: A Comparative Evaluation of Deordering and Reordering Strategies

    Md. Monjurul Islam, Sabah Binte Noor, Fazlul Hasan Siddiqui, Gahangir Hossain

    cs.AI

    This study covers foundational concepts for enhancing plan-execution flexibility, including partial-order planning, the producer-consumer-threat formalism, and a range of deordering and reordering strategies. Creating a partial-order plan from a sequential one by removing unnecessary ordering constraints is a practical way to improve execution flexibility, and several methods have been proposed for this task. This study analyzes their...

    arxiv.org/abs/2609.16822 · PDF

  31. 31

    Can We Do Interpretable NLI with Graphs Based on Atomic Propositions?

    Younes Boufouss, Luc Pommeret, Thomas Gerald, Patrick Paroubek, Sophie Rosset

    cs.AI · cs.IR

    While Large Language Model (LLM)-based Natural Language Inference (NLI) systems achieve high accuracy, their decision-making processes lack auditable structures. This paper explores whether NLI can be performed using only interpretable, graph-based representations of evidence. We introduce a fully graph-based pipeline where the classifier never directly processes the input text. Instead, sentences are decomposed into atomic propositions,...

    arxiv.org/abs/2609.16814 · PDF

  32. 32

    Layers, Sinks, and Scaling: Adaptive Evidence Selection for Multimodal Large Language Models

    Zhenbin Wang, Lei Zhang, Lituan Wang, Wei Huang, Yan Wang, Zhenwei Zhang

    cs.AI

    Multimodal large language models (MLLMs) can answer knowledge-intensive visual questions by combining visual evidence from images with facts retrieved from external sources. However, MLLMs may overlook relevant evidence in both modalities, attending weakly to the textual sentences or visual regions needed for the correct answer. Recent efforts address this by highlighting retrieved text and marking visual regions before generation, but apply...

    arxiv.org/abs/2609.16795 · PDF

  33. 33

    Integrating the Analytic Hierarchy Process with Large Language Models for Transparent Multi-Criteria Decision-Making

    Han Zhiguang, Farah Benamara, Pascale Zaraté

    cs.AI

    LLMs are increasingly employed in a wide range of decision-making tasks. However, the opacity of their internal reasoning makes it difficult to validate or interpret their outputs, and the need for interpretability becomes especially critical in high-stakes settings. This study examines the decision-making capabilities of LLMs through the Analytic Hierarchy Process (AHP), a classical and widely used multicriteria decision-making framework. We...

    arxiv.org/abs/2609.16779 · PDF

  34. 34

    Coverage-Aware Virtual IMU Augmentation for Low-Resource Human Activity Recognition

    Jiayuan Gao, Yingwei Zhang, Ziyao Tang, Yuejia Ma, Yuanzhe Chen, Shuchao Song, Boshi Tang

    cs.AI

    IMU-based human activity recognition (HAR) enables continuous, privacy-friendly monitoring of daily activities using wearable sensors. However, building reliable HAR models that generalize across diverse users and real-world conditions requires large amounts of labeled IMU data, which are expensive and difficult to collect. Existing approaches mainly rely on augmentation or synthesis to expand available data, but indiscriminately adding...

    arxiv.org/abs/2609.16768 · PDF

  35. 35

    Turn-level Multiscale Density Ratio Estimation for LLM Agents

    Zishuo Zhao, Kai Chen, Ao Li, Yuan Liu

    cs.AI

    With the rapid development of Large language model (LLM), agent systems enhanced by LLMs show huge potential in being able to deal with complex tasks, especially involving multi-step thinking or interaction with tools. For applying LLM techniques with a well-designed agent paradigm, post-training of LLM in multiple agent scenarios is necessary to achieve better performance. Among the variable post-training techniques, alignment methods such...

    arxiv.org/abs/2609.16760 · PDF

  36. 36

    Beyond Episodic AI: Cognitive Field Networks for Biologically Inspired Persistent Cognition

    Byung Gyu Chae

    cs.AI

    Cognitive Field Theory (CFT) proposes that cognition arises from memory-dressed collective dynamics that generate a persistent macroscopic cognitive field. Here we develop a Cognitive Field Network (CFN), a recurrent Transformer in which the organized hidden field re-enters subsequent inference through \[ Φ_{n+1}=F_θ(X_{n+1},Φ_n). \] Rather than prescribing an explicit memory operation, the CFN allows new information to act on an already...

    arxiv.org/abs/2609.16752 · PDF

  37. 37

    LSREP: A Longitudinal State-Replay Protocol for Evaluating Conversational Memory, with ICE v2 as an Audited Local-First Architecture

    Deepesh Sonar

    cs.AI · cs.CL · cs.IR

    Conversational memory changes during use, so endpoint question answering alone cannot establish how a persistent state accumulates, ages, or incorporates revisions. We introduce LSREP, a Longitudinal State-Replay Evaluation Protocol combining ordered replay, explicit lifecycle schedules, repeated probes, evolving reference answers, and mechanism-fidelity checks. Its architectural case study is ICE v2, a local-first memory middleware with...

    arxiv.org/abs/2609.16730 · PDF

  38. 38

    VideoMM: Adaptive Macro-Micro Inference for Efficient Video MLLMs

    Haoyu Guo, Yuan Feng, Junlin Lv, Mingjun Xiao, S Kevin Zhou, Xike Xie

    cs.AI · cs.CL · cs.CV · cs.MM

    Scaling Multimodal Large Language Models (MLLMs) to long-form video understanding is bottlenecked by the explosion of visual tokens, which saturates context windows and incurs prohibitive costs. Current solutions predominantly rely on auxiliary models for token reduction but face a fundamental dilemma: lightweight encoder-driven approaches often overlook critical semantic information, whereas heavyweight MLLM-driven reduction negates the...

    arxiv.org/abs/2609.16722 · PDF

  39. 39

    little m: An AI Agent for Industrial Process Optimization

    Yongchao Ye, Xinyu He, Dutliff Boshoff, Way Kuo, Lishuai Li

    cs.AI

    Manufacturing consumes one third of global energy and still has significant room for improvement in terms of energy efficiency. Optimal process control is essential for this purpose. However, synthesizing mathematical optimization models from messy, real-world industrial specifications requires bridging unstructured natural language and spatial diagrams with rigorous mathematical syntax. This poses a profound challenge for general-purpose...

    arxiv.org/abs/2609.16680 · PDF

  40. 40

    AI for Games in the Foundation Model Era

    Meng Luo, Yanlin Li, Hao Li, Hongzhan Lin, Pengfei Zhou, Tianjie Ju, Ran Zhang, Yeying Jin, Mong-Li Lee, Wynne Hsu

    cs.AI

    Foundation models, alongside advances in learned game-world models, are reshaping AI across the game lifecycle. Beyond playing games, recent systems model players and game dynamics, support design and development, adapt player-facing experiences at runtime, and evaluate resulting artifacts. Yet these directions have evolved largely separately, obscuring which capabilities transfer across settings and which remain tied to particular games,...

    arxiv.org/abs/2609.16679 · PDF

  41. 41

    ANIMASK: What the Model Contributes to Role Play in Simulated Story Worlds

    Xiucheng Zhang, Zhuoning Xu, Hanjun Luo, Yankai Chen, Hanan Salam, Xue Liu

    cs.AI

    When a language model plays a character, the observed behavior reflects both the assigned persona and the default dispositions of the actor model itself. Existing evaluations test persona fidelity or model defaults in isolation, but neither says, at a specific choice with consequences, what the persona changed and what the model's default kept. We introduce ANIMASK, a simulation framework that freezes books and scripts into story worlds whose...

    arxiv.org/abs/2609.16667 · PDF

  42. 42

    ReDraft, Don't Just Distill: Reference-Driven Revision for Continual VLLM Post-Training

    Zhihao Zhang, Mingqi Wu, Qiaole Dong, Enyu Zhou, Shuo Li, Boyang Liu, Jiazheng Zhang, Honglin Guo, Xin Guo, Shaofan...

    cs.AI

    Continual post-training of large multimodal models should add new capabilities while preserving those from pre-training, and the two goals pull in opposite directions. SFT gives explicit target supervision that learns a task from near-zero accuracy, but its off-policy targets move the model far enough to cause forgetting; on-policy methods such as RLVR and self-distillation preserve policy proximity yet supply little signal when the policy...

    arxiv.org/abs/2609.16639 · PDF

  43. 43

    EchoPath: Execution-Level Replayable Memory for GUI Agents

    Yao Zhao, Aditya Shanmugham, Swastik Roy, Yanxun Xu

    cs.AI

    Computer-use agents increasingly operate browsers, software, and desktop applications via CLI or API portals, but graphical user interface (GUI) still plays an important role in common industrial production scenarios. GUI agents commonly employ fresh observe-plan-ground-act loops, which is inefficient for enterprise tasks that repeatedly update records, process forms, configure tools, and export reports. We introduce EchoPath, a...

    arxiv.org/abs/2609.16635 · PDF

  44. 44

    A Framework for Generating Valid Context-Specific Benchmarks through Expert Guidance

    Kimberly Le Truong, Nari Johnson, Anna Kawakami, Hoda Heidari

    cs.AI · cs.CY

    This paper presents an end-to-end approach for generating context-specific large language model (LLM) benchmark datasets by combining expert input with synthetic data generation. Existing benchmark construction methods often trade off validity and scalability: datasets designed with domain experts can produce high-quality evaluations but are slow and costly to create, while synthetically generating data may scale efficiently but often results...

    arxiv.org/abs/2609.16592 · PDF

  45. 45

    Do LLMs Have Values? A Quantitative Analysis and Alignment Framework for Values in Large Language Models

    Keqing Zhang, Jingyu Chen, Yufan Liu, Yongqiang Zhu, Nai Ding, Lai Jiang, Congyan Lang, Bing Li, Weiming Hu

    cs.AI

    As Large Language Models (LLMs) increasingly handle complex subjective tasks, aligning their intentions and behaviors with human values has become a critical scientific challenge. However, current efforts are confounded by a striking behavioral paradox: they fluctuate unpredictably under minor wording changes ("swing"), yet stubbornly ignore explicit instructions to correct ingrained biases ("rigidity"). Resolving this duality is critical for...

    arxiv.org/abs/2609.16589 · PDF

  46. 46

    Query-Aware Source-Risk Triage for Retrieval-Augmented Generation

    Kainan Zhou, Gangzhen Qian, Chuhong Xu, Lu Yi

    cs.AI

    Retrieval-augmented generation (RAG) pipelines may omit a source's material relationship to the query. We study a pre-generation triage layer that treats this relationship as query dependent. The method routes canonical query families for enhanced review and assigns retrieved pages to pass, contextualize, exclude, or review. It combines a four-dimension page score, rank-discounted family aggregation, intent-preserving query mutations, and a...

    arxiv.org/abs/2609.16564 · PDF

  47. 47

    QueryFormer: Winning Solution for KDD Cup 2026 Tencent UniRec Challenge

    Yuanzhe Zhou, Zhaoyang Zeng

    cs.AI

    Post-click conversion rate (pCVR) prediction requires jointly modeling feature interactions and sequential user behaviors. The KDD Cup 2026 Tencent UniRec Challenge calls for a unified architecture addressing both. We observe that existing unified architectures often generate query tokens---the central information hub---with projection-based multi-layer perceptrons (MLPs), without explicit token-to-query attention for refining the query side....

    arxiv.org/abs/2609.16548 · PDF

  48. 48

    AquiLLM: Evaluating Faithfulness in Open-Weight RAG-LLM Systems for Scientific Research

    Bernie Boscoe, Srinath Saikrishnan, Vikram Seenivasan, Jack Stark, Andrew Lizarraga, Morgan Himes, Jonathan Soriano,...

    cs.AI

    Scientific research increasingly relies on large, heterogeneous data sources, motivating interest in retrieval-augmented generation (RAG) systems that provide natural language access to scientific knowledge and research workflows. Researchers are exploring the viability of these systems as natural language interfaces for document search and for generating analysis code and pipeline components. At the same time, concerns about data privacy and...

    arxiv.org/abs/2609.16519 · PDF

  49. 49

    From Manual Construction to AI-Driven Scenario Emergence: Rethinking Catastrophe Risk Modeling

    Hang Gao

    cs.AI · cs.LG

    Traditional catastrophe (CAT) risk models rely on costly manual construction to generate extreme weather scenarios, an approach largely unchanged since the 1990s. As climate extremes intensify, this creates mounting challenges to the entire risk transfer chain. This study proposes the TAISE framework, which repurposes AI weather forecasting models to produce coherent extreme weather sequences at a fraction of traditional costs. Through...

    arxiv.org/abs/2609.16493 · PDF

  50. 50

    Skill-based Agentic Evaluation for Real-time Data Science Tasks

    Aniruddha Tamhane, Raghavendra Addanki, Ayushi Aggarwal, Aditya Bansal, Rui Wang, Charles Menguy, Swati Jain

    cs.AI · cs.LG · cs.MA

    We present a framework for evaluating data-science agents on live, continuously updated data using executable ground truth and format-agnostic factoid scoring. Consider this example query: "what were last week's audience sizes"---the reference answer changes as the underlying data changes, so static references become outdated and standard LLM-as-a-judge pipelines cannot verify responses against a fixed ground truth. Our central contribution,...

    arxiv.org/abs/2609.16487 · PDF

  51. 51

    Fine-Tuning Fixes Mode Collapse and Over-Dispersion in LLMs

    Kirill Skobelev, Eric Fithian, X. Y. Han

    cs.AI

    Recent work by Doshi and Hauser (2024), Bisbee et al. (2024), and Xie et al. (2026) raises concerns that outputs from large language models (LLMs) tend to be under-diverse: they repeat or resemble one another more often than responses from the population they are meant to represent, a phenomenon known as mode collapse. In this work, we show that whether mode-collapse, or its opposite, occurs depends on the specific model and dataset used....

    arxiv.org/abs/2609.16454 · PDF

This edition is part of The Daily Abstract — cs.AI archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.