cs.AI · 2026-09-03 · No. 104

Artificial Intelligence, 2026-09-03.

46 new papers in cs.AI. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

46 entries
  1. 01

    Discriminative World Models for Web Agents

    Kelvin Li, Dhruv Pendharkar, Anish Pahilajani, Chuyi Shang, Leon Oks, Leonid Karlinsky, Rogerio Feris, Trevor...

    cs.AI · cs.LG

    Recent web agents use world models for test-time action selection by sampling candidate actions, predicting the resulting web states, and ranking them with a ranker model or a Process Reward Model (PRM). These world models are typically trained via supervised next-state prediction to generate fixed representations like HTML or AXTree snapshots. However, this objective is misaligned with the downstream ranker, which relies on predicted states...

    arxiv.org/abs/2609.02885 · PDF

  2. 02

    AI Contextual Measurement for Recovering Individual and Group-Level Effects: Validation Against Survey Measures and an Occupational Application

    Wenxin Jiang, Xuyang Wang, Yuxiao Wu

    cs.AI · cs.LG

    Researchers increasingly use artificial intelligence to construct measures of social, organizational, and occupational characteristics that are absent from conventional surveys. We propose AICOME, AI COntextual MEasurement, a framework for evaluating whether AI-derived respondent-level measures can recover individual and group-level effects in contextual models. The key idea is that an AI measure constructed at the respondent level can be...

    arxiv.org/abs/2609.02821 · PDF

  3. 03

    Large Language Models (LLMs) for Telecom Root Cause Analysis (RCA): A Structured Reasoning Framework for Evidence-Grounded Diagnosis

    Hao Zhou, Mandar Kulkarni, Hao Chen, Yan Xin, Charlie, Zhang

    cs.AI

    Root cause analysis (RCA) is a critical task in telecom network operations, but diagnosing performance degradations in modern 5G and emerging 6G networks remains challenging due to complex cross-layer dependencies. While large language models (LLMs) offer promising capabilities for reasoning and knowledge integration, directly applying vanilla LLMs to telecom RCA often leads to hallucination, unstable reasoning, and poor alignment with...

    arxiv.org/abs/2609.02805 · PDF

  4. 04

    SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment

    Qinghua Mao, Wanying Qu, Dadi Guo, Leitao Yuan, Qingyu Liu, Yu Li, Guanxu Chen, Yanwei Fu, Xi Lin, Xia Hu, Dongrui Liu

    cs.AI · cs.CR

    The performance of LLM-based agents is jointly shaped by the base model and the harness used when interacting with the environment. This exposes them to safety risks in both harmful final responses and multi-step execution trajectories. Existing safety alignment mechanisms often rely on either external harness updates or policy optimization, yet applying either paradigm in isolation fails to bridge runtime control with intrinsic safety. We...

    arxiv.org/abs/2609.02786 · PDF

  5. 05

    Measurement-Driven Sub-Network Selection for On-Premise Retrieval-Augmented Factory Agents

    Vasileios Rizeakos, Georgios Paisios, Alexandros Machairas, Michael Birbas, Athanasios Bachoumis

    cs.AI

    On-premise assistants can give factory workers conversational access to machine documentation, but models capable of the task rarely fit shop-floor hardware. We show that after structural compression and retrieval-grounded adaptation, model size is no longer a reliable predictor of adapted answer quality: general capability falls almost linearly with parameter count, while judged retrieval-augmented answer quality does not. We therefore treat...

    arxiv.org/abs/2609.02760 · PDF

  6. 06

    Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems

    Yihang Chen, Yuxiang Chen, Yuxuan Huang, Meng Fang, Weilin Luo, Jun Wang

    cs.AI

    Multi-agent LLM systems commonly use an orchestrator to decompose a task for a team of workers and then improve through textual reflection. Despite strong empirical results, these systems lack a unified account of coordination, memory improvement, and the role of external verification. We model orchestrator-worker interaction as a bilevel coordination game: under bounded coupling, the workers' local-update game is an approximate potential...

    arxiv.org/abs/2609.02750 · PDF

  7. 07

    Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills

    Jianlyu Chen, Yuyang Hu, Hongjin Qian, Jiawei Liu, Wenqing Wei, Xiaolong Chen, Defu Lian, Zhicheng Dou, Chaozhuo Li,...

    cs.AI · cs.CL

    Autonomous agents are beginning to carry out machine-learning (ML) research end to end. These agents combine a model backbone with a harness for planning, execution, memory, and verification, but this architecture still leaves domain-specific know-how outside the agent. We call this missing layer operational knowledge, the know-how that separates knowing a method from making it work. That knowledge is not absent from the field. It appears in...

    arxiv.org/abs/2609.02749 · PDF

  8. 08

    Door-in-the-Face Requests and Refusal Behaviour in Large Language Models

    Til Jordan

    cs.AI · cs.CL

    Does the door-in-the-face technique work on language models? In humans, a large request that is refused makes a smaller follow-up request more likely to be granted. We test this on nine production models from three providers: each model refuses a large request, then receives a smaller version of the same request, and we compare its compliance with asking directly. The answer depends on the model. On Anthropic's frontier models the technique...

    arxiv.org/abs/2609.02707 · PDF

  9. 09

    Loom: Weaving Diagnostic Strands into Free-Text Consensus via Embedding-Space Reweighting

    Ron Begleiter, Katya Egert Berg, Gilad Saban, Gil Shabat

    cs.AI · cs.CL · cs.LG

    Aggregating noisy, conflicting textual hypotheses into a reliable consensus is a fundamental challenge when deploying NLP systems in real-world industrial settings. While monolithic Large Language Model (LLM) agents offer unbounded expressivity for tasks like Root Cause Analysis (RCA), they suffer from context limits, compounding hallucinations, and prohibitive inference latency. Traditional weak supervision offers statistical rigor but is...

    arxiv.org/abs/2609.02649 · PDF

  10. 10

    Collective creativity in hybrid societies

    Mason Youngblood, Katie Mudd, Manuel Anglada-Tort, Cameron Jones, Elena Miu, Diana Omigie, Margaret Schedel

    cs.AI · cs.CY · cs.MA

    Generative AI is changing how cultural artifacts are created and circulated, and with it our understanding of creativity itself. Researchers disagree about whether these tools enrich or impoverish culture, and we argue that much of that disagreement comes from conflating two distinct components of creativity: novelty, a property of single artifacts, and diversity, a property of populations. We argue further that creativity in the context of...

    arxiv.org/abs/2609.02620 · PDF

  11. 11

    CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI

    Austin Tudor David Andrews, Liam Wilkinson, Jamie Heagerty, Harry Coppock, Jakob Nicolaus Foerster, Rui Ponte Costa

    cs.AI

    We present CivBench, an open-source benchmark for evaluating language model agents in long-horizon, tool-mediated environments through the Model Context Protocol (MCP). A single episode spans 300+ turns and produces thousands of tool calls over a large action space, requiring sustained planning, state monitoring, and execution under partial observability. The environment exposes 76 MCP tools and a narration layer that converts visual game...

    arxiv.org/abs/2609.02459 · PDF

  12. 12

    UTP-Bench: Uncertainty-aware Travel Planning Benchmark

    Etcharla Revanth Rao, Priyanshu Karmakar, Shubhojit Mallick, Manish Gupta, Shreya Ghosh, Abhik Jana

    cs.AI · cs.CL

    Large Language Models (LLMs) have recently demonstrated strong capabilities in automated travel itinerary generation. However, real- world travel planning is inherently uncertain: transportation delays, crowd fluctuations, and unexpected stochastic delays frequently inval- idate otherwise feasible schedules. Existing benchmarks like TravelPlanner and TripCraft assume deterministic environments, evaluating only static constraint satisfaction...

    arxiv.org/abs/2609.02421 · PDF

  13. 13

    Contrastive Explanations in Quantitative Bipolar Argumentation Frameworks

    Xiang Yin, Nico Potyka, Antonio Rago, Francesca Toni

    cs.AI

    Argumentation frameworks are useful tools for representing and reasoning with information in a variety of settings, e.g. in supplementing AI models as they perform classification tasks, with a notable benefit of providing additional explainability. In this paper, we introduce contrastive explanations for Quantitative Bipolar Argumentation Frameworks (QBAFs), one such formalism. Unlike most existing explanations for QBAFs, which explain the...

    arxiv.org/abs/2609.02399 · PDF

  14. 14

    Diagnosing with Insights: Structured Analysis of Agent Failures via Behavioral Abstractions

    Jiayi Bi, Yanjie Gao, Yuanmin Xie, Liqun Li, Tianyin Xu, Fan Yang, Mao Yang

    cs.AI

    With the proliferation of LLM agents, the ability to understand and diagnose failures in agents is essential to achieving superior effectiveness and trustworthiness. As agent failures often manifest via long and complex trajectories, manually finding the needles in the haystack is untenable. However, traditional diagnosis techniques for software bugs can hardly address LLM agent failures, while completely relying on LLMs as the judge yields...

    arxiv.org/abs/2609.02371 · PDF

  15. 15

    SALA: Semantic-Aware Logical Alignment for Complex Reasoning in In-Context Learning

    Zhao Ji, Wenqing Chen, Zhixuan Chu, Jianxing Yu, Jingping Liu, Shanhe Zhao, Zibin Zheng

    cs.AI · cs.CL

    Effective in-context learning (ICL) for complex reasoning relies on selecting the right demonstrations. Traditional retrieval methods based on surface similarity fail to capture the underlying problem-solving logic. Recent logic-based methods address this by matching predefined reasoning steps, but the rigid rules and exact-match criteria is improper to handle flexible or diverse reasoning processes. To address the problem, we propose SALA, a...

    arxiv.org/abs/2609.02336 · PDF

  16. 16

    Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds

    Axel Ahlqvist, Richard Guan, Juan-Pablo Rivera, Adeline Kassler, Dmitrii Troitskii, Alexandra Souly, Kai Fronsdal,...

    cs.AI · cs.CL · cs.LG

    A core obstacle to alignment evaluation is evaluation awareness: capable models can tell when they are being tested rather than deployed, weakening the conclusions a safety evaluation can support. We present two techniques that make simulated alignment evaluations harder to distinguish from real deployments. Our first technique, critique refinement, spends additional inference-time compute on each simulator action: the simulator generates...

    arxiv.org/abs/2609.02302 · PDF

  17. 17

    SCX Router: Streaming Zero-Shot Model Selection with a Decoder-KV Classifier and a Real-World Task Ontology

    Ihor Stepanov, Aleksandr Smechov, Mykhailo Shtopko, Dmytro Vodianytskyi, Oleksandr Lukashov

    cs.AI · cs.CL

    The rapid proliferation of large language models (LLMs) and the growing diversity of their applications presents a unique optimization opportunity: selecting the right model for the task, while optimizing for speed, cost, and quality at a per-task level. However, inference endpoints can vary widely in quality, price, latency, context support, tool use, domain expertise, and reasoning behavior. This heterogeneity makes manual heuristics...

    arxiv.org/abs/2609.02292 · PDF

  18. 18

    CoMerge: Conflict-Driven Preference Optimization for Multi-Task Model Merging

    Mingjie Zheng, Zihao Chen, Wenqing Chen, Weile Yuan, Zhixuan Chu, Jianxing Yu, Zibin Zheng

    cs.AI · cs.CL

    Model merging provides an efficient paradigm for constructing multi-task large language models (LLMs) without full model retraining, yet it remains challenged by parameter interference. While existing methods aim to preserve the capabilities of individual expert models and mitigate interference, they generally do not directly learn from the potentially degraded behaviors exposed by naive merging. In this paper, we propose a conflict-driven...

    arxiv.org/abs/2609.02273 · PDF

  19. 19

    Codebook Agent: Amortized Topology Design for LLM Multi-Agent Systems

    Jinxi Yu, Yubei Li, Eric Hanchen Jiang, Zhi Zhang, Dong Liu, Wenxiao Zhao, Levina Li, Kai-Wei Chang, Ying Nian Wu

    cs.AI · cs.LG · cs.MA

    Adapting the communication topology of an LLM multi-agent system to each query improves both accuracy and efficiency, yet current designers treat this as conditional graph generation: a variational, autoregressive, or diffusion decoder searches the $N \times N$ adjacency space, and a graph-network proxy trained on utility and a structural cost such as edge count ranks the sampled candidates. We argue that this formulation is misaligned with...

    arxiv.org/abs/2609.02264 · PDF

  20. 20

    APEx: Distillation of Agent Procedural Experience for Adaptive Deep Research Question Answering

    Jie Ding, Rui Sun, Xinyuan Zhang, Zeyu Zhang, Xin Liu

    cs.AI · cs.CL

    Deep research agents augment large language models with external tools to answer complex, long-horizon questions through multi-turn reasoning. Learning from prior experience is crucial for continual improvement, yet existing methods either retrieve verbose task-specific traces that burden decision-making, or distill procedural skills that remain decoupled from downstream policy adaptation. We propose APEx, a hierarchical experience...

    arxiv.org/abs/2609.02253 · PDF

  21. 21

    LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails

    Vansh Wahi

    cs.AI · cs.LG

    Self-improving agent pipelines have a problem at their center. An optimizer rewrites prompts to score higher, and the score comes from a judge that is itself an LLM. That judge has the last word on whether the system is getting better, and our position is that it has not earned it. The judge should be demoted from oracle to advisor: its verdict becomes one input among several, and every change is gated instead by a deterministic verification...

    arxiv.org/abs/2609.02246 · PDF

  22. 22

    Task-Level Natural Language Priors as Learning Signals for Low-Resource LLM Training

    Jian Gao, Xiao Zhang, Xun Zhu, Miao Li, Ji Wu

    cs.AI

    Large language models (LLMs) often struggle when low-resource training data are ambiguous or incomplete. Task-level natural-language priors can provide useful guidance in such settings, but existing approaches usually treat these priors as input context rather than as learning signals during training. We propose Prior-Guided Tuning (PGT), a training perspective that incorporates natural-language priors as auxiliary learning signals for...

    arxiv.org/abs/2609.02244 · PDF

  23. 23

    Propose to Learn, Learn to Propose: Evaluability-Aware Assistance under Bounded Rationality

    Yifan Zhu, Sammie Katt, Samuel Kaski

    cs.AI · cs.HC · cs.MA

    AI assistants often collaborate by proposing candidate edits, plans, or designs that users evaluate before adoption. Existing assistance methods focus on proposal quality or user-goal inference, often assuming that the user can reliably evaluate any proposal, which can fail in practice because of bounded rationality. We study evaluability-aware proposal planning, where proposals serve both as task interventions and as probes for learning...

    arxiv.org/abs/2609.02242 · PDF

  24. 24

    PGPO: Potential-Guided Policy Optimization for Multi-Turn Agentic Tasks

    Yuyao Zheng, Haipeng Sun, Junwei Bao, Lemao Liu, Hongfei Jiang, Yang Song, Dejing Dou

    cs.AI

    Group-based reinforcement learning (RL) has become an effective paradigm for LLM post-training, but in multi-turn agentic tasks with sparse terminal rewards, it often provides coarse credit for intermediate actions. To obtain more fine-grained credit assignment, recent work such as GiGPO introduces step-level advantages for intermediate actions. However, these step-level signals still rely on the final outcome of each individual trajectory....

    arxiv.org/abs/2609.02236 · PDF

  25. 25

    PhoenixNest-Video: Evidence-Grounded Multimodal Agent Framework for Automated Video Interview Assessment

    Fan Yuxuan, Huang Miaojun, Zhang Haimei, Wu Jingshen, Liu Hao

    cs.AI

    Interview assessment requires per-criterion judgments grounded in behavioral evidence, yet surging applicant volumes have made human-only evaluation costly and inconsistent, while existing AI approaches yield opaque scores without traceable rationale. We introduce PhoenixNest-Video, an evidence-grounded multimodal agent framework for automated video interview assessment. It builds a semantic video graph as structured working memory, performs...

    arxiv.org/abs/2609.02231 · PDF

  26. 26

    SkillGLoW: Procedural-Family Skill Consolidation for Self-Improving Agents on Long-Horizon Task Streams

    Ao Yan, Xin Zhang, Jiawei Du, Joey Tianyi Zhou

    cs.AI

    LLM agents increasingly self-improve by writing and reusing textual skills, kept either as one global document or as a flat pool of per-task entries, though most of the evidence comes from domains with structurally similar tasks. On long-horizon workloads where each task demands a different solution, the two forms fail in opposite ways: the document collapses into generic discipline, while the pool inflates and its entries stay bound to the...

    arxiv.org/abs/2609.02217 · PDF

  27. 27

    PEARL: Path-Entity Aligned Relational Learning with Contextual Subgraphs for Inductive Knowledge Graph Completion

    Yunchi Yang, Longlong Li, Cunquan Qu

    cs.AI

    Inductive knowledge graph completion (IKGC) aims to predict missing links involving entities unseen during training, requiring models to learn transferable relational and structural patterns. Existing subgraph- and path-based approaches often encode relational paths independently of their surrounding query subgraphs, although the predictive relevance of a path may vary across structural contexts. We propose PEARL, a Path-Entity Aligned...

    arxiv.org/abs/2609.02216 · PDF

  28. 28

    ASCII Attack: Recontextualising Harmful Requests as Artistic Critique in Large Language Models

    Da Cheng Gu, Yifei Dong, Xinghao Yang, Yongshun Gong, Wei Liu

    cs.AI

    Safety alignment trains large language models to refuse harmful requests stated plainly, but that training is applied mostly to surface form. Requests that only recontextualise the same operational content, changing how the model reads it, are therefore only weakly covered. The ASCII Attack is one such recontextualisation. It is single-turn and black-box: one message, with no access to model internals. It embeds a fully legible harmful...

    arxiv.org/abs/2609.02215 · PDF

  29. 29

    Examining the Vulnerability of Multi-Agent Medical Systems to Human Interventions for Clinical Reasoning

    Benjamin C Liu, Dillon Mehta, Rishi Malhotra, Adam Zobian, Yong Ying Tan, Samir Chopra, Daniella Rand, Natalie Pang,...

    cs.AI

    Human interventions at fault points can alter the diagnostic accuracy of multi-agent medical systems. We defined fault points as moments in AI agent conversations, in which an agent's reasoning became most vulnerable to external influence. Using the MedQA dataset, this study analyzed simulated doctor-patient conversations to measure how interventions shifted reasoning and accuracy. Correct intervention methods showed an improvement in...

    arxiv.org/abs/2609.02191 · PDF

  30. 30

    FUSE: An Evaluating Framework for Dangerous Capabilities of LLMs

    Zhengyi Jin, Ru Zhang, Xiao Chen, Xinbo Liu, Jiaxuan Lin, Jia Huang, Jianyi Liu, Zhen Yang

    cs.AI

    Fragmented safety evaluation undermines the governance of dangerous AI capabilities. We present a modular framework that evaluates each model through three orthogonal pipelines---Knowledge ($K$), Defense ($D$), and Harm ($H$)---under a unified protocol, aggregating results into a standardized dangerous-capability profile $φ$. Pluggable modules supply scenario seeds, knowledge banks, hazard queries, and judge rubrics, while the core evaluation...

    arxiv.org/abs/2609.02168 · PDF

  31. 31

    EmoStance: Response-Side Affective-Orientation Control for Empathetic Response Generation via Emoji Weak Supervision

    Ziyuan Jin, Yuxuan Ge, Zheng Tian

    cs.AI · cs.CL

    Empathetic response generation requires models to decide not only what to say, but also how to respond to the previous speaker's affective situation. We formulate this as response-side affective-orientation control and use multi-annotator emoji distributions as weak affective--attitudinal evidence, rather than as output symbols or gold labels, to induce a latent control space that operationally approximates listener stance. We construct...

    arxiv.org/abs/2609.02133 · PDF

  32. 32

    Beyond Context Windows: Persistent Discovery Context for Data-Centric Agents

    Jalal Mahmud

    cs.AI · cs.IR

    Data-centric agents repeatedly perform a discovery step before planning or execution: identifying the data objects relevant to a task. Yet successful discovery outcomes are typically discarded rather than reused. We introduce persistent discovery context, a lightweight memory layer that stores prior intent-to-object mappings and reuses them to augment future retrieval. Across three structured data environments, persistent discovery context...

    arxiv.org/abs/2609.02129 · PDF

  33. 33

    Semantic Signal-Assisted Inspection and Recovery Allocation in Reverse Logistics

    Jiani He, Dingyan Shang, Yihua Xu, Shiqi Huang, Yan Lyu, Jize Li, Shangjing Tang

    cs.AI

    Reverse-logistics operators often decide how to inspect and route returned assets before their condition is fully observed, while full inspection consumes scarce labor. Semantic Signal-Assisted Decision Support converts return notes into a condition factor and a signal-quality score that guide inspection depth and recovery allocation under shared labor capacity. We evaluate the framework in three synthetic benchmark scenarios spanning...

    arxiv.org/abs/2609.02116 · PDF

  34. 34

    READY or Not: Reliable Enterprise Agent Deployment

    Veronica Chatrath, Bryan Zhu, Jingxuan Fan, George Pu, Soham Dinesh Tiwari, Soham Dan, Ryan Young, Yuan, Li, Yuang...

    cs.AI

    An AI agent can perform well on benchmarks and still be unsuitable for deployment. Existing AI-agent benchmarks measure whether an agent can complete realistic professional work, whereas enterprise deployment asks a different question: whether an agent can meet a required reliability level, under acceptable human oversight, and at tolerable cost. We introduce Reliable Enterprise Agent Deployment (READY), a framework for qualifying AI agents...

    arxiv.org/abs/2609.02095 · PDF

  35. 35

    MASkills: Continual Skills Optimization for Multi-Agent LLM Systems

    Huaiyuan Yao, Xiaoou Liu, Charles Fleming, Tianlong Chen, Hua Wei

    cs.AI · cs.CL

    LLM-based multi-agent systems have shown strong performance on complex tasks, yet continual improvement from interaction experience remains challenging. Existing self-reflection methods build experience memories, but memories are mostly hard to invoke, refine, or scale, while agent skills offer a more actionable unit: structured procedural knowledge that specifies when to act, how to act, and which resources or tools to use. We introduce...

    arxiv.org/abs/2609.02094 · PDF

  36. 36

    Beyond Outcome Gaps: Process-Aware Fairness Diagnosis for LLM-based Multi-Agent Decision Systems

    Yiran Zhao, Lu Zhou, Liming Fang, Yufei Chen, Jiafei Wu, Zhe Liu, Xiaogang Xu

    cs.AI

    LLM-based multi-agent systems (MAS) are increasingly considered for high-stakes decision-making, yet outcome-based fairness audits can miss where risks arise within the decision trajectory. We present SCOPED-Hiring, a process-aware fairness diagnosis pipeline for LLM-based hiring MAS. SCOPED-Hiring constructs controlled resume variants, runs role-based hiring committees, logs over 311K structured decision trajectories, and converts trajectory...

    arxiv.org/abs/2609.02092 · PDF

  37. 37

    CHIME: Credit-Aware Hierarchical Memory Evolution for Long-Horizon Agentic Planning

    Yongshi Ye, Tian Lan, Feihu Jiang, Muyang Ye, Bin Zhu, Qianghuai Jia, Longyue Wang, Zhao Xu, Weihua Luo, Xiaodong Shi

    cs.AI

    Planning is a central capability that enables agents to decompose complex long-horizon tasks into manageable steps. Test-time search and training-based methods improve planning but incur high inference costs or require expensive training data. Self-evolving memory instead accumulates reusable experience from agent interaction outcomes into an external memory bank, so planning capability keeps improving at inference time without parameter...

    arxiv.org/abs/2609.02074 · PDF

  38. 38

    ToolGate: An Executable Acceptance Pipeline for Tool-Dependent Scientific Benchmark Construction

    Ke Zhang, Yankang Liu, Roya Zandi, Maziar Raissi

    cs.AI · cs.MS · cs.SE

    Scientific benchmarks are commonly built by domain experts who write tasks and cross-check one another's work, or who adapt existing material from textbooks, published papers, and online resources. These routes can produce strong evaluations, but they require substantial per-item labor. Language models can reduce this repeated work by proposing candidates quickly. The remaining problem is acceptance. We target scientific questions whose...

    arxiv.org/abs/2609.02067 · PDF

  39. 39

    MineTRACE: An Evidence-Grounded Interactive Reasoning System for Mineral Prospectivity

    Yiran Zhang, Jinwen Liu, Daniel Su, Yisu Chen, Qiang Sun, Chris Gonzalez, Eun-Jung Holden, Marco Fiorentini, Wei...

    cs.AI

    Mineral exploration requires integrating heterogeneous geochemical, geophysical, and geological evidence, yet existing prospectivity systems often provide only opaque scores or heatmaps. We present MineTRACE, a web-based system for evidence-grounded exploration of eight commodities: Cu, Au, Ni, W, Sn, Co, Ta, and Mn. Users can explore prospectivity maps, query locations or regions, inspect supporting evidence, and interact through natural...

    arxiv.org/abs/2609.02060 · PDF

  40. 40

    DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents

    Zhuoran Yu, Le Thien Phuc Nguyen, Jaden Park, Xinyi Gu, Zexue He, Soochahn Lee, Rogerio Feris, Yong Jae Lee

    cs.AI · cs.CV · cs.LG

    Multimodal Large Language Models (MLLMs) have achieved strong performance on structured visual understanding tasks such as chart and document question answering. However, existing benchmarks typically evaluate these domains in isolation, leaving underexplored a key capability: whether models can use textual context to determine how chart evidence should be selected, interpreted, and aggregated. We introduce DocHop, a benchmark for integrated...

    arxiv.org/abs/2609.02059 · PDF

  41. 41

    Monitoring Web Agents Without Internal Signals: Observable Trajectories and Key-Step Supervision

    Sitong Pan, Yipeng Shen, Yilin Lu, Caiwen Ding, Lu Cheng, Qianwen Wang

    cs.AI

    Reliable web-agent monitoring is difficult when model-internal uncertainty signals such as token logits are unavailable. In this work, we study prefix-level risk prediction for web agents using observable trajectory signals: given an evolving prefix, estimate whether the current execution remains on track or is tending toward failure. We derive two observable trajectory representations: Macro features summarize cross-step agent--environment...

    arxiv.org/abs/2609.02057 · PDF

  42. 42

    HeadWiseKV: Budgeted Per-Head Cache Residency for Hybrid Long-Context Language Models

    Renjie Xie, Juncheng Yang, Aoting Hu, Mingxi Zhang, Liyao Wu, Zheheng Hong, Wei Xu

    cs.AI

    Long-context inference retains a growing key--value (KV) cache during decoding, which consumes substantial GPU memory and can reduce generation throughput. This bottleneck remains in hybrid language models because their residual global-attention layers can dominate context-dependent cache demand. We study how to allocate this state under an aggregate KV-residency budget. We introduce HeadWiseKV, a training-free framework that compresses the...

    arxiv.org/abs/2609.02029 · PDF

  43. 43

    ClaimReceipt: Verifying Evidence Sufficiency and Coverage in Agent Evaluations

    Peiying Zhu, Sidi Chang

    cs.AI · cs.CR · cs.MA

    Agent evaluations face two distinct evidentiary questions: whether a reported claim is recomputable from retained evidence (sufficiency), and whether the retained records cover the committed experiment set (coverage). Generic logs and hash-linked transcripts answer neither reliably. We introduce ClaimReceipt, a claim-relative receipt specification and selective verifier that binds typed transaction evidence to a signed experiment manifest and...

    arxiv.org/abs/2609.01992 · PDF

  44. 44

    When Agents Implement Systems: A Case Study in Defects, Detection, and Evaluation Rigor

    Phanindra Reddy Madduru

    cs.AI · cs.MA

    As LLM coding agents increasingly perform end-to-end engineering work, we lack empirical characterization of how they behave on systems-level requirements: schema design, async orchestration, configuration correctness, and retrieval-filtering trade-offs. We present a case study of one such agent implementing a multi-component data system against a detailed pre-existing specification. Storage technologies, schema, entity-resolution algorithm,...

    arxiv.org/abs/2609.01985 · PDF

  45. 45

    Benchmarking Language Models for Statistical Problem Formulation

    Chen Wang, Junzhe Zhao, Xin Cong, Wanlu Deng, Ke Deng

    cs.AI

    Large language models (LLMs) are increasingly used as assistants for statistical and data science work, yet existing evaluations largely assume the analysis target is already specified. In practice, users arrive with informal goals and heterogeneous data, leaving the model to decide what statistical task is implied and which data are relevant. We first formalize this upstream step as Statistical Problem Formulation and decompose it into two...

    arxiv.org/abs/2609.01982 · PDF

  46. 46

    Post-Training Ternarization of Qwen3-4B Capability, Effective Bit Budget, Storage Compression, and Deployment

    Anirudh Malik, M Sparsh Mehra, Poojith Devan

    cs.AI · cs.LG

    Ultra-low-bit language models can reduce storage and memory bandwidth, but a nominal "1.58-bit" label does not fully describe the stored representation, retained capability, or runtime behavior. We study an end-to-end post-training conversion of Qwen, an instruction-tuned 4B-parameter model, using KOTMS rotation, E2M-ATQ ternarization, and GPTQ-style error compensation from TWLA. The experiment is weight-only: activations remain at 16-bit...

    arxiv.org/abs/2609.01962 · PDF

This edition is part of The Daily Abstract — cs.AI archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.