cs.AI · 2026-06-09 · No. 18

Artificial Intelligence, 2026-06-09.

46 new papers in cs.AI. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

46 entries
  1. 01

    Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting

    Avijit Ghosh, Anka Reuel, Jenny Chim, Wm. Matthew Kennedy, Srishti Yadav, Jennifer Mickel, Yanan Long, Andrew Tran,...

    cs.AI

    AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs. The cost is interpretive: readers cannot reliably compare results across sources, identify what a report omits, or trace an aggregate claim to its underlying evidence. Recent efforts address isolated components but leave three gaps: they cover only narrow slices of the evaluation lifecycle and do not...

    arxiv.org/abs/2606.09809 · PDF

  2. 02

    SIGA: Self-Evolving Coding-Agent Adapters for Scientific Simulation

    Matthew Ho, Brian Liu, Jixuan Chen, Audrey Wang, Lianhui Qin

    cs.AI · cs.CL

    Advanced scientific simulators expose specialized input languages that turn simulation goals into executable configurations, but learning them can cost domain scientists hours to days. We study simulator setup as a problem of agent-tool interface grounding: what minimal simulator-specific adaptations are needed for an off-the-shelf coding agent to operate real scientific software? Our intuition is that coding agents already know how to...

    arxiv.org/abs/2606.09774 · PDF

  3. 03

    Collaborative Human-Agent Protocol (CHAP)

    Arsalan Shahid, Gordon Suttie, Philip Black

    cs.AI · cs.CL · cs.HC

    Foundation models are moving from response generation into operational roles. They plan across steps, call tools, request human input, coordinate with other agents, and increasingly carry responsibility for work that affects customers, claims, code, contracts, and clinical decisions. Production deployments are no longer one human supervising one model. They are multi-human, multi-agent collaborations that cross teams, time zones, and trust...

    arxiv.org/abs/2606.09751 · PDF

  4. 04

    Multi-Turn Evaluation of Deep Research Agents Under Process-Level Feedback

    Rishabh Sabharwal, Hongru Wang, Amos Storkey, Jeff Z. Pan

    cs.AI · cs.CL · cs.LG

    Existing benchmarks for deep research agents (DRAs) assess only single-shot outputs, ignoring a key question: can DRAs improve their reports when guided by feedback? To investigate this, we conduct a multi-turn evaluation of DRAs under two feedback settings: self-reflection, in which the agent revises its report without any external diagnostic signal, and process-level feedback, in which the agent receives guidance targeting gaps in its...

    arxiv.org/abs/2606.09748 · PDF

  5. 05

    SearchSwarm: Towards Delegation Intelligence in Agentic LLMs for Long-Horizon Deep Research

    Pu Ning, Quan Chen, Kun Tao, Xinyu Tang, Tianshu Wang, Qianggang Cao, Xinyu Kong, Zujie Wen, Zhiqiang Zhang, Jun Zhou

    cs.AI

    Large language models are increasingly expected to handle complex, long-horizon real-world tasks whose context demands can grow without bound, yet model context windows remain inherently finite. Recent work explores a paradigm where a main agent decomposes tasks and dispatches subtasks to subagents, which execute and return only summarized results, conserving the main agent's context budget. However, performing this well requires delegation...

    arxiv.org/abs/2606.09730 · PDF

  6. 06

    Beyond Probabilistic Similarity: Structural, Temporal, and Causal Limitations of Retrieval-Augmented Generation in the Legal Domain

    Hudson de Martim

    cs.AI

    Retrieval-Augmented Generation (RAG) has become a standard architectural response to unreliability in legal AI, yet high-profile failures, including fabricated citations submitted to courts and anachronistic legal content presented as current, continue to appear across jurisdictions. We argue that these failures are not residual confabulations to be eliminated by scaling language models, but symptoms of an architectural mismatch between...

    arxiv.org/abs/2606.09724 · PDF

  7. 07

    Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization

    Mohammad Beigi, Ming Jin, Lifu Huang

    cs.AI · cs.LG

    Reward hacking is usually studied after it becomes visible, once a model earns high proxy reward while failing the intended task. We instead study what proxy RL teaches before that failure appears. We introduce Proxy Reward Internalization and Mechanistic Exploitation (PRIME), a learned capability to assess task correctness, predict proxy acceptance, and reason about exploitable proxy--gold gaps. In coding RL environments with exploitable...

    arxiv.org/abs/2606.09711 · PDF

  8. 08

    (Auto)formalization is supposed to be easy: Trellis process semantics for spelling out rigorous proofs

    Wesley Pegden

    cs.AI · cs.LO · math.CO

    We present Trellis: an autoformalization system that leverages LLM agents in a deterministically constrained workflow to enforce incremental progress in Lean autoformalization tasks through iterative refinement of natural language proofs. Our approach is motivated by the common mathematician's notion of what it means to have a rigorous proof in the first place: namely, that it would be routine to elaborate any part of the proof in further...

    arxiv.org/abs/2606.09674 · PDF

  9. 09

    Correlation Is Not Enough: Embedding Human Metadata for Individual Causal Discovery

    Suraj Biswas, Saurabh Gupta, Pritam Mukherjee

    cs.AI · cs.CL · cs.LG · cs.PF · q-bio.QM

    Ask a pretrained biomedical language model whether "cortisol 28 ug/dL" and "stock-market volatility" are related, and it returns a cosine similarity of 0.83 on a scale where 1.0 means identical. The two share no mechanism. This is not a corner case: every off-the-shelf biomedical encoder we tested (BioBERT, PubMedBERT, BioM-ELECTRA) scores unrelated cross-domain pairs between 0.76 and 0.92 when the answer should be near zero. Accuracy on...

    arxiv.org/abs/2606.09672 · PDF

  10. 10

    SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks

    Hongcheng Gao, Hailong Qu, Jingyi Tang, Jiahao Wang, Zihao Huang, Hengkang Qiao, Shihong Huang, Junming Yang, Yi Li,...

    cs.AI · cs.CL

    Spatial reasoning is a foundational capability for multimodal large language models (MLLMs) to perceive and operate within the physical world. However, existing benchmarks predominantly rely on passive evaluation (e.g., static VQA) or simulator-specific pipelines, failing to assess general interactive spatial understanding. We introduce SpatialWorld, a unified benchmark designed specifically for evaluating the interactive spatial...

    arxiv.org/abs/2606.09669 · PDF

  11. 11

    Frequency-based Constrained Sampling for Interval Patterns

    Djawad Bekkoucha, Abdelkader Ouali, Bruno Crémilleux

    cs.AI

    Output space pattern sampling is a powerful alternative to exhaustive pattern mining for exploring large pattern spaces, as it enables users to focus on representative patterns drawn according to a chosen interestingness measure. In this paper, we address the problem of sampling interval patterns under user-defined syntactic constraints. We introduce CFips, a sampling approach that incorporates constraints directly into the sampling...

    arxiv.org/abs/2606.09666 · PDF

  12. 12

    From 0-to-1 to 1-to-N: Reproducible Engineering Evidence for MetaAI Recursive Self-Design

    Dun Li, Jiatao Li, Hongzhi Li

    cs.AI

    Recursive self-design refers to AI-assisted modification of the mechanisms by which an AI system is built, evaluated, and improved. This paper treats MetaAI not as a mature paradigm, but as a working term for a human-seeded, AI-expanded development pattern in which the design space itself becomes a target of modification. We propose an operational evidence framework with four criteria: inspectable target system, meta-level modifier,...

    arxiv.org/abs/2606.09663 · PDF

  13. 13

    Next-Token Prediction Learns Generalisable Representations of Sleep Physiology

    Jonathan F. Carter, Lionel Tarassenko

    cs.AI

    Foundation models offer a promising route to compress multi-modal physiological signals into compact representations of human health, with broad applications across sleep medicine, cardiology, neurology and other healthcare domains. Existing models have typically been trained with masked-reconstruction or contrastive objectives. However, masked reconstruction may be poorly suited to the stochastic nature of these signals, while contrastive...

    arxiv.org/abs/2606.09605 · PDF

  14. 14

    Optical Reasoning: Rethinking Images as an Expressive Reasoning Medium Beyond Text

    Yutong Bian, Dongjie Cheng, Heming Xia, Yongqi Li, Wenjie Li

    cs.AI

    Chain-of-Thought (CoT) improves the performance of Large Language Models (LLMs) and has been extended to Multimodal Large Language Models (MLLMs). More recent work further moves from text-based multimodal reasoning toward interleaved-modal reasoning, where intermediate steps can incorporate both textual rationales and visual evidence. In this work, we propose a bolder and more ambitious idea: could images alone serve as the reasoning medium...

    arxiv.org/abs/2606.09585 · PDF

  15. 15

    TABVERSE: Benchmarking Cross-Format Table Understanding in LLMs and VLMs

    Momina Ahsan, Sarfraz Ahmad, Ming Shan Hee, Roy Ka-Wei Lee, Preslav Nakov

    cs.AI · cs.CL · cs.IR

    Large Language Models (LLMs) and Vision-Language Models (VLMs) are increasingly evaluated on table reasoning tasks, but the role of table representation remains under-explored. In practice, the same table content may appear in different structural formats, such as HTML, Markdown, and LaTeX, or as rendered images. However, existing evaluations often let content, format, layout, and modality vary together, making it difficult to isolate...

    arxiv.org/abs/2606.09578 · PDF

  16. 16

    Self-Explainability in Self-Adaptive and Self-Organising Systems: Status and Research Directions

    Tom Beyer, Svea Wisy, Sven Tomforde

    cs.AI

    The growing complexity of self-adaptive and self-organising systems, fuelled by advances in Artificial Intelligence (AI), has made them increasingly difficult to understand and trust. While Explainable AI aims to provide insight into AI decision-making, a more advanced goal is for systems to explain themselves - an ability referred to as Self-Explainability (SX). This article presents a systematic literature review on SX, analysing existing...

    arxiv.org/abs/2606.09568 · PDF

  17. 17

    PRISM: Recovering Instruction Sets from Language Model Activations

    Gilad Gressel, Rahul Pankajakshan, Julia Diament, Efim Hudis, Krishnashree Achuthan, Yisroel Mirsky

    cs.AI · cs.LG

    As LLMs are deployed as agents, reliable monitoring requires knowing not only what they output, but which instructions are steering their behavior. This is difficult when models infer unintended subgoals, follow contextual cues, or are influenced by prompt injections and hidden objectives. While activation-to-language methods suggest that hidden states can reveal natural-language information, existing approaches are not designed to recover...

    arxiv.org/abs/2606.09563 · PDF

  18. 18

    AI Scientists Are Only as Good as Their Evidence: A Stratified Ablation of Proprietary Data and Reasoning Skills in Drug-Asset Valuation

    Yinan Wang

    cs.AI

    AI Scientist agents are often evaluated as if capability were mainly a function of model quality, prompting, or reasoning scaffolds. We test a different hypothesis in drug-asset valuation: for knowledge-intensive scientific decisions, the limiting factor is often the evidence substrate the agent can access. We run a controlled three-arm ablation on a production valuation agent: A is a plain web-only LLM analyst, B adds public structured tools...

    arxiv.org/abs/2606.09556 · PDF

  19. 19

    From Rigid to Dynamic: Entropy-Guided Adaptive Inference for Long-Context LLMs

    Zhanchao Xu, Haoyang Li, Qingfa Xiao, Fei Teng, Chen Jason Zhang, Lei Chen, Qing Li

    cs.AI · cs.CL

    Existing sparse attention and KV cache compression methods for long-context LLM inference typically apply fixed sparsity patterns or uniform budgets across all attention heads, overlooking the substantial variation in attention behavior among heads and contexts. We observe two distinct entropy patterns among attention heads: Rigid Heads, whose entropy stays near zero across input segments, and Dynamic Heads, whose entropy fluctuates...

    arxiv.org/abs/2606.09508 · PDF

  20. 20

    Deterministic Integrity Gates for LLM-Assisted Clinical Manuscript Preparation: An Auditable Biomedical Informatics Architecture

    Yoojin Nam, Jinhoon Jeong, Namkug Kim

    cs.AI · cs.DL

    Objective. Large language models (LLMs) increasingly draft clinical research manuscripts, but their fluency can hide fabricated citations, numbers that drift from source tables, and unmet reporting-guideline items. Existing tools generate text without verifying it, and self-critique inherits the blind spots that produce confident fabrication. We describe an architecture that pairs generation with verification. Methods. The design rests on...

    arxiv.org/abs/2606.09500 · PDF

  21. 21

    LLM-Orchestrated Conformance Checking in Stroke Care Without Computer-Interpretable Guidelines

    Giorgio Leonardi, Stefania Montani, Manuel Striani, Alessandro Canessa, Delfina Ferrandi

    cs.AI

    Objective: Conformance checking in healthcare seeks to assess whether patient care pathways adhere to clinical guidelines. However, its practical application often depends on the availability of formal, machine-interpretable representations of guidelines, such as Computer-Interpretable Guidelines (CIGs), which are seldom available in real-world clinical settings. Methods: This work introduces a modular framework based on the orchestration of...

    arxiv.org/abs/2606.09489 · PDF

  22. 22

    Emergent alignment and the projectability of ethical personas

    Guillermo Del Pinal, Youngchan Lee, Cameron McNamara, Alejandro Perez Carballo

    cs.AI · cs.LG

    Work on `emergent misalignment' shows that finetuning LLMs on narrow tasks can induce broadly misaligned behavior. This supports the `persona selection' (PSM) hypothesis: during pre-training, LLMs learn to simulate different characters and perspectives, which can be elicited and refined during post-training. This paper investigates the converse phenomenon, `emergent alignment', and uses it to support and refine the PSM and motivate a novel...

    arxiv.org/abs/2606.09475 · PDF

  23. 23

    TheoremBench: Evaluating LLMs on Theorem Proving in Formal Mathematics

    QuocViet Pham, Elvir Karimov, Andrey Galichin, Ivan Oseledets

    cs.AI

    LLMs have recently achieved strong results on formal proving benchmarks. However, existing evaluations remain heavily concentrated on competition-style problems and often fail to capture how models behave on longer, more dependency-rich mathematical developments. We introduce TheoremBench, a Lean4 benchmark designed to evaluate theorem provers beyond contest settings. The benchmark is built from nearly one hundred classical theorems and is...

    arxiv.org/abs/2606.09450 · PDF

  24. 24

    AliyunConsoleAgent: Training Web Agents in Real-World Cloud Environments via Distillation and Reinforcement Learning

    Bojie Rong, Zheyu Shen, Qiaoping Wang, Pengfei Kang, Yang Xu, Yawen Wei, Hanyu Wu, Zhi Zhao, Leihao Pei, Linquan Jiang

    cs.AI

    We present AliyunConsoleAgent, a web agent framework for automated documentation verification in real-world cloud consoles. Major cloud platforms encompass hundreds of products with rapid feature iteration, causing console UIs to frequently diverge from their corresponding documentation. Verifying that documented procedures accurately reflect the current console and can be executed end-to-end demands an estimated 4 million recurring...

    arxiv.org/abs/2606.09447 · PDF

  25. 25

    SIFT: Selective-Index For Fast Compute of RAG Prefill by Exploiting Attention Invariance

    Rya Sanovar, Srikant Bharadwaj, Hritvik Taneja, Moinuddin Qureshi

    cs.AI · cs.AR

    Retrieval-Augmented Generation (RAG) injects LLM queries with relevant documents to improve response quality. This injection increases prompt length and slows time to first token (TTFT). Unlike standard queries, RAG queries have a unique property of context reuse where the same documents recur across user queries. Thus, fully recomputing documents for every RAG query does redundant compute and increases TTFT. Prior works precompute KV tensors...

    arxiv.org/abs/2606.09441 · PDF

  26. 26

    Bayesian Selective Latent Inference for Wastewater-First Influenza Monitoring

    Yixuan Zhang, Yang Song, Hao Wang, Samir Bhatt, Hengguan Huang

    cs.AI

    Wastewater influenza surveillance can reveal community circulation before clinical reporting, but wastewater alone is not a fully identifiable proxy for human burden. Existing wastewater models assume a fixed evidence set, while generic evidence-acquisition methods treat official surveillance streams as interchangeable costly features. We cast wastewater-first influenza monitoring as a selective decision problem: starting from mandatory...

    arxiv.org/abs/2606.09433 · PDF

  27. 27

    WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces

    Wanli Li, Bowen Zhou, Yunyao Yu, Zhou Xu, Yifan Yang, Dongsheng Li, Caihua Shan

    cs.AI

    Computer-use agents (CUAs) increasingly operate in runtimes that combine visual desktop control, command-line execution, code editing, browsers, and external tools. Existing benchmarks, however, often evaluate these interfaces as separable capabilities, leaving long-horizon cross-interface orchestration under-tested. Thus, we introduce WeaveBench, a long-horizon hybrid-interface benchmark with 114 tasks across 8 real-world work domains,...

    arxiv.org/abs/2606.09426 · PDF

  28. 28

    Correct Looks Better: Pairwise Comparisons Reveal Accuracy Rankings

    Mina Remeli, Moritz Hardt

    cs.AI · cs.CL · cs.LG

    Pairwise comparisons combined with aggregation methods like Elo have become central to evaluating generative models, yet concerns remain that they reward superficial stylistic cues or display judge biases. In a more positive turn, we show that model rankings from pairwise comparisons strongly agree with ground-truth-based accuracy rankings when such ground truth is available for comparison. By converting five well-known benchmarks into...

    arxiv.org/abs/2606.09409 · PDF

  29. 29

    Capacity, Not Format: Rethinking Structured Reasoning Failures

    Hengxin Fan

    cs.AI · cs.CL

    Prior work treats structured output as a reasoning tax, but this framing is incomplete: the cost of formatting depends strongly on a model's spare capacity. Using information-matched prose controls and a four-level schema complexity gradient, we separate format-specific effects from prompt-length confounds across 4 models and 5 benchmarks with 0% parse failures on successfully generated responses. We find that structured formats are...

    arxiv.org/abs/2606.09410 · PDF

  30. 30

    RunAgent SuperBrowser: A Theory of Autonomous Web Navigation Grounded in Human Browsing Behaviour

    Radeen Mostafa, Sawradip Saha

    cs.AI

    We present SUPERBROWSER, an autonomous web-navigation agent designed against a single guiding hypothesis: a web agent should browse the way a person browses. A human reading a page does not retain every pixel they have seen; they look at a few candidate targets, decide on one, and remember only what is needed to keep the goal alive. We operationalize this perception-cognition-action triad as three coupled mechanisms. First, a vision-first...

    arxiv.org/abs/2606.09399 · PDF

  31. 31

    From Coarse to Fine: Managing Temporal Granularity in Spatio-Temporal Data for Fine-Grained Traffic Prediction

    Shuhao Li, Weidong Yang, Yue Cui, Zizhuo Xu, Lipeng Ma, Fan Zhang, Xiaofang Zhou

    cs.AI

    Efficient acquisition, storage, and utilization of traffic data are critical challenges in spatio-temporal data management. Most traffic data systems collect and store observations at fixed, coarse-grained temporal intervals to reduce storage and computation costs. However, such coarse-grained data severely limits downstream applications that require predictions at a finer temporal granularity. Collecting and maintaining fine-grained traffic...

    arxiv.org/abs/2606.09392 · PDF

  32. 32

    Capability-Aligned Hierarchical Learning for Tool-Augmented LLMs

    Haotong Yang, Ting Long, Yi Chang

    cs.AI

    Tool learning enables LLMs to invoke external tools to accomplish tasks. Prior studies have demonstrated the effectiveness of a hierarchical structure: a high-level policy handles global planning and decomposes tasks into manageable sub-tasks, and a low-level policy focuses on invoking tools to solve these sub-tasks. However, these works typically optimize the high-level and low-level policies separately, leading to planner-executor...

    arxiv.org/abs/2606.09371 · PDF

  33. 33

    Experience Makes Skillful: Enabling Generalizable Medical Agent Reasoning via Self-Evolving Skill Memory

    Haoran Sun, Wenjie Li, Yujie Zhang, Zekai Lin, Fanrui Zhang, Kaitao Chen, Xingqi He, Yichen Li, Mianxin Liu, Lei...

    cs.AI · cs.CL

    Medical agent systems are increasingly expected to support interactive clinical decision making rather than only static question answering. In such settings, effective agents must reuse prior experience across evolving cases, yet existing memory mechanisms often retain raw historical traces that are redundant, noisy, and difficult to govern. More importantly, they rarely distinguish which memories are truly useful for future reasoning. This...

    arxiv.org/abs/2606.09365 · PDF

  34. 34

    Leveraging Structural Constraints for Diffusion-based Neural TSP Solvers

    Mickaël Basson, Philippe Preux

    cs.AI

    Neural combinatorial optimization has recently achieved strong results on the Euclidean Traveling Salesman Problem (TSP) using generative models such as diffusion and consistency models. State-ofthe-art approaches like FT2T combine fast consistency-based prediction with gradient-based inference time refinement. However, gradient search often incurs significant computational overhead and may not align with the discrete structure of feasible...

    arxiv.org/abs/2606.09343 · PDF

  35. 35

    TRL-Bench: Standardizing Cross-Paradigm Representation-Level Evaluation of Tabular Encoders

    Wei Pang, Xiangru Jian, Hehan Li, Zhixuan Yu, Alex Xue, Jinyang Li, Zhengyuan Dong, Xinjian Zhao, Hao Xu, Chao...

    cs.AI · cs.DB

    Tabular encoders are usually evaluated inside task-specific end-to-end pipelines, so models from different training paradigms are difficult to compare directly even when they operate on similar tabular signals. We introduce TRL-Bench, a multi-granular tabular representation learning (TRL) benchmark that standardizes cross-paradigm representation-level evaluation: each encoder exports row-, column-, or table embeddings through its supported...

    arxiv.org/abs/2606.09323 · PDF

  36. 36

    Anything2Skill: Compiling External Knowledge into Reusable Skills for Agents

    Qianjun Pan, Yutao Yang, Junsong Li, Jie Zhou, Kai Chen, Xin Li, Qin Chen, Liang He

    cs.AI

    Retrieval-augmented generation (RAG) enables agents to access external knowledge at inference time, but it primarily retrieves fragmented declarative evidence, leaving agents to repeatedly infer task procedures from passages, manuals, examples, logs, or trajectories. This raises a fundamental question: can skills extracted from external knowledge bases be installed into an agent, enabling it to rapidly approximate domain expertise? In this...

    arxiv.org/abs/2606.09316 · PDF

  37. 37

    FF-JEPA: Long-Horizon Planning in World Models with Latent Planners

    Sergi Masip, Jonathan Swinnen, Yutong Hu, Renaud Detry, Tinne Tuytelaars

    cs.AI

    Joint Embedding Predictive Architectures (JEPAs) have shown promising world modeling capabilities, enabling planning in latent space by optimizing action trajectories using methods like the Cross-Entropy Method (CEM). These methods are, however, too computationally expensive and ineffective for long-horizon planning. Furthermore, these methods typically require an explicit image of the goal state, which is not always possible in real-world...

    arxiv.org/abs/2606.09311 · PDF

  38. 38

    MASS: Deep Research for Social Sciences with Memory-Augmented Social Simulation

    Yongrui Liu, Deyi Xiong

    cs.AI

    Deep Research agents powered by Large Language Models (LLMs) have exhibited extraordinary potential in automated paper writing tasks. However, existing systems rely heavily on literature retrieval and synthesis through internet and local knowledge bases, often resulting research in lacking insight and creativity in social science. To address this issue, we propose "Memory-Augmented Social Simulation (MASS)", an innovative paradigm that...

    arxiv.org/abs/2606.09198 · PDF

  39. 39

    IMUG-Bench: Benchmarking Unified Multimodal Models on Interleaved Understanding and Generation

    Lingyi Meng, Zecong Tang, Haoran Li, Tengju Ru, Zhejun Cui, Weitong Lian, Qi Kang, Hangshuo Cao, Yichen Zhu, Yechi...

    cs.AI · cs.CV · cs.MM

    In recent years, unified multimodal models (UMMs) have emerged to support both understanding and generation within a single framework. Mastering dynamic, multi-turn interleaved image-text dialogues is a crucial task for UMMs in real-world applications. However, existing benchmarks fail to evaluate this important task, as they are often limited to single-turn or static settings, and typically overlook exposure bias in multi-turn interactions....

    arxiv.org/abs/2606.09169 · PDF

  40. 40

    Reliable to Expressive: A Curriculum for Rubric-Following Safety Judges

    Yongtaek Lim, Hyeji Choi, Minwoo Kim

    cs.AI

    Safety judges are increasingly deployed to evaluate model outputs against evolving criteria, yet recent meta-evaluation work shows they remain brittle under prompt and rubric variation, with false negative-rate swings of up to 0.24 reported for stylistic perturbations alone. We argue that safety judgment is fundamentally a rubric-following problem: a robust judge must apply the given evaluation criteria consistently across rubric formulations...

    arxiv.org/abs/2606.09165 · PDF

  41. 41

    Vision Language Model Helps Private Information De-Identification in Vision Data

    Tiejin Chen, Pingzhi Li, Kaixiong Zhou, Tianlong Chen, Hua Wei

    cs.AI

    Visual Language Models (VLMs) have gained significant popularity due to their remarkable ability. While various methods exist to enhance privacy in text-based applications, privacy risks associated with visual inputs remain largely overlooked such as Protected Health Information (PHI) in medical images. To tackle this problem, two key tasks: accurately localizing sensitive text and processing it to ensure privacy protection should be...

    arxiv.org/abs/2606.09132 · PDF

  42. 42

    Late-Layer Fusion is Enough: Dual-Path Vision Token Routing for Multimodal Large Language Models under Visual Saturation

    Siyuan Liu, Jinyang Wu

    cs.AI · cs.CL · cs.CV · cs.LG

    Multimodal large language models (MLLMs) commonly inherit the deep, symmetric Transformer backbone designed for unimodal text modeling, and apply the same computation uniformly to image and language tokens. This design overlooks a key modality asymmetry: image and text tokens differ substantially in information density, redundancy, and required reasoning depth. Through a layer-wise analysis of LLaVA-1.5, we observe that vision tokens tend to...

    arxiv.org/abs/2606.09131 · PDF

  43. 43

    A Regret Minimization Framework on Preference Learning in Large Language Models

    Suhwan Kim, Taehyun Cho, Geon-Hyeong Kim, Yu Jin Kim, Youngsoo Jang, Moontae Lee, Jungwoo Lee

    cs.AI

    Reinforcement learning with verifiable rewards (RLVR) has enabled progress on reasoning-intensive tasks by relying on task-specific verifiers that provide automated correctness signals. However, many realistic language tasks are difficult to equip with reliable verifiers, motivating a growing reliance on reinforcement learning from human feedback (RLHF). In this setting, we argue that a closer examination of how human feedback should be...

    arxiv.org/abs/2606.09124 · PDF

  44. 44

    ComplexConstraints and Beyond: Expert Rubrics for RLVR

    Sushant Mehta, Liudas Panavas, Edwin Chen

    cs.AI

    As LLM capabilities advance rapidly, the evaluation methods used to assess them increasingly lag behind. Traditional benchmarks relied on programmatic verification of narrow, surface-level constraints, but real-world instruction following and agentic tasks demand assessment of nuanced, context-dependent behaviors that resist simple scripted checks. We present a systematic analysis of expert-curated rubric-based evaluation as an alternative...

    arxiv.org/abs/2606.09118 · PDF

  45. 45

    Graph2Idea:Retrieval-Augmented Scientific Idea Generation with Graph-Structured Contexts

    Xu Li, Hanzhe Tu, Xun Han

    cs.AI

    Generating novel, feasible, and high-quality research ideas is an important yet challenging task in scientific discovery.Recent Large Language Model (LLM)-based methods often ground idea generation with retrieved literature, but the retrieved evidence is usually provided as flat text, such as titles, abstracts, or summaries. Such flat contexts may contain redundant or weakly relevant information, while making cross-paper relations among...

    arxiv.org/abs/2606.09105 · PDF

  46. 46

    DynaOD: Dynamic Origin-Destination Flow Generation with Discrete-to-Continuous Temporal Semantic Modeling

    Jie Zhao, Xianqi Dai, Jie Feng, Huandong Wang, Yong Li

    cs.AI

    Dynamic origin-destination (OD) flow generation seeks to synthesize realistic mobility dynamics from temporal context alone, without relying on historical OD observations. A key challenge is to translate semantic temporal signals into temporally coherent OD patterns while preserving the inherent spatial heterogeneity of urban regions. We propose DynaOD, a semantic-driven framework that models temporal dynamics through two complementary...

    arxiv.org/abs/2606.09086 · PDF

This edition is part of The Daily Abstract — cs.AI archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.