cs.CL · 2026-06-15 · No. 24

Computation and Language, 2026-06-15.

8 new papers in cs.CL. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

8 entries
  1. 01

    SIMMER: Benchmarking Latent Failures in LLM Executable Planning with a World Model

    Xiaoxin Lu, Ranran Haoran Zhang, Rui Zhang

    cs.CL · cs.AI

    Large language models (LLMs) are increasingly deployed as planners for autonomous agents in household environments. While existing benchmarks evaluate whether LLM-generated plans execute successfully, they overlook a critical type of failure: latent failures. Unlike immediate failures that trigger instant feedback at execution time and enable timely correction, latent failures do not immediately halt plan execution but silently compromise...

    arxiv.org/abs/2606.14574 · PDF

  2. 02

    Fodor and Pylyshyn's Systematicity Challenge Still Stands

    Michael Goodale, Salvador Mascarenhas

    cs.CL · cs.AI

    The recent successes of neural networks producing human-like language have caused significant stir in cognitive science, with many researchers arguing that classical puzzles about human cognition and challenges to artificial intelligence are being solved by neural networks. A notable case is the argument from systematicity due to Jerry Fodor and Zenon Pylyshyn, argues that humans display systematic biconditional dependencies. For example,...

    arxiv.org/abs/2606.14512 · PDF

  3. 03

    MoDiCoL: A Modular Diagnostic Continual Learning Dataset for Robust Speech Recognition

    Theresa Pekarek Rosin, Matthias Kerzel, Stefan Wermter

    cs.CL · cs.AI · cs.SD

    Modern Automatic Speech Recognition (ASR) systems have made remarkable progress on standard benchmarks, yet performance gaps have emerged under real-world distribution shifts, caused by recording conditions, accents, speech impairments, and noise. Existing datasets and benchmarks typically isolate these factors, which overlooks their co-occurrence in real-world applications. In this paper, we argue that model robustness can be treated as a...

    arxiv.org/abs/2606.14459 · PDF

  4. 04

    Learning to Hear Hesitation: Continual Learning for Disfluency-Aware ASR

    Henri-Leon Kordt, Theresa Pekarek Rosin, Jae Hee Lee, Stefan Wermter

    cs.CL · cs.AI · cs.SD

    Despite advances in large-scale Automatic Speech Recognition (ASR), disfluent speech remains challenging, as state-of-the-art systems are often optimized to omit disfluencies, leading to information loss and hallucinations. Prior work has focused on verbatim transcription and the integration of disfluency markers, but adapting models on limited datasets can lead to catastrophic forgetting of general-domain knowledge. We address this gap by...

    arxiv.org/abs/2606.14391 · PDF

  5. 05

    Achieving Precise Text-To-Cypher Via Grounded Knowledge Graph Data Generation

    Francesco Cazzaro, Jessica Lennon, Ariadna Quattoni

    cs.CL · cs.AI

    Property Graphs are rapidly being adopted as database frameworks for representing heterogeneous data sources. To enable precise access to the information contained in them we need conversational interfaces based on Text-To-Cypher (Text2Cypher) parsers. This paper presents an automatic synthetic data generation method that can be leveraged to fine-tune small LLMs for this task. We conduct experiments on all the major Text-To-Cypher benchmarks,...

    arxiv.org/abs/2606.14325 · PDF

  6. 06

    OdysSim: Building Foundation Models for Human Behavior Simulation

    Xuhui Zhou, Weiwei Sun, Weihua Du, Jiarui Liu, Haojia Sun, Qianou Ma, Tongshuang Wu, Yiming Yang, Maarten Sap

    cs.CL · cs.AI · cs.LG

    Large language models are increasingly deployed as human simulators for interactive evaluation and social simulation. Yet helpfulness-driven post-training pulls them toward a homogeneous, overly agreeable assistant register, creating a behavioral Sim2Real gap. We present OdysSim, the largest open systematic investigation of behavioral foundation models, i.e., models trained to simulate human behavior at scale. We propose SOUL, a taxonomy of...

    arxiv.org/abs/2606.14199 · PDF

  7. 07

    Implicit Reasoning for Large Language Model-based Generative Recommendation

    Yinhan He, Liam Collins, Bhuvesh Kumar, Jundong Li, Neil Shah, Donald Loveland

    cs.CL · cs.AI

    Large Language Models (LLMs) are increasingly adopted as backbones for Generative Recommendation (GR), promising access to pretrained world knowledge. Yet reliably invoking this knowledge for GR remains poorly understood. A key obstacle is that LLM-based GR typically represents items with Semantic IDs (SIDs), disrupting LLMs' natural-language reasoning interface because these tokens are unseen by the LLM during pretraining. Existing...

    arxiv.org/abs/2606.14142 · PDF

  8. 08

    SANA: What Matters for QA Agents over Massive Data Lakes?

    Austin Senna Wijaya, Jiaxiang Liu, Haonan Wang, Eugene Wu

    cs.CL · cs.AI · cs.DB

    Exploratory question answering (EQA) over data lakes requires an LLM agent to discover relevant sources, analyze retrieved data, and adapt its actions based on intermediate results. End-to-end accuracy alone cannot distinguish failures in search, planning, data analysis, or the agent's Action Policy: its decisions about what to do next and when to submit an answer. We present SANA (Search Agent Navigation Ablation framework), a diagnostic...

    arxiv.org/abs/2606.13904 · PDF

This edition is part of The Daily Abstract — cs.CL archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.