cs.CL · 2026-07-08 · No. 47

Computation and Language, 2026-07-08.

19 new papers in cs.CL. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

19 entries
  1. 01

    RSF-GLLM: Bridging the Semantic Gap in Multi-Hop Knowledge Graph QA via Recurrent Soft-Flow and Decoupled LLM Generation

    Sambaran Bandyopadhyay, Ananth Muppidi

    cs.CL · cs.AI

    Multi-hop Question Answering over Knowledge Graphs faces a critical challenge: traditional retrieve-then-read pipelines break differentiability, preventing the retriever from learning to bridge the semantic gap where intermediate nodes lack lexical overlap with the query. To address this, we propose RSF-GLLM, a framework decoupling differentiable graph reasoning from answer generation. Our Recurrent Soft-Flow (RSF) module employs a GRU-guided...

    arxiv.org/abs/2607.06527 · PDF

  2. 02

    Pitwall: Faithful Natural-Language Race-Strategy Briefings from a Calibrated Real-Time Monte Carlo Engine

    Juan S. Santillana

    cs.CL · cs.AI · cs.LG · stat.AP

    Live sports commentary is grounded generation under a deadline: statements concern real, named athletes, the grounding state changes every few seconds, and no reference text exists at generation time. We present Pitwall, a production system that generates natural-language Formula 1 strategy briefings in English, Spanish, and Portuguese, treating faithfulness as an architectural property rather than an aspiration: every published sentence is...

    arxiv.org/abs/2607.06495 · PDF

  3. 03

    Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities

    So Hasegawa, Shailaja Keyur Sampat, Lei Liu, Wei-Peng Chen

    cs.CL · cs.AI

    Current benchmarks for evaluating Large Language Models (LLMs) in data analysis often fail to reflect real-world settings. They typically focus on fact retrieval from small tables and overlook the challenges of large multi-tabular datasets, external knowledge integration, and exploratory insight discovery. We introduce DataGovBench, a benchmark derived from governmental open data designed to evaluate LLMs in practical scenarios. The benchmark...

    arxiv.org/abs/2607.06482 · PDF

  4. 04

    From Voting to Agent Collaboration: Answer-Type-Aware LLM Pipelines for BioASQ 14b

    Taeyun Roh, Eunha Lee, Wonjune Jang, Sohyun Chung, Junha Jung, Jaewoo Kang

    cs.CL · cs.AI

    Biomedical question answering requires not only accurate extraction of information from scientific literature but also reliable integration of evidence across multiple documents. This study presents a question-type-specific large language model (LLM) framework for BioASQ 14b Task B, designed to improve answer robustness and evidence grounding in biomedical question answering. Rather than applying a single prompting strategy to all questions,...

    arxiv.org/abs/2607.06452 · PDF

  5. 05

    Estimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMs

    Andrea Alfarano, Andrea Bacciu, Saab Mansour, Amin Mantrach, Marcello Federico

    cs.CL · cs.AI

    Uncertainty estimation (UE) enables LLM-powered systems to recognize when to abstain, yet existing research has predominantly focused on English. We present the first large-scale evaluation of UE methods across 22 languages, spanning high-, mid-, and low-resource settings. Using two human-curated Q\&A datasets, we compare open and closed box UE methods (nine in total) across different model sizes and architectures while eliciting long-form...

    arxiv.org/abs/2607.06327 · PDF

  6. 06

    Spider 2.0-AIFunc: Extending Real-World Text-to-SQL to AI-Native SQL Workflows

    Tianyang Liu, Canwen Xu, Fangyu Lei, Nikki Lijing Kuang, Jixuan Chen, Tao Yu, Julian McAuley, Zhewei Yao, Yuxiong He

    cs.CL · cs.AI · cs.DB

    Major cloud data platforms now expose large language model capabilities as native SQL functions, enabling analysts to perform classification, filtering, sentiment analysis, extraction, similarity search, and aggregation within ordinary SQL queries. Yet existing text-to-SQL benchmarks evaluate only conventional SQL and provide no signal on whether models can generate such AI-native SQL. We introduce Spider 2.0-AIFunc, a benchmark of 465...

    arxiv.org/abs/2607.06229 · PDF

  7. 07

    Improving LLM-Generated Process Model Quality Through Reinforcement Learning: The Role of Reward Function Design

    Alexander Rombach, Chantale Lauer, Nijat Mehdiyev

    cs.CL · cs.AI · cs.LG

    Large language models (LLMs) can generate BPMN process models from natural-language descriptions, yet supervised fine-tuning (SFT) limits their output quality to the patterns present in the training data. Reinforcement learning (RL) can optimize beyond this ceiling using external quality measures, but how the reward function should be designed when quality is multi-dimensional remains unexplored. We present a systematic investigation of...

    arxiv.org/abs/2607.06175 · PDF

  8. 08

    LongCrafter: Towards Diverse Long-Context Understanding via Evidence-Graph-Guided Instruction Synthesis

    Chenhao Yuan, Yinhao Xu, Shuwen Xu, Xizhi Yang, Jiaxiang Liu, Chenxi Zhou, Shaoping Huang, Haolin Ren, Pengfei Cao,...

    cs.CL · cs.AI

    Synthesizing long-context supervised fine-tuning (SFT) data is a scalable way to enhance the long-context understanding of large language models (LLMs), yet existing approaches share three limitations: narrow task coverage, insufficient instruction difficulty, and a lack of faithfulness supervision. We propose \textbf{LongCrafter}, a structured synthesis framework that couples a hierarchical task taxonomy with an evidence-grounded pipeline....

    arxiv.org/abs/2607.06160 · PDF

  9. 09

    LLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial Observability

    Chenxu Wang, Yongkun Yang, Boyuan Du, Shiwei Lin, Huaping Liu

    cs.CL · cs.AI

    Deliberation plays a crucial role in collaboration; when humans work together, they naturally engage in communication to align information and reach an agreement. In this paper, we investigate deliberative large language model (LLM) agents under partially observable joint decision-making tasks. We formalize deliberative collaboration as a cooperative joint decision problem with partial and asymmetric observations, and introduce a scalable...

    arxiv.org/abs/2607.06157 · PDF

  10. 10

    From Blueprint to Reality: Modeling and Applying Putnam's Social Capital Theory with LLM-based Multi-agent Simulations

    Shiyi Ling, Zhi Zheng, Hui Zheng, Wenjun Xue, Feng Ye, Tong Xu

    cs.CL · cs.AI · cs.SI

    Putnam's Social Capital Theory is a foundational framework for collective action and community prosperity. However, traditional empirical methods face practical limits on control and replication. Meanwhile, LLM-based social simulations are typically behavior-driven and lack theory-aligned environments for modeling Putnam's core propositions. To address these gaps, we introduce SocaSim, an LLM-based multi-agent simulation framework to study...

    arxiv.org/abs/2607.06080 · PDF

  11. 11

    PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages

    Daryna Dementieva, Nikolay Babakov, Kathy Hämmerl, Ilseyar Alimova, Jindřich Libovický, Shu Okabe, Miras Baisbay,...

    cs.CL · cs.AI

    Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), yet existing benchmarks remain heavily biased toward high-resource languages, with English and Chinese dominating both pre-training corpora and evaluation suites. The recently released PolyMath (Wang et al., 2025) dataset represents a significant step forward, yet its coverage is still limited to 18 only high-resource languages....

    arxiv.org/abs/2607.05992 · PDF

  12. 12

    InfluMatch: Frontier-Quality KOL Search at 4B-Model Cost

    Krittanon Kaewtawee, Petmongkon Pornpichitsuwan, Natchaya Temyingyong, Nutnicha Laplamoon, Wachiravit Modecrua,...

    cs.CL · cs.AI · cs.IR

    Matching influencers (KOLs) to free-form, multi-part Thai marketing criteria is today served either by keyword search over structured profiles, which misses semantic fit, or by prompting frontier LLMs over every candidate, which is accurate but slow and expensive. We present InfluMatch, a low-cost three-stage cascade -- retrieval $\rightarrow$ rerank $\rightarrow$ reason -- built entirely from small open-weight models: dense retrieval returns...

    arxiv.org/abs/2607.05968 · PDF

  13. 13

    Is Domain Adaptation Always Helpful? A Frozen-Backbone Study of Cross-Domain Sentiment Transfer

    Phat Tran, Artin Lahni, Pranav Kulkarni, Yaolun Zhang

    cs.CL · cs.LG

    Sentiment analysis with frozen pre-trained language model (PLM) backbones has become a common paradigm, yet the practical benefit of explicit domain adaptation remains unclear, particularly when backbones encode varying degrees of target-domain knowledge. We present a preliminary case study evaluating a controlled family of frozen embedding backbones (Qwen3-Embedding 0.6B, 4B, 8B), alongside RoBERTa-base and FinBERT. We train a lightweight...

    arxiv.org/abs/2607.05937 · PDF

  14. 14

    Mitigating Factual Hallucination in Large Reasoning Models via Mixed-Mode Advantage Regularization

    Kaishen Wang, Tong Zheng, Xuehao Cui, Ruibo Chen, Tianyi Xiong, Heng Huang

    cs.CL · cs.LG

    Large reasoning models (LRMs) improve language model capabilities by generating explicit thinking traces before final answers. In factuality-oriented question answering (QA), such thinking often improves overall performance by helping the model recover relevant knowledge and refine its answers. However, we find that this benefit is not uniform at the instance level: explicit thinking can also overturn correct non-thinking answers and lead to...

    arxiv.org/abs/2607.05861 · PDF

  15. 15

    When Should LLMs Search? Counterfactual Supervision for Search Routing

    Minho Kim

    cs.CL · cs.AI

    Search-augmented language models can use external evidence to compensate for limitations in parametric knowledge, but search is not uniformly beneficial: models may call search for questions they can already answer, or rely on noisy evidence when correction, clarification, or abstention would be more appropriate. We formulate this as an instance-level search-routing problem: deciding whether search is needed to improve task success relative...

    arxiv.org/abs/2607.05752 · PDF

  16. 16

    Where to cut, how deep: BPE and Unigram-LM on chemistry SMILES

    Hunter Heidenreich

    cs.CL · cs.LG · q-bio.BM

    Every chemical language model reading SMILES begins with a tokenizer, yet the field has inherited byte-pair encoding (BPE) from natural language with little scrutiny. In natural language, BPE's principal alternative, Unigram-LM, is known to build structurally different vocabularies. Whether that contrast survives in chemistry was open. We report a controlled comparison of BPE and Unigram-LM over a fixed 165-token chemistry base, at the small...

    arxiv.org/abs/2607.05691 · PDF

  17. 17

    RPAM: A Principled Metric for Evaluating Associations in Language Models with High Predictive Validity in Downstream Outputs

    Damian Hodel, Jevin West, Aylin Caliskan

    cs.CL · cs.AI

    Language models (LMs) exhibit problematic biases, such as stereotypes. Effectively analyzing and mitigating such biases requires accurate and generalizable evaluation methods of the underlying associations. Some existing approaches focus on downstream metrics that analyze associations in generated text. Since generated text content can vary drastically across LMs, such metrics often require specialized evaluation datasets, which limits the...

    arxiv.org/abs/2607.05679 · PDF

  18. 18

    Do It Right! A Methodology for Successful NLP System Development

    Olga V. Patterson, Brett South, T. Elizabeth Workman, Scott L DuVall

    cs.CL · cs.AI

    Natural language processing (NLP) is a common method for supplying data to clinical research and decision making by extracting information from electronic medical records. Numerous textbooks and tutorials describe specific algorithms and applications for text processing, yet algorithmic knowledge is only one ingredient of a successful NLP project. Drawing on the available literature, this paper presents a stepwise approach that applies the...

    arxiv.org/abs/2607.05644 · PDF

  19. 19

    BaFCo: A Document Understanding Benchmark for Complex Bangla Form Comprehension

    Abu Tyeb Azad, Ishita Sur Apan, Fahim Ahmed, Sumaiya Karim Katha, Ezharuddin Jubaer, Armun Alam, Pranjal Kumar...

    cs.CL · cs.AI · cs.CV

    Document comprehension is a challenging yet impactful task for Multimodal Large Language Models, especially as these systems see growing adoption in real-world, human-centric applications. However, this adoption is limited for low-resource languages such as Bangla due to the scarcity of high-quality annotated data. To address this gap, we introduce BaFCo, a benchmark dataset for Bangla form comprehension with a focus on Document Layout...

    arxiv.org/abs/2607.05614 · PDF

This edition is part of The Daily Abstract — cs.CL archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.