cs.CL · 2026-08-23 · No. 93

Computation and Language, 2026-08-23.

14 new papers in cs.CL. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

14 entries
  1. 01

    G-CARL: Grounded Checklist-Aligned Reward Learning for Patient-Oriented Medical Report Interpretation

    Shiao Xie, Siyu Chen, Jianwei Lv, Bo Yuan, Yujin Wang, Xiandong Li

    cs.CL · cs.AI · cs.CV

    Personalized interpretation of medical reports has emerged as an increasingly important need among patients. Addressing this need requires both evidence-grounded medical factuality and context-dependent patient communication, yet existing medical vision-language tasks do not adequately capture these dual requirements. To bridge this gap, we introduce Patient-oriented Medical Report Interpretation (PMRI), a novel open-ended multimodal...

    arxiv.org/abs/2608.20331 · PDF

  2. 02

    Inducing Task Models from Computer-Use Traces

    Yucheng Jiang, Zora Zhiruo Wang, Ruishi Chen, Diyi Yang

    cs.CL · cs.AI

    Naturalistic computer-use traces, passively recorded screenshots and mouse or keyboard actions, are a valuable resource for deriving symbolic, auditable, and reusable models of how everyday work is done. Such models matter as computer-use agents enter real work, where agents need to learn how tasks are actually performed, and organizations need to audit and reuse that knowledge. However, inducing such task models is challenging, as activity...

    arxiv.org/abs/2608.20319 · PDF

  3. 03

    Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization

    Qian Kou, Xiaofeng Shi, Xiaosong Qiu, Hua Zhou

    cs.CL · cs.AI

    Large language models often fail to answer questions about a bounded document collection when the source documents are not retrieved at inference time. We study this setting as document knowledge internalization: converting a fixed corpus into usable parametric knowledge for retrieval-free question answering. We propose IAR (Inject, Align, and Recover), a three-stage post-training framework that separates structured document knowledge...

    arxiv.org/abs/2608.20281 · PDF

  4. 04

    Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection

    Atsuyuki Miyai, Kiyoharu Aizawa, Toshihiko Yamasaki

    cs.CL · cs.AI · cs.LG

    We present a novel approach to efficient LLM agent harness optimization through adaptive validation task selection. Harness optimization iteratively rewrites the harness code based on validation performance, enabling substantial performance gains without updating the underlying model weights. Existing approaches, however, evaluate a fixed validation set in full at every iteration, incurring substantial evaluation costs even on tasks that...

    arxiv.org/abs/2608.20169 · PDF

  5. 05

    SABET-QA: Temporal Knowledge Graph Question Answering

    Brahim Touayouch, Mirette Moawad, Dmitry Akulov

    cs.CL · cs.AI

    Question Answering over Temporal Knowledge Graphs (TKGQA) requires reasoning over time-sensitive facts, yet existing embedding-based methods struggle with multi-step queries due to single-pass reasoning pipelines. We propose SABET-QA, a framework that iteratively refines reasoning states across multiple hops via a bidirectional entity-temporal scoring mechanism and a slot-aware contextualization module that aligns question semantics with...

    arxiv.org/abs/2608.20083 · PDF

  6. 06

    Auditing Cross-Lingual Fairness in Language Model Watermarking

    Alexander Nemecek, Osama Zafar, Debargha Ganguly, Vikash Singh, Vipin Chaudhary, Erman Ayday

    cs.CL · cs.CR · cs.LG

    Watermarking schemes for large language model output are evaluated almost exclusively on English text using each scheme's detection threshold and a narrow set of quality measurements. Multilingual deployment exposes evaluation-design choices that are inconsequential on English but determine conclusions cross-lingually. We propose an evaluation framework with four components: detection thresholds calibrated empirically per deployment context,...

    arxiv.org/abs/2608.20047 · PDF

  7. 07

    Interrupting the Loop: Periodic Subject Changes Raise Judged Surprise and Connection in Base Language Models

    Roberto I. Ono Filho

    cs.CL · cs.AI

    Where does the novelty a base language model produces with no task come from, and what can an LLM judge of a long stream actually see? We dismantle a cognitively inspired generation loop over 24 conditions on three base models. Most of its effect lives in one operation: a new subject injected every few hundred tokens (an interruption) into a stream whose literal repetition is damped (habituation). We judge windows of generated text only, with...

    arxiv.org/abs/2608.19893 · PDF

  8. 08

    A knowledge-guided agentic framework for mitigating patient-context ambiguity in health queries

    Mahyar Abbasian, Saba A. Farahani, Arshia Ilaty, Hung Cao, Ramesh Jain, Amir M. Rahmani

    cs.CL · cs.AI

    Patients often submit short, underspecified queries to healthcare chatbots that lack the patient-specific information needed to determine an appropriate response. Although these queries may be linguistically clear, they can support multiple plausible answers depending on undisclosed factors such as symptoms, diagnoses, medications, allergies, or dietary restrictions. A language model answering such a query directly may therefore rely on...

    arxiv.org/abs/2608.19875 · PDF

  9. 09

    LoRA-GA$^2$: Low Rank Adaptation with Multi-step Gradient Adaptive Alignment

    Haonan He, Xinyue Fan

    cs.CL · cs.AI

    Low-Rank Adaptation (LoRA) is a prominent fine-tuning method for large models, achieving competitive performance with reduced memory overhead. However, a persistent performance gap remains between LoRA and full fine-tuning. Recent studies have sought to narrow this gap by employing one-step gradient approximations of pretrained weights to align LoRA updates with the principal directions or intrinsic dimensionalities of full fine-tuning...

    arxiv.org/abs/2608.19800 · PDF

  10. 10

    Projector Is All You Train

    Nyx Iskandar, Saathvik Selvan, Slater Victoroff

    cs.CL · cs.CV · cs.LG

    The typical training process of a multimodal large language model (MLLM) involves adapting both the language model backbone and the projector between the backbone and a modality-specific encoder. We ask whether fine-tuning the backbone of an MLLM is necessary to adapt it to a new modality. Through experiments on 3D MLLMs, we find that training only the projector is sufficient to achieve strong multimodal performance relative to existing...

    arxiv.org/abs/2608.19726 · PDF

  11. 11

    Forking Fast: Efficiently Estimating Uncertainty Dynamics in Text Generation

    Eric Bigelow, Amir Zur, Satchel Grant, Tal Haklay, Can Rager, Owen Lewis, Thomas McGrath, Jack Merullo, Ekdeep Singh...

    cs.CL · cs.AI · cs.LG

    LLM reasoning is stochastic, and so understanding a model requires grappling with the distribution of reasoning chains that it might produce for a given question, i.e., its uncertainty. Resampling-based analyses characterize this distribution, revealing which steps of a rollout determine how the model arrives at its answer. However, a major limitation of these approaches is that resampling text sequences at every token or sentence in a...

    arxiv.org/abs/2608.19611 · PDF

  12. 12

    When Machines Speak: A Unified Generative Framework for Integrating Machine-Native Symbols into Pretrained Large Language Models

    Su Yan, Rakesh Iyer

    cs.CL · cs.AI

    Many real-world AI systems represent entities, behaviors, and structured information using discrete machine-native symbols rather than natural language. While these representations are compact and preserve task-relevant structure, they lie outside the linguistic token space of pretrained large language models (LLMs), creating a fundamental divide between language modeling and structured prediction. We introduce UniLang, a unified generative...

    arxiv.org/abs/2608.19529 · PDF

  13. 13

    Automated Summarization of Financial News Using Large Language Models and Retrieval-Augmented Generation: An Early Empirical Study (Fall 2023)

    Pranav Chandaliya

    cs.CL · cs.AI · cs.IR

    Stock market analysts and investors face a daily challenge: too much financial news, too little time. Manually reading and synthesizing hundreds of company-specific articles is impractical, yet missing key information can directly affect investment decisions. This project, conducted at George Washington University in Fall 2023, explores whether Large Language Models can automate this process reliably. We built a pipeline that pulls news...

    arxiv.org/abs/2608.19526 · PDF

  14. 14

    Are LLMs becoming similarly creative? Evidence from three years of models

    Nirav Patel, Josiah Crossman, Eva Aggarwal, Emily Wenger

    cs.CL · cs.AI · cs.CY

    Many benchmarks track Large Language Model (LLM) performance on tasks with verifiable answers, but less is known about how LLM performance is evolving on open-ended tasks, where creativity, originality and diversity may matter as much as quality. As LLMs increasingly support human ideation and creative work, understanding trends in LLM performance on open-ended tasks is critical. This paper presents a preliminary analysis of LLM creative...

    arxiv.org/abs/2608.19437 · PDF

This edition is part of The Daily Abstract — cs.CL archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.