cs.CL · 2026-06-18 · No. 27

Computation and Language, 2026-06-18.

18 new papers in cs.CL. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

18 entries
  1. 01

    Freeing the Law with LOCUS: A Local Ordinance Corpus for the United States

    Denis Peskoff, Joe Barrow, Christopher Vu, Diag Davenport

    cs.CL · cs.CY · cs.LG

    Progress in legal AI increasingly depends on access to authoritative legal text at scale. Yet one of the most consequential layers of American law remains largely absent from existing machine-readable corpora: local ordinances. Local codes govern zoning, housing, business licensing, public health, noise, animal control, and many other domains of everyday regulation, but they are fragmented across vendor platforms designed for human browsing...

    arxiv.org/abs/2606.19334 · PDF

  2. 02

    Trade-offs in Medical LLM Adaptation: An Empirical Study in French QA

    Ikram Belmadani, Oumaima El Khettari, Carlos Ramisch, Frederic Bechet, Richard Dufour, Benoit Favre

    cs.CL · cs.AI

    The development of large language models (LLMs) has led to an increased focus on their adaptation to specialized domains and languages, yet the effectiveness of domain adaptation strategies remains unclear. We present a study of medical domain adaptation using French medical question-answering (QA) as a case study. We compare continual pretraining (CPT), supervised fine-tuning (SFT), and their combination across three model families, multiple...

    arxiv.org/abs/2606.19266 · PDF

  3. 03

    Language Models as Interfaces, Not Oracles: A Hybrid LLM-ML System for Pediatric Appendicitis

    Soheyl Bateni, Maryam Abdolali

    cs.CL · cs.AI

    Large language models (LLMs) can make clinical decision support more accessible by interpreting free-text documentation, but their direct use as diagnostic engines is limited by sensitivity to prompts, information order, and plausible but incorrect outputs. Structured machine-learning models offer more stable risk prediction, yet they require tabular inputs that are difficult to integrate with narrative clinical workflows. We present ClaMPAPP...

    arxiv.org/abs/2606.19183 · PDF

  4. 04

    Leadership as Coordination Control: Behavioral Signatures and the Recovery-Advantage Boundary in Multi-Agent LLM Teams

    Haewoon Kwak

    cs.CL · cs.AI · cs.MA

    Team science holds that leadership is contingent: it helps only under specific conditions, and capable, autonomous teams may need none at all. We ask the analogous question for multi-agent LLM teams: under what measurable conditions does process-level coordination control add value, and do those conditions match what team science predicts? We use behavioral signatures (majority lock-in, exploration, recovery from an incorrect round-0...

    arxiv.org/abs/2606.19111 · PDF

  5. 05

    Sumi: Open Uniform Diffusion Language Model from Scratch

    Mengyu Ye, Keito Kudo, Wataru Ikeda, Ryosuke Matsuda, Keisuke Sakaguchi, Jun Suzuki

    cs.CL · cs.LG

    Diffusion models have become a promising alternative to autoregressive models. Among these, uniform diffusion language models (UDLMs) permit any token to be updated at any step, in principle enabling more flexible generation. However, no UDLM has yet been pretrained from scratch at both large parameter scale and large token budget. Both autoregressive modeling and masked diffusion modeling already have capable models at scale that the...

    arxiv.org/abs/2606.19005 · PDF

  6. 06

    G-IdiomAlign: A Gloss-Pivoted Benchmark for Cross-Lingual Idiom Alignment

    Fengying Ye, Yanming Sun, Runzhe Zhan, Zheqi Zhang, Lidia S. Chao, Derek F. Wong

    cs.CL · cs.AI

    Idioms are difficult to transfer across languages due to their non-compositionality and weak surface-form grounding, making literal mappings unreliable. We present G-IdiomAlign, a gloss-pivoted benchmark where each idiom is anchored by an English gloss from Wiktionary. We further construct a high-confidence reference alignment set for reproducible evaluation. G-IdiomAlign supports two protocols: (1) a controlled Multiple-Choice Idiom...

    arxiv.org/abs/2606.18989 · PDF

  7. 07

    Beyond Tokenization: Direct Timestep Embedding and Contrastive Alignment for Time-Series Question Answering

    Yafeng Wu, Huu Hiep Nguyen, Thin Nguyen, Hung Le

    cs.CL · cs.AI

    Recent advances in large language models (LLMs) have given rise to time-series question answering (TSQA), which formulates time-series analysis as natural-language question answering. However, directly feeding raw numerical series into LLMs suffers from a tokenization bottleneck: Byte Pair Encoding fragments continuous values into unstable tokens whose embeddings lack meaningful metric structure, resulting in the loss of magnitude, scale, and...

    arxiv.org/abs/2606.18986 · PDF

  8. 08

    As Easy as Rocket Science: Assessing the Ability of Large Language Models to Interpret Negation in Figurative Language

    Jasmine Owers, Edwin Simpson, Martha Lewis

    cs.CL · cs.AI

    Figurative language and negation are two areas that challenge current language models, however, both are widely used throughout written and spoken language. Large language models (LLMs) are also widely used in everyday contexts where they cannot necessarily be tuned for a specific dataset. It is therefore essential to understand the ability of LLMs to correctly interpret text that includes both negation and figurative language. To investigate...

    arxiv.org/abs/2606.18922 · PDF

  9. 09

    Approximate Structured Diffusion for Sequence Labelling

    Nicolas Floquet, Joseph Le Roux, Nadi Tomeh

    cs.CL · cs.LG

    Sequence labelling, a core task of Natural Language Processing (NLP), consists in assigning each token of an input sentence a label. From a Machine Learning point of view, sequence labelling is often cast as a Linear-Chain Conditional Random Field (CRF) parametrised by a neural network. While this approach gives good empirical results, CRFs assume a finite decision span (eg label bigrams) which can limit their expressivity and hurt...

    arxiv.org/abs/2606.18856 · PDF

  10. 10

    Aligning Implied Statements for Implicit Hate Speech Generalizability with Context-Bounded Semi-hard Negative Mining

    Wicaksono Leksono Muhamad, Yunita Sari

    cs.CL · cs.AI

    Classifying implicit hate speech remains a challenge, as intent is often masked through insinuation and context rather than explicit slurs. Prior supervised contrastive approaches improve in-domain detection but can overfit surface cues and struggle to transfer across datasets. We propose ImpSH, a triplet-based framework that aligns posts with implied statements when available and uses context-bounded semi-hard negatives to focus learning on...

    arxiv.org/abs/2606.18852 · PDF

  11. 11

    Beyond Reward Engineering: A Data Recipe for Long-Context Reinforcement Learning

    Xiaoyue Xu, Sikui Zhang, Xiaorong Wang, Xu Han, Chaojun Xiao

    cs.CL · cs.AI

    Long-context reasoning is an essential capability for large language models, particularly when they are deployed as autonomous agents that must reason over lengthy trajectories. Reinforcement learning (RL) has recently emerged as a dominant paradigm for improving this ability, yet existing work largely focuses on reward engineering while diverse training data remains scarce. We revisit this problem from a data-centric perspective and show...

    arxiv.org/abs/2606.18831 · PDF

  12. 12

    RedactionBench

    Sean Brynjólfsson, Shashvat Jayakrishnan, Esha Sali, Diptanshu Purwar, Madhav Aggarwal

    cs.CL · cs.AI

    Large Language Models are increasingly applied to sensitive domains that require redaction of personally identifiable information (PII). While redacting PII is a data cleaning prerequisite, existing benchmarks conflate extraction mechanics with privacy semantics. A public phone number is not equivalent to a phone number in a medical record. Whether information constitutes a violation depends heavily on who holds it, why, and in what context,...

    arxiv.org/abs/2606.18782 · PDF

  13. 13

    Morpheus: A Morphology-Aware Neural Tokenizer and Word Embedder for Turkish

    Tolga Şakar

    cs.CL · cs.AI

    Turkish is agglutinative: meaning is carried by morphemes, yet the subword tokenizers that drive modern language models split words by corpus statistics, fragmenting semantically loaded suffixes and -- in the case of WordPiece and rule-based analyzers -- failing to decode their output back to the original text. This paper presents \textbf{Morpheus}, a neural morpheme-boundary model for Turkish that is at once a lossless, morphology-aware...

    arxiv.org/abs/2606.18717 · PDF

  14. 14

    TW-LegalBench: Measuring Taiwanese Legal Understanding

    Fei-Yueh Chen, Chun Huang Lin, Chan Wei Hsu, Kuan Hsuan Yeh, Zih-Ching Chen, Kuan-Ming Chen, Patrick Chung-Chia Huang

    cs.CL · cs.AI · cs.IR

    Large language models (LLMs) have shown impressive capabilities across diverse tasks, yet their performance on jurisdiction-specific legal reasoning remains underexplored. We present TW-LegalBench that utilizes Taiwanese legal system's rich official corpus open to the public to fill the gap in evaluating LLMs on Taiwanese law, among common-law benchmarks that focus on English sources and civil-law benchmarks focusing on sources of Simplified...

    arxiv.org/abs/2606.18699 · PDF

  15. 15

    PEC-Home: Interpretation of Progressively Elliptical Commands in Smart Homes

    Yingyu Shan, Zeming Liu, Silin Li, Boao Qian, Jiashu Yao, Yuhang Guo, Haifeng Wang

    cs.CL · cs.AI

    Recent advancements in Large Language Models (LLMs) have empowered home assistants with natural language interaction capabilities. However, current assistants overlook the progressive omission that occurs in human dialogue as shared context accumulates, leading to more elliptical expressions for efficient communication. Thus, current assistants still struggle to interpret such elliptical expressions accurately, which limits their...

    arxiv.org/abs/2606.18636 · PDF

  16. 16

    BCL: Bayesian In-Context Learning Framework for Information Extraction

    Haoliang Liu, Chengkun Cai, Xu Zhao, Han Zhu, Shizhou Huang, Xinglin Zhang, Tao Chen, Jenq-Neng Hwang, Zhang Huaping, Lei Li

    cs.CL · cs.AI

    Existing information extraction (IE) tasks increasingly adopt in-context learning (ICL) with large language models. However, current approaches either show inconsistent performance across model scales or lack systematic optimization and generalizability. Building on this, we propose BCL (Bayesian In-Context Learning Framework for Information Extraction), the first optimization framework that uses particle filtering with Bayesian updates to...

    arxiv.org/abs/2606.18620 · PDF

  17. 17

    Are LLMs Ready to Assist Physicians? PhysAssistBench for Interactive Doctor-Patient-EHR Assistance

    Tianming Du, Peijie Yu, Sihan Shang, Danli Shi, My Linh Nguyen, Shengbo Gao, Guangyuan Li, Yinghong Yu, Yan Jiang,...

    cs.CL · cs.AI

    The most plausible near-term role of medical LLMs is to assist rather than replace physicians, yet current evaluations often test isolated capabilities: clinical knowledge, EHR system interaction, or patient communication. Physician assistance instead requires coordinating these capabilities within the same interaction, where physicians issue underspecified requests, patients describe symptoms ambiguously, and EHR systems demand precise tool...

    arxiv.org/abs/2606.18613 · PDF

  18. 18

    Steerable Cultural Preference Optimization of Reward Models

    Minsik Oh, Advit Deepak, Sophie Wu, Douwe Kiela, Ekaterina Shutova

    cs.CL · cs.AI

    It is essential for large language model (LLM) technology to serve many different cultural sub-communities in a manner that is acceptable to each community. However, research on LLM alignment has so far predominantly focused on predicting a unified response preference of annotators from certain regions. This paper aims to advance the development of alignment models with a more global outlook, that are able to accurately represent the...

    arxiv.org/abs/2606.18606 · PDF

This edition is part of The Daily Abstract — cs.CL archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.