cs.CL · 2026-07-03 · No. 42

Computation and Language, 2026-07-03.

13 new papers in cs.CL. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

13 entries
  1. 01

    LACUNA: A Testbed for Evaluating Localization Precision for LLM Unlearning

    Matteo Boglioni, Thibault Rousset, Siva Reddy, Marius Mosbach, Verna Dankers

    cs.CL · cs.AI · cs.LG

    LLMs memorize sensitive training data, including personally identifiable information (PII), creating a pressing need for reliable post hoc removal methods. Unlearning has emerged as a promising solution, with state-of-the-art(SOTA) methods often following a localize-first, unlearn-second paradigm that targets specific model parameters. However, existing benchmarks evaluate unlearning solely at the output level, leaving open the question of...

    arxiv.org/abs/2607.02513 · PDF

  2. 02

    Reasoning LLM Improves Speaker Recognition in Long-form TV Dramas

    Yuxuan Li, Lingxi Xie, Xinyue Huo, Jihao Qiu, Jiacheng Shao, Pengfei Chen, Jiannan Ge, Kaiwen Duan, Qi Tian

    cs.CL · cs.AI · cs.CV

    Long-form TV dramas present a formidable challenge for comprehensive video understanding, where deciphering complex storyline often relies on \textbf{speaker recognition}, the task of accurately attributing each spoken utterance to its respective character. In this paper, we advance this field through two primary contributions. (1) We introduce \textbf{DramaSR-532K}, a large-scale benchmark comprising 532K annotated dialogue lines across more...

    arxiv.org/abs/2607.02504 · PDF

  3. 03

    World Wide Models: Literary Tools for Cultural AI

    Nina Begus

    cs.CL · cs.AI

    LLMs stage a new form of cultural encounter that is massive, automated, and monolingual. Literary disciplines have always negotiated cultural struggles with comparative reading of literature, narratological and poetic analysis, critical theory, world literature, and translation. These tools have now become indispensable for building culturally literate AI. The essay develops a layered framework toward more nuanced textual models and...

    arxiv.org/abs/2607.02369 · PDF

  4. 04

    On the Role of Directionality in Structural Generalization

    Zichao Wei

    cs.CL · cs.LG

    Several SLOG test categories explicitly involve directional distinctions (modifier position shifts, argument extraction positions), yet AM-Parser, the previous SOTA, uses an AM algebra whose operations do not encode direction. We redesign the symbolic backend around CCG directed types (deterministic CKY + single linear decoder, 30K learnable parameters). Under the same BERT-base encoder, the system achieves 75.9$\pm$6.4% LF exact match,...

    arxiv.org/abs/2607.02307 · PDF

  5. 05

    Challenges and Recommendations for LLMs-as-a-Judge in Multilingual Settings and Low-Resource Languages

    A. Seza Doğruöz, Xixian Liao, Verena Blaschke, Jakob Prange, Senyu Li, David Ifeoluwa Adelani

    cs.CL · cs.AI

    LLM-as-a-Judge has become the dominant evaluation paradigm for many natural language generation tasks, due to shortcomings of conventional metrics and high correlations with human judgment, albeit mostly in English. There are now attempts to extend LLM-as-a-Judge to multilingual settings including low-resource languages. However, LLMs have limited proficiency in low-resource languages, and there is often no adequate human validation in these...

    arxiv.org/abs/2607.02235 · PDF

  6. 06

    HaloGuard 1.0: An Open Weights Constitutional Classifier for Multilingual AI Safety

    Navaneeth Sangameswaran, Preetham S, Ashmiya Lenin

    cs.CL · cs.CR · cs.LG

    We present HaloGuard 1.0, an open-weights implementation of the constitutional-classifier paradigm for input safety. It achieves state-of-the-art performance on English and multilingual prompt-safety benchmarks at roughly one-tenth the model size of current leading open guard models. The safety constitution is the organising structure of the corpus: a natural-language constitution of 46 policies and 2,940 subcategories drives synthetic data...

    arxiv.org/abs/2607.02079 · PDF

  7. 07

    SPLIT: Cross-Lingual Empathy and Cultural Grounding in English and Ukrainian LLM Responses

    Anna Chorna

    cs.CL · cs.AI · cs.CY

    Large Language Models are increasingly deployed in emotional-support contexts and crisis-related situations. Nevertheless, their cross-lingual abilities in these circumstances remain underexplored. Existing benchmarks emphasize multilingual performance but rarely examine crisis-related empathy and cultural grounding in low-to-mid-resource languages. We introduce SPLIT, a 500-prompt benchmark designed to evaluate LLM consistency in generating...

    arxiv.org/abs/2607.02049 · PDF

  8. 08

    OpenSafeIntent: Evaluating Intent-Calibrated Safe Completion Across Dual-Use Prompt Sets

    Rheeya Uppaal, Seungwoo Lyu, Selina Sung, Junjie Hu

    cs.CL · cs.AI

    Safe completion requires models to provide useful assistance without enabling harm, but this behavior is difficult to evaluate with isolated prompts. We introduce OpenSafeIntent, a benchmark of controlled prompt-sets that vary intent while holding the underlying task fixed. Each datapoint contains benign, dual-use, and malicious variants of the same task. This design lets us evaluate whether models calibrate assistance across intent shifts,...

    arxiv.org/abs/2607.02047 · PDF

  9. 09

    Object Aligner: A Configurable JSON Schema Similarity Score for Graphs, Applied to LLM Prompt Optimization

    Jan Drchal

    cs.CL · cs.AI · cs.LG

    Large language models (LLMs) are often asked to produce JSON conforming to a fixed schema, powering information extraction, tool calling, agentic planning, and knowledge-graph construction. Measuring how closely an output matches a gold reference is essential yet surprisingly hard: exact match is brittle, text similarity ignores structure, and an LLM judge is expensive, opaque, and non-deterministic. We address this with Object Aligner (OA),...

    arxiv.org/abs/2607.01972 · PDF

  10. 10

    Towards a Phonology-Informed Evaluation of Multilingual TTS

    Sneha Ray Barman, Neeraj Kumar Sharma, Shakuntala Mahanta

    cs.CL · cs.ET · cs.LG

    Neural TTS systems can sound natural across languages, but naturalness does not guarantee the preservation of sound contrasts that distinguish words from their grammatical forms. Standard metrics like MOS do not test for this. We propose a classifier-based framework that audits TTS output against language-specific phonological patterns using human speech as a benchmark. Testing Assamese advanced tongue root (ATR) vowel harmony with Meta's MMS...

    arxiv.org/abs/2607.01965 · PDF

  11. 11

    AIriskEval-edu: New Dataset for Risk Assessment in AI-mediated K-12 Educational Explanations

    Javier Irigoyen, Roberto Daza, Francisco Jurado, Julian Fierrez, Ruben Tolosana, Alvaro Ortigosa, Enrique Blas,...

    cs.CL · cs.AI · cs.DB

    This work introduces AIriskEval-edu-db2, a new dataset designed to train and evaluate auditors based on LLMs for an explainable pedagogical risk assessment in instructional content for grades K-12. The dataset comprises 1,639 explanations from 170 curated ScienceQA questions, covering science, language arts, and social sciences. For each question, the dataset includes an explanation written by a human teacher alongside 11 explanations...

    arxiv.org/abs/2607.01934 · PDF

  12. 12

    TUDUM: A Turkish-Thinking Reasoning Pipeline for Qwen3.5-27B

    Baran Bingol, Bahaeddin Turkoglu

    cs.CL · cs.AI

    This paper presents TUDUM (Türkçe Düşünen Üretken Model), a project pipeline for adapting a Qwen-family 27B thinking model toward Turkish reasoning. The central problem is not only to answer Turkish prompts in Turkish, but to make the explicit reasoning trace itself Turkish. A thinking model may translate a Turkish prompt into an English-centered internal or visible scratchpad, solve the problem mostly in English, and only localize the final...

    arxiv.org/abs/2607.01927 · PDF

  13. 13

    PARTREP: Learning What to Repeat for Decoder-only LLMs

    Andikawati P Widjaja, Yongjun Kim, Hyounghun Kim, Jaeho Lee

    cs.CL · cs.LG

    While decoder-only LLMs excel at a vast array of natural language tasks, it suffers from an asymmetric information flow induced by causal attention: later tokens are richer in contextual grounding than earlier ones. A simple and effective remedy is prompt repetition -- just appending a second copy of prompt before generation can redistribute grounding across positions and improve reasoning performance. However, full repetition of the original...

    arxiv.org/abs/2607.01792 · PDF

This edition is part of The Daily Abstract — cs.CL archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.