cs.CL · 2026-07-23 · No. 62

Computation and Language, 2026-07-23.

21 new papers in cs.CL. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

21 entries
  1. 01

    Generative AI floods and dilutes the market for books

    Tuhin Chakrabarty, Xinyue Liu, Jane C. Ginsburg, Paramveer Dhillon

    cs.CL · cs.AI · cs.CY

    Generative AI can produce book-length works of fiction at near-zero cost. These books are often dismissed as low-quality ``slop'' that buyers will ignore, and are assumed to carry little commercial weight. We test that assumption with full-text AI detection across 14,419 self-published genre-fiction books sold on Amazon from 2023 to 2026, matched to daily sales records through June 2026. None of these books disclose whether or not they...

    arxiv.org/abs/2607.20349 · PDF

  2. 02

    Sound Probabilistic Safety Bounds for Large Language Models

    Mahdi Nazeri, Anne-Kathrin Schmuck, Sadegh Soudjani, Alessandro Abate

    cs.CL · cs.AI

    We propose a novel framework for computing rigorous bounds on the probability that a large language model (LLM) generates harmful output to a given prompt. We study a new application of the Clopper-Pearson confidence intervals to obtain probably approximately correct (PAC) bounds for this problem. As our main technical contribution, we propose an algorithm that leverages features in the latent space to prioritize exploring branches in the...

    arxiv.org/abs/2607.20286 · PDF

  3. 03

    The Maskability Index: Predicting Task-Objective Alignment in Pretrained Language Models

    Ahmad Pouramini, Mahsa Afsharzadeh

    cs.CL · cs.AI

    Large-scale pretrained language models such as T5 and BERT have demonstrated strong capabilities for generating structured knowledge. However, their performance depends on how closely the prompting strategy matches the objectives used during pretraining. We introduce the Maskability Index (MI), a quantitative metric that estimates whether a knowledge relation is better suited to masked-style prompting or prefix-style prompting in few-shot...

    arxiv.org/abs/2607.20265 · PDF

  4. 04

    On the Systematic Challenges of Culturally Loaded Machine Translation: Dream of the Red Chamber as the Cultural Lens

    Yiming Wang, Jiayuan Di

    cs.CL · cs.AI

    Culturally loaded translation poses unique challenges for machine translation (MT), as meanings are deeply embedded in socio-cultural contexts beyond surface linguistic forms. Although large language models (LLMs) have enabled MT systems to achieve human-like quality in many scenarios, their ability to handle culturally loaded expressions remains underexplored. In this study, we systematically investigate the challenges posed by culturally...

    arxiv.org/abs/2607.20241 · PDF

  5. 05

    SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD

    Dongfang Li, Xiaodong Luo, Ruoyu Sun, Xuhui Chen, Linyuan Qiu, Jian Meng, Zhengxuan Lu, Yiting Wang, Yucheng Xie,...

    cs.CL · cs.AI

    Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and inefficient kernel execution. While most large-scale LLM training systems are built around GPU-based clusters, this report presents an end-to-end optimization practice on the Ascend NPU SuperPOD. Using the...

    arxiv.org/abs/2607.20145 · PDF

  6. 06

    Reinforcement Learning for Large Language Model Selective Evidence Adoption from Contaminated Retrieval Results

    Yanyu Chen, Yue Li, Yongyi Cui, Dongsheng Shi, Lichang Dai

    cs.CL · cs.AI

    Retrieval-augmented large language models frequently face contexts that interleave useful evidence with misleading statements or instruction-like content. Blanket refusal discards valid evidence, whereas uncritical adoption yields incorrect or unsafe answers. The ability to selectively adopt relevant information while rejecting deceptive or harmful content is therefore critical for reliable deployment in real-world retrieval settings. We...

    arxiv.org/abs/2607.20090 · PDF

  7. 07

    Language-Specific versus Cross-Lingual Knowledge Graphs for Implicit Aspect Identification in Arabic: A Comparative Study of Reasoning and Adaptation Strategies

    Lujain A. Alawwad

    cs.CL · cs.AI

    Aspect-based sentiment analysis (ABSA) in Arabic must recover both explicitly stated aspects and implicit aspects that are never named in the text. Implicit identification typically relies on an auxiliary knowledge source (e.g., a knowledge graph (KG)) linking opinion cues to aspect categories, but for a lower-resource language the practitioner faces a design choice: reuse a mature English KG through multilingual embeddings, or build a...

    arxiv.org/abs/2607.20056 · PDF

  8. 08

    TINY_SCHILLER: A Drop-In German Drama Corpus for Small Language Models

    Mark Schutera

    cs.CL · cs.AI · cs.DL · cs.HC

    tiny_schiller closes the small-language-model prototyping, fine-tuning, education, and research gap for German literary text, providing a single-file, drop-in counterpart to Karpathy's tiny_shakespeare. The available German literary corpora are larger and richer, but require parser engineering before a single line of training or fine-tuning code can run. tiny_schiller is a 2.07-megabyte single file of eleven public-domain Schiller dramas,...

    arxiv.org/abs/2607.19992 · PDF

  9. 09

    When Does Knowledge Distillation Hurt? Reliability-Aware Distillation for Low-Resource Language Summarization

    Dipto Sumit, Ankan Kumar Roy Srizon, Sadia Khair Rodela, Atia Haque Asha, Mourchona Afrin, Niloy Farhan, Farig Sadeque

    cs.CL · cs.AI

    Knowledge distillation (KD) is a standard approach for compressing sequence-to-sequence models, but its per-sample effects are rarely examined. On the BanSum Bangla summarization benchmark, we find that standard KD improves ROUGE-L by only +0.0003 over a cross-entropy baseline, and that approximately 51.3% of training samples are estimated to actively harm student validation loss under standard KD. We propose two complementary...

    arxiv.org/abs/2607.19956 · PDF

  10. 10

    Overview of FinMMEval 2026 Task 2: Multilingual Financial Short-Answer Question Answering

    Zhuohan Xie, Xueqing Peng, Georgi Georgiev, Dimitar Dimitrov, Yuyang Dai, Rania Elbadry, Vanshikaa Jani, Lingfei...

    cs.CL · cs.AI · cs.CE

    FinMMEval 2026 Task 2 evaluates short-answer financial question answering over multilingual evidence. Each final-test item pairs an English question with financial statements and news in English, Chinese, Japanese, Spanish, and Greek. Participating systems submit one concise answer per item in JSONL format. The final-test set contains 256 items, split evenly between easy and expert tiers; each tier contains four question templates...

    arxiv.org/abs/2607.19867 · PDF

  11. 11

    Overview of FinMMEval 2026 Task 1: Multilingual Financial Multiple-Choice Question Answering

    Zhuohan Xie, Yuyang Dai, Rania Elbadry, Vanshikaa Jani, Georgi Georgiev, Dimitar Dimitrov, Fan Zhang, Xueqing Peng,...

    cs.CL · cs.AI · cs.CE

    FinMMEval 2026 Task 1 evaluates multilingual financial multiple-choice question answering in English, Chinese, Arabic, and Hindi. The task tests whether systems can select the correct answer to finance questions involving domain terminology, numerical interpretation, and conceptual financial reasoning across languages and scripts. The final-test set contains 800 questions, with 200 questions per language; gold answers were withheld during...

    arxiv.org/abs/2607.19856 · PDF

  12. 12

    Sentence Splitter: Uncovering Latent Factual Structure for Self-Supervised Learning

    Ahmad Pouramini, Mahsa Afsharizadeh

    cs.CL · cs.AI

    This paper introduces Sentence Splitter, a self-supervised framework built upon a T5-based encoder--decoder architecture for uncovering the latent factual structure of natural language sentences. The proposed method identifies the semantic boundary between a descriptive prefix (head) and its factual completion (tail) by formulating sentence splitting as a discrete segmentation problem, where a sentence of length $N$ admits $N$ possible split...

    arxiv.org/abs/2607.19845 · PDF

  13. 13

    TriAgent: Divergence-Aware Multi-Agent Committees for Cost-Efficient Financial Sentiment Analysis

    Isabel Xu, Cynthia Xu, Rachel Ren, Cong Guo, Jiacheng Ding

    cs.CL · cs.CE · cs.DB · cs.LG

    Production LLM-based financial sentiment analysis faces a structural cost trap: most queries are trivially classifiable, yet expensive cloud reasoners process them all, and the bill scales linearly with user count. We present TriAgent, a multi-agent committee stratified by contextual granularity -- a word-level lexicon (VADER), a sentence-level domain transformer (FinBERT), and a cross-sentence reasoner (Qwen2.5, 0.5B-14B-4bit, with...

    arxiv.org/abs/2607.19794 · PDF

  14. 14

    Learning the Arabic Dialect Continuum as a Continuous Space: A Regression Approach to Speaker Origin Prediction

    Mohamed Aziz Khadraoui, Adel Ammar, Bilel Benjdira, Zahid Khan, Skander Turki, Wadii Boulila

    cs.CL · cs.AI · cs.CY · cs.LG · cs.NE

    We present a regression-based approach to Arabic dialect geolocation that models dialectal variation as a continuous geographic space rather than discrete categories. Speaker origin is predicted as continuous latitude-longitude coordinates using a hierarchical neural architecture that fuses frame-level XLS-R-300M and Whisper-large-v3 encoder representations with phonotactic descriptors through a Transformer encoder and a learnable...

    arxiv.org/abs/2607.19751 · PDF

  15. 15

    SLPO: Scaling Latent Reasoning via a Surrogate Policy

    Runyang You, Zhiyuan Liu, Yongqi Li, Wenjie Li

    cs.CL · cs.AI · cs.LG

    Reinforcement learning with verifiable rewards has become the predominant recipe for eliciting test-time scaling in explicit Chain-of-Thought reasoners. Yet this scaling path remains computationally costly, since every intermediate step must be decoded as a language token. Latent reasoning instead carries intermediate computation as continuous vectors and already matches or surpasses explicit CoT at far shorter horizons. Despite this promise,...

    arxiv.org/abs/2607.19691 · PDF

  16. 16

    Multi-Mask Diffusion Language Models for Few-Step Generation

    Sijin Chen, Yinuo Ren, Heyang Zhao, Ziheng Cheng, Quanquan Gu, Lexing Ying

    cs.CL · cs.LG

    Masked diffusion models (MDMs) are a promising family of language generators, but achieving high-quality few-step generation remains challenging. In MDMs, all forward trajectories collapse to a single fully masked state, leaving no terminal entropy for consistency-style few-step generation. While recent few-step alternatives based on uniform-state diffusion avoid this degeneracy, it becomes harder to distinguish clean tokens from noise than...

    arxiv.org/abs/2607.19686 · PDF

  17. 17

    Reference-Free Evaluation of Reasoning in Open-Ended Question Answering

    Guneet Singh Kohli, Yuxiang Zhou, Michael Sejr Schlichtkrull, Gregory E Dean, Maria Liakata

    cs.CL · cs.AI · cs.LG

    AI-generated answers in high-stakes domains are often fluent but difficult to verify, especially when they contain multi-step reasoning rather than a single final answer. We propose a reasoning-based, reference-free framework for auditing LLM-generated outputs. The method decomposes a generated reasoning trace into segments, labels local premise-target relations using Natural Language Inference (NLI), and organizes these relations into a...

    arxiv.org/abs/2607.19678 · PDF

  18. 18

    Adaptive Capitulation: A Structural Failure Mode of LLM Responses in Vulnerability Contexts

    Eunna Lee

    cs.CL · cs.AI · cs.MA

    Large language models operating in emotionally sensitive contexts face a structural trilemma: when users in vulnerable states request information that may reinforce maladaptive attribution, current response architectures resolve the tension through protective restriction, uninflected facilitation, or unintegrated co-presence of both imperatives -- each preserving one objective at the cost of the other. Administering a three-turn escalating...

    arxiv.org/abs/2607.19629 · PDF

  19. 19

    Scaling Laws for Hypernetwork-Based Knowledge Injection in Large Language Models

    Nischay Dhankhar, Dos Baha, Abulhair Saparov

    cs.CL · cs.LG

    Injecting factual knowledge into large language models (LLMs) reliably and at scale remains an open challenge. Hypernetworks provide a promising solution to large-scale knowledge injection. Although hypernetworks are typically applied for test-time adaptation, we explore their use in train-time knowledge injection, where, given a large corpus of facts, we train a hypernetwork to generate a fixed LoRA adapter that, when inserted into the...

    arxiv.org/abs/2607.19604 · PDF

  20. 20

    On the Computational Complexity of Structural Generalization

    Zichao Wei

    cs.CL · cs.LG

    Structural generalization has been measured repeatedly by several benchmarks, yet it has never been formally defined. We give a definition that translates the two premises (compositional structure and unbounded generalization) into mathematical language. The definition itself is neutral: a compiler that hard-codes the rules satisfies it just as well. But structural generalization becomes a scientific question only insofar as the capacity can...

    arxiv.org/abs/2607.19573 · PDF

  21. 21

    Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning

    Lizhe Fang, Weizhou Shen, Tianyi Tang, Yisen Wang

    cs.CL · cs.AI

    Large language models that generate step-by-step reasoning traces have achieved strong performance on complex tasks, and extending them to long-context settings has emerged as an important frontier. However, we identify a critical failure mode in this regime: \emph{repetitive copying}, where models extensively copy text from the input into their reasoning traces rather than productively solving the problem. We show that this behavior is...

    arxiv.org/abs/2607.19345 · PDF

This edition is part of The Daily Abstract — cs.CL archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.