cs.CL · 2026-10-06 · No. 135

Computation and Language, 2026-10-06.

15 new papers in cs.CL. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

15 entries
  1. 01

    MemPilot: Orchestrating On-Demand Multimodal Memory Curation for LLM Agents

    Haozhen Zhang, Haodong Yue, Quanyu Long, Jianzhu Bao, Qingyuan Liu, Tao Feng, Bohan Liu, Weida Liang, Wenya Wang

    cs.CL · cs.AI · cs.LG

    Memory has become integral to the LLM agent ecosystem, supporting information retention and reuse across interactions. However, most existing agent memory systems construct memory in a query-agnostic manner, which can incur unnecessary preprocessing cost and discard details that later prove essential. Recent studies have begun shifting memory processing toward runtime adaptation, but typically specialize in particular operations or fixed...

    arxiv.org/abs/2610.06830 · PDF

  2. 02

    CLIFT: Conformal Self-Verification for Web Agent Training and Test-Time Scaling

    Yifan Zhang, Yutong Dai, Viraj Prabhu, Zhiyuan Hu, Ran Xu, Zeyuan Chen

    cs.CL · cs.AI · cs.LG

    Open-source web agents are now strong enough to execute realistic browser tasks, but training them with reinforcement learning still depends on weak supervision: binary task success is too sparse for credit assignment, while frontier-language-model judges are too expensive to call at every step and cannot be assumed available at deployment. We introduce CLIFT, a training and test-time scaling method built around conformal self-verification....

    arxiv.org/abs/2610.06829 · PDF

  3. 03

    IdeaLens: Detecting AI Ideas in Long-form Writing

    Rishanth Rajendhran, Minjoon Choi, Jenna Russell, Ramya Namuduri, Deniz Bölöni-Turgut, Marzena Karpinska, John...

    cs.CL · cs.AI · cs.LG

    While modern AI detectors identify who wrote the words, emerging policies on AI use increasingly hinge on a different question: who came up with the ideas? We introduce IdeaLens, a detector that identifies whether a document's ideas came from a human or AI (idea provenance), regardless of who wrote its words. To focus IdeaLens on ideas rather than prose, we represent documents as outlines: lists of items that each pair a discourse role with a...

    arxiv.org/abs/2610.06778 · PDF

  4. 04

    Balancing Memory Pathways: Analyzing and Improving Memory Utilization in Hybrid LMs

    Hyunji Lee, Joykirat Singh, Zaid Khan, Justin Chih-Yao Chen, Elias Stengel-Eskin, Alessandro Sordoni, Arman Cohan,...

    cs.CL · cs.AI

    Recurrent-attention hybrid language models (LMs), which interleave attention and recurrent layers, are increasingly used to combine the efficiency of the recurrent layers with the strong performance of attention layers. Prior work suggests that attention and recurrent layers offer complementary pathways to use past information: attention supports precise memory recall from earlier tokens, while recurrent layers support consolidation of...

    arxiv.org/abs/2610.06750 · PDF

  5. 05

    ufakzeka-karar: An Open Turkish Typed-Decision Model with Order-Invariant Option Scoring

    Sait Furkan Teke

    cs.CL · cs.LG

    ufakzeka-karar is an open Turkish decision model with 182,494,466 parameters. Given a Turkish text and questions of a fixed answer type (a choice, a level on an ordered scale, or yes or no), it returns a temperature-scaled probability for every option and an expected error that serves as a "not sure" signal, without generating text and in one CPU forward pass for up to ten options. Built on the lab's ufakzeka-1-base, its head scores each...

    arxiv.org/abs/2610.06744 · PDF

  6. 06

    Representation-Space MMD for Diffusion Language Models

    Ilya Drobyshevskiy, Ilia Sudakov, Maksim Semenov, Denis Kuznedelev, Maksim Ignatov, Pavel Temirchev, Nikita...

    cs.CL · cs.LG

    We introduce a post-training method for diffusion language models (DLMs) that minimizes Maximum Mean Discrepancy (MMD) between generated and reference distributions in the feature space of a frozen pretrained DLM. To estimate MMD, we retain contextual features at individual token positions, obtaining multiple observations per sequence from a single extractor pass. We optimize this objective using policy gradients for discrete models and...

    arxiv.org/abs/2610.06648 · PDF

  7. 07

    Word-Level Text Unmixing via Evidence-Preserving Ownership Routing with Language Models

    Jinglin He, Siyang Jiang, Lixing He, Guoliang Xing, Hongkai Chen

    cs.CL · cs.AI

    Text from multiple sources can become interleaved into a single sequence when attribution metadata is lost, such as overlapping speech transcripts, document reading flows, or concurrent agent streams. We formalize this challenge as Word-Level Text Unmixing: given an interleaved lexical stream and source count K, recover the original source sequences while preserving every word occurrence and its within-source order exactly. Directly...

    arxiv.org/abs/2610.06603 · PDF

  8. 08

    Anatomy of LLM Sycophancy: What a Flip Rate Hides

    Haonan Huang

    cs.CL · cs.AI · cs.LG

    A model under pushback can correct itself, capitulate, or hold, and one flip rate counts a correction and a capitulation alike. Using SycoLens, a modular replay protocol, we test how user pressure and evaluation settings shape measured flip rates. Each measurement is one stateless replay of an item, a committed answer, and one scripted user line in a fixed form. Every effect is read against a matched control with the line deleted. Pushback...

    arxiv.org/abs/2610.06522 · PDF

  9. 09

    SOL: Measuring Gaps between Text Distributions by Double Sliced Wasserstein Metrics

    Gregor Kornhardt, Moritz Piening, Jannis Chemseddine, Gabriele Steidl

    cs.CL · cs.LG · stat.ML

    Evaluating text generation requires measuring how well the generated distribution matches the data distribution. For autoregressive models, this is done by the perplexity. Diffusion and flow-based language models can only provide a likelihood bound, whose tightness differs between model families. Sample-based substitutes such as generative perplexity with entropy do not consider the distribution fit. We propose SOL, a distance between text...

    arxiv.org/abs/2610.06513 · PDF

  10. 10

    The Assistance Dilemma: Learning to Teach via Multi-Turn Reinforcement Learning

    Jakub Macina, Manu Kapur, Mrinmaya Sachan

    cs.CL · cs.AI

    Large language models (LLMs) trained to answer questions are natively poor at teaching. Reinforcement Learning (RL) against a simulated student is a promising approach to improve their pedagogy, but existing RL-trained tutors reward the student's success on the tutored problem with the tutor's words still in context. The reward is then easiest to raise by telling the student the answer, and a tuned penalty is needed to reduce telling. Drawing...

    arxiv.org/abs/2610.06446 · PDF

  11. 11

    Ontology Concept Overlap as a Training Signal: Knowledge-Grounded Reinforcement Learning for Clinical Question Answering

    Aditya Tanna, Abhishek Jindal

    cs.CL · cs.LG

    Reinforcement learning post-training for language models relies on two reward designs: human preferences (RLHF, DPO) and binary verifiers (RLVR). Clinical question answering fits neither. Near-correct answers differ by a single substituted entity, and no executable check decides clinical correctness. We instantiate a soft verifier from a maintained controlled vocabulary: UMLS Concept Unique Identifier overlap (via scispaCy, set-level F1)...

    arxiv.org/abs/2610.06360 · PDF

  12. 12

    Agentic schema-guided extraction of materials process knowledge from scientific literature

    Sameer Sadruddin, Jennifer D'Souza

    cs.CL · cond-mat.mtrl-sci · cs.AI · cs.DL · cs.ET

    Materials literature contains detailed experimental knowledge, but procedures, chemical entities and measurements remain difficult to aggregate because they are reported in heterogeneous forms and depend on process-specific context. We present SciKGExtract, a schema-guided framework that combines large-language-model extraction with chemical normalization and agent-based evaluation and refinement before knowledge-graph integration. We...

    arxiv.org/abs/2610.06322 · PDF

  13. 13

    DialectSentEval 2026: Arabic Dialect Sentiment Analysis and Swapping Shared Task

    Saad Ezzini, Shadi Abudalfa, Maram Alharbi, Salmane Chafik, Hind Alatawi, Mo El-Haj, Ahmed Abdelali, Osamah...

    cs.CL · cs.AI

    Sentiment analysis is a fundamental problem in Natural Language Processing (NLP). Standard sentiment classification for the Arabic language remains challenging due to the high volume of dialectal Arabic. To advance research in this area, this paper proposes the Shared Task on Sentiment Analysis and Swapping in Arabic Dialects (DialectSentEval), hosted with the Arabic Natural Language Processing Conference (ArabicNLP 2026). This shared task...

    arxiv.org/abs/2610.06298 · PDF

  14. 14

    From Abusive Language Classification to Sequence Labeling Identification

    Nicolas Zampieri, Ignacio Lopez, Manon Girard, Jeremy Auguste

    cs.CL · cs.LG

    Industrial content moderation must process massive message streams under tight latency constraints, yet most abusive language (AL) detection systems rely on sentence-level classification (ALC), which neither localizes abusive spans nor identifies who is targeted. We define Abusive Language Identification (ALI) as a sequence-labeling task that jointly extracts AL spans and target mentions, and assess whether this approach can be used for text...

    arxiv.org/abs/2610.06287 · PDF

  15. 15

    Cross-lingual Calibration of Pre-Generation Success Probes for Multilingual LLM Routing

    Andrea Paganelli, Stefano Civelli, Pietro Bernardelle, Gianluca Demartini

    cs.CL · cs.AI

    Pre-generation success probes estimate response correctness from a language model's hidden activations before decoding, enabling cost-aware routing. While prior work has demonstrated their utility primarily on English inputs, we study their reliability across languages along three dimensions: (1) whether they preserve the ranking of likely successes and failures (DISCRIMINATION); (2) whether they retain probabilities that match observed...

    arxiv.org/abs/2610.06216 · PDF

This edition is part of The Daily Abstract — cs.CL archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.