cs.CL · 2026-10-02 · No. 131

Computation and Language, 2026-10-02.

13 new papers in cs.CL. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

13 entries
  1. 01

    KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards

    Pengfei Li, Naufal Suryanto, Sicheng Zhang, Muzammal Naseer

    cs.CL · cs.AI · cs.CR

    LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agentic tasks, and do not directly measure LLMs' ability to generate executable commands for real-world cybersecurity tools. This gap is critical because cybersecurity operations rely on strict command-line interfaces (CLIs),...

    arxiv.org/abs/2610.02206 · PDF

  2. 02

    Hierarchical Continuous Diffusion Language Models

    Hui Ren, Zihan Li, Chang Liu, Huidong Liu, Alexander Schwing

    cs.CL · cs.AI · cs.LG

    Discrete diffusion language models offer a compelling alternative to autoregressive generation for tasks demanding bidirectional reasoning and global constraint satisfaction. Yet they share a structural bottleneck: when decoding in parallel, each token is sampled independently from its marginal, severing the statistical dependencies among the tokens decoded together. Continuous diffusion language models avoid this by denoising a shared...

    arxiv.org/abs/2610.02193 · PDF

  3. 03

    From Knowledge Access to Source Learning: Developing Source-Specific Competence

    Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin, Jinjin He, Xiyuan Yang, Haoxin Liu, Ye Yu, Haibo Jin, Yijia Xiao,...

    cs.CL · cs.AI · cs.LG

    Large language model (LLM) agents increasingly rely on persistent external sources to solve sequences of knowledge-intensive tasks. Existing methods improve how source content is accessed and organized, while agent-memory systems preserve reusable knowledge from prior interactions, but repeated use of the same source is still largely treated as repeated access rather than an opportunity to progressively improve understanding of that source....

    arxiv.org/abs/2610.02150 · PDF

  4. 04

    Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

    Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma

    cs.CL · cs.AI · cs.DB

    Real-world enterprise data science and analytics workflows require reasoning across dozens of tables, performing statistical analyses, and acting on the results. Established text-to-SQL benchmarks evaluate query generation alone, and audits have found their answer keys frequently wrong. Because real enterprise warehouses are too sensitive to release, these benchmarks are built on public datasets where a business event fits in a single table....

    arxiv.org/abs/2610.02122 · PDF

  5. 05

    Scalable, Transferable Meta-network for Data Selection Requires a Different Loss (and Why the Obvious Choice is Problematic)

    Zilin Du, Bowen Yang, Boyang Albert Li

    cs.CL · cs.AI · cs.LG

    Data selection is critical for training large language models on massive and heterogeneous corpora. Meta-learning for Training-data Selection offers a principled alternative to heuristic scoring by learning data weights from a target validation objective, but existing methods face a trade-off between fine-grained valuation and transferability to unseen data. A natural solution is to replace per-sample weights with a selection network....

    arxiv.org/abs/2610.02092 · PDF

  6. 06

    Mem++: Non-Destructive Memory for Long-Term Organizational LLM Agents

    Ahmad Yehia, Aly O. Abdelkareem, Islam Ahmed, Hesham Omran, Khaled Alashmouny, Christian Claudel, Abduallah Mohamed

    cs.CL · cs.AI

    Large Language Model (LLM) agents now take part in organizational work, where many authors record decisions across documents over months. Because a revised decision arrives as a new document rather than an edit, answering a question requires knowing which version held at a given time. However, most memory systems compress the record at write time. By distilling each document into facts, notes or graph edges, these methods fix what can be...

    arxiv.org/abs/2610.02002 · PDF

  7. 07

    Universal Byte-Level Encoding: UTF-8/UTF-16 Routing to Reduce Cross-Script Token-Budget Disparities

    Hyunsik Kim, Youngmoon Jung

    cs.CL · cs.LG

    Byte-level byte-pair encoding (BBPE) tokenizers are attractive for multilingual large language models (LLMs) because they cover all Unicode text. In UTF-8-based BBPE, however, many scripts start from a higher fallback cost than English: when no learned merges can be applied, a multibyte character requires multiple byte-derived symbols. We call this worst-case pre-merge cost the encoding floor. A higher floor can increase token counts and...

    arxiv.org/abs/2610.01984 · PDF

  8. 08

    A rubric landscape for evaluating clinical reasoning in large language models: what exists, what is missing, and what needs to be combined

    Zhangshu Joshua Jiang, Zina Ibrahim, James T. Teo

    cs.CL · cs.AI

    Exam-style accuracy does not establish whether large language models (LLMs) reason well over clinical records. We define clinical reasoning as integrating and updating evidence across time and sources to form, revise and justify a patient's problem representation and a defensible plan. This structured narrative review maps three literatures: medical education assessment instruments, clinical LLM benchmarks published from 2023 onwards, and...

    arxiv.org/abs/2610.01938 · PDF

  9. 09

    Cross-Lingual Alignment for Decoder-Only Models using MoE Routers

    Lucas Bandarkar, Clark Peng, Ahmed Haj Ahmed, Aditi Khandelwal, Nanyun Peng

    cs.CL · cs.AI

    Cross-lingual contrastive learning has been a core component of multilingual encoder training, but the ability to explicitly align representations is not possible in decoder-only LLMs because of varying multilingual tokenization. However, a growing amount of research suggests that even in LLMs, higher cross-lingual representational alignment leads to improved cross-lingual transfer. In this paper, we propose a novel approach to reimagine...

    arxiv.org/abs/2610.01921 · PDF

  10. 10

    What Makes Something Hard(er)? Explaining Question Difficulty in Natural Language

    Peng Cui, Qiaoyuan Zheng, Rudolf Debelak, Mrinmaya Sachan

    cs.CL · cs.AI

    Difficulty is one of the most fundamental properties of a question: it determines whether the question can meaningfully discriminate between models of differing ability. Although a variety of methods can now estimate or predict difficulty automatically, they yield only a single descriptive number, with no account of the underlying factors that make a question difficult in the first place. In this work, we propose a data-driven approach that...

    arxiv.org/abs/2610.01627 · PDF

  11. 11

    Can LLMs Reliably Annotate Bioassay Metadata to Improve Data Readiness?

    Laura van Weesep, Riccardo Tedoldi, Jens Sjölund, Hossein Azizpour, Susanne Winiwarter, Ola Engkvist, Jon Paul...

    cs.CL · cs.AI · cs.DB · q-bio.QM

    The emergence of foundation models for molecular property prediction requires a high degree of AI data readiness, including reliable metadata annotation. However, both public repositories and industrial screening databases suffer from missing, inconsistent, or conflated assay annotations. In this work, we quantify the extent of missing annotations in PubChem for the BioAssay Ontology (BAO) assay format and physical detection method fields and...

    arxiv.org/abs/2610.01616 · PDF

  12. 12

    Which LLM to pick? Online Active Model Selection for Large Language Models

    Alessandro Turrin, Patrik Okanovic, Torsten Hoefler, Nezihe Merve Gürel

    cs.CL · cs.LG

    Large Language Models (LLMs) are increasingly applied to process streaming data, with practitioners relying on benchmarks to select the best model even though these signals only approximate real performance. While oracle annotations can provide reliable feedback, they are often costly and difficult to obtain at scale. To address this challenge, we propose ONLINE LLM PICKER, the first framework for active model selection for LLMs in online...

    arxiv.org/abs/2610.01592 · PDF

  13. 13

    AURAL: Adaptive Latent Reasoning with Joint Chunk for Speech Language Models

    Yuxiang Wang, Kunyu Feng, Yuancheng Wang, Zihang Liu, Shengbo Cai, Qinke Ni, Wan Lin, Tao Feng, Yingda shen,...

    cs.CL · cs.LG · cs.SD

    Model intelligence and fast response jointly shape the quality of interaction with speech language models, yet remain difficult to achieve together. Explicit chain-of-thought (CoT) improves reasoning and audio understanding, but generating intermediate reasoning tokens delays responses. Describing fine-grained acoustic cues further lengthens CoT and increases latency. Latent reasoning can reduce this overhead, yet existing methods often trail...

    arxiv.org/abs/2610.01560 · PDF

This edition is part of The Daily Abstract — cs.CL archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.