cs.CL · 2026-09-29 · No. 128

Computation and Language, 2026-09-29.

13 new papers in cs.CL. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

13 entries
  1. 01

    Telescopic Language Models

    Zhilin Guo, Boqiao Zhang, Hakan Aktas, Kyle Fogarty, Nursena Koprucu Aslan, Wenzhao Li, Canberk Baykal, Albert Miao,...

    cs.CL · cs.AI

    One deployed language model must often serve many compute budgets, yet serving each budget still means a separate training or compression run per point. We train a Telescopic Language Model (TLM) to be that continuum: a nested-capacity Transformer supervised by stochastic prefix supervision with a full anchor. At every step, one randomly truncated prefix of the capacity axis is trained against the full next-token target, alongside one...

    arxiv.org/abs/2609.35769 · PDF

  2. 02

    Improving Test-Time Scaling with Adaptive Looped Transformers

    Yichen You, Tianyu Fu, Aosong Feng, Xingtai Lv, Xuefei Ning, Ning Ding, Yu Wang

    cs.CL · cs.LG

    Looped transformers have demonstrated promising parameter efficiency by reusing layers for latent computation. Prior studies compare looped and non-looped models at matched parameters or per-token FLOPs. However, to the best of our knowledge, whether looping improves test-time scaling as outputs grow longer remains underexplored. Through post-training looped transformers, we study the accuracy-compute slope, measured as the accuracy gain per...

    arxiv.org/abs/2609.35748 · PDF

  3. 03

    Harness Learning Enables Generalizable Test-Time Adaptation

    Alvin Zhang, Xuecheng Liu, Zixuan Wang, Fahim Tajwar, Daman Arora, Ruslan Salakhutdinov, Daniel Khashabi, Yuda Song,...

    cs.CL · cs.LG

    A language-model agent is jointly defined by its model and its harness, the executable program that organizes model calls, tool use, and information flow. Because different tasks call for different ways of organizing these operations, the harness needs to be adapted using feedback from the task at hand. We introduce harness learning, which trains a proposer model to revise a solver's harness using execution feedback. We formulate this process...

    arxiv.org/abs/2609.35738 · PDF

  4. 04

    MS-GLA: Multi-Scale Gated Linear Attention for Addressing Representational Bottlenecks via Multi-Temporal Resolution

    Prasoon Dev, Anirudh Sankar, Vasudeva Varma

    cs.CL · cs.AI

    Gated Linear Attention (GLA) Transformers advance linear recurrent models through data-dependent gating, but face a core limitation: the fixed-capacity memory matrices across all heads operate at a single temporal resolution, where each token is processed individually, forcing them to simultaneously encode local syntactic patterns and long-range semantic structure, creating a representational bottleneck that gating alone is insufficient to...

    arxiv.org/abs/2609.35664 · PDF

  5. 05

    Which the Eye Fears: Writing with Read-Blindness Explains Massive Activations in Transformers

    Swagatam Mukhopadhyay, Vishal Vivek Saley, Vraj Parikh, Mausam

    cs.CL · cs.LG

    Massive activation features (MAs) in Transformers are extreme-value residual-stream features that persist across layers despite the model's ability to suppress them. Why do they survive? Our investigation using an operator-level mechanistic analysis of attention and feed-forward (FFN) blocks reveals that these blocks systematically ignore MA coordinates while reading, but not while writing; creating a read-write asymmetry that blocks...

    arxiv.org/abs/2609.35630 · PDF

  6. 06

    FactorEngram: Factorized N-gram Memory with Basis-Level Gating for Language Models

    Bowen Yang, Jingbo Zhou, Qinghong Miao, Hua Wu

    cs.CL · cs.AI

    Lookup-based memory has been a promising way to scale the parameters of large language models (LLMs). It retrieves learned representations of local token patterns, such as n-grams, instead of reconstructing them through successive layers of computation. However, existing designs such as Engram treat each retrieved embedding as a monolithic unit. Each embedding is stored in its own hashed slot and modulated by a single scalar gate. As a...

    arxiv.org/abs/2609.35578 · PDF

  7. 07

    Almieyar: A Culturally Grounded Benchmark for Multi-Dialect Arabic Speech Recognition

    Omid Ghahroodi, Anas Madkoor, Dima Faris Al Saudi, Fagr Tahir, Malak Annan, Talha shahid javad allah rakha, Omar...

    cs.CL · cs.AI

    Arabic speech technology has largely focused on Modern Standard Arabic, leaving the living dialects spoken by hundreds of millions under-served. We introduce ALMIEYAR, a culturally grounded ASR benchmark covering 17 Arabic dialects across six families, built entirely from newly recorded speech unseen by existing models. Dialect-community coordinators selected culturally relevant images across 10 topics, and native speakers described them...

    arxiv.org/abs/2609.35564 · PDF

  8. 08

    Less Sycophancy, Stronger Refusal? Lessons for AI Safety from Mechanistic Interpretability

    Xu Wang, Difan Zou, Xuansheng Wu

    cs.CL · cs.AI · cs.LG

    Reliable refusal of harmful requests is essential to the safe deployment of language models. Because excessive eagerness to please users may undermine existing refusal capabilities, reducing sycophancy offers a potential route to stronger refusal beyond the harmful scenarios covered by safety training. We investigate this possibility using compensatory feature injection (CFI), a training technique designed to limit the acquisition of a target...

    arxiv.org/abs/2609.35544 · PDF

  9. 09

    Beyond Token Scale: Chunk-Level Sparse Autoencoders for Reliable Semantic Feature Discovery

    Xu Wang, Yifan Yang, TingHao YU, Difan Zou

    cs.CL · cs.AI · cs.LG

    Sparse autoencoders (SAEs) expose features that help us understand and steer language models, but faithful reconstruction does not guarantee informative concepts. Token-level objectives reward lexical and formatting details alongside semantic content, all competing for a limited sparse budget. We introduce a family of chunk-level SAEs that encode mean-pooled activations over chunks, each a contiguous span of tokens: Mean-Chunk reconstructs...

    arxiv.org/abs/2609.35521 · PDF

  10. 10

    Spontaneous Context Restoration: How Language Models Recover from Corrupted Inputs

    Pranjal Garg, Jacob Beck

    cs.CL · cs.AI

    Language models sometimes produce correct outputs even when their inputs are corrupted by deletion, replacement, or misspelling. We study the internal processes accompanying this behavior, which we call context restoration, in controlled attention-only transformers and five pretrained LLMs (1B-32B parameters) across arithmetic, reading comprehension, and multiple-choice reasoning tasks. In the attention-only transformers, restoration emerges...

    arxiv.org/abs/2609.35475 · PDF

  11. 11

    AwarenessBench: Assessing Cognitive Capabilities of Language Models

    Xiaojian Li, Rongwu Xu, Tianyun Zhang, Yue Wang, Shuo Chen, Qiner Lyu, Briana Zhang, Peiran Yang, Kyle Xue Chen,...

    cs.CL · cs.AI

    As language models (LMs) exhibit increasingly consciousness-like behaviors, evaluating their cognitive abilities becomes essential. We introduce AwarenessBench, the first comprehensive benchmark for assessing the cognitive abilities of LMs in four dimensions: metacognition, self-awareness, social awareness, and situational awareness, covering 15 cognitive functions and 14,381 samples. Evaluating 18 state-of-the-art LMs, we find that all...

    arxiv.org/abs/2609.35409 · PDF

  12. 12

    Multilinguality in Hybrid Attention LLMs

    Lucas Bandarkar, Junlin Hu, Chenyuan Yang, Mohsen Fayyaz, Nanyun Peng

    cs.CL · cs.AI

    In response to the growing demand for long sequences in agentic and reasoning use cases, many state-of-the-art LLMs combine multiple variants of attention to mitigate the quadratic complexity of traditional softmax attention. These hybrid attention LLMs aim to balance the strengths and limitations of full attention and alternatives based on recurrence. This work presents a first study of how hybrid attention impacts the multilinguality of...

    arxiv.org/abs/2609.35378 · PDF

  13. 13

    From Normative Frameworks to Alignment Data: Constructing and Evaluating SFT and Preference Data

    Husrev Taha Sencar, Rezart Beka, Danish Naeem, Seda Ozalkan, Majd Hawasly, Ji Lucas, Ala AlFuqaha, Mohamed Abdallah,...

    cs.CL · cs.AI

    Aligning language models with a specified normative framework requires translating abstract principles into concrete examples and preference signals from which models can learn. We present an expert-driven methodology for constructing such alignment data and apply it to a normative framework grounded in Islamic ethical, theological, and jurisprudential traditions. Over approximately one year, seven domain experts systematically probed...

    arxiv.org/abs/2609.35201 · PDF

This edition is part of The Daily Abstract — cs.CL archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.