cs.CL · 2026-08-15 · No. 85

Computation and Language, 2026-08-15.

17 new papers in cs.CL. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

17 entries
  1. 01

    LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure

    Fanfei Li, Jana Zeller, Manuel Prada-Corral, Thaddäus Wiedemer, Prasanna Mayilvahanan, Ryan Cotterell, Wieland Brendel

    cs.CL · cs.AI · cs.LG

    Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skill acquisition is difficult, as prior exposure to related content is hard to characterize. To address this challenge, we introduce LITTLECURRICULUM, a curated 88B-token pretraining corpus tailored to U.S. elementary school material, explicitly excluding concepts, facts, and vocabulary taught above Grade 5. Training a...

    arxiv.org/abs/2608.13545 · PDF

  2. 02

    DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data

    Peter Schneider-Kamp, Jacob Nielsen, Gianluca Barmina, Kenneth Enevoldsen, Lukas Galke Poech

    cs.CL · cs.AI

    Current large language model development relies on massive, often non-permissible datasets, creating a high barrier for researchers committed to open-source and ethically sourced data. We introduce Mimir v1, a 1-billion-parameter language model based on the Hierarchical Reasoning Model (HRM) architecture, that is trained from scratch and delivers highly competitive performance for English and sets a new state of the art for Danish using only...

    arxiv.org/abs/2608.13517 · PDF

  3. 03

    Toward a Gricean Retreat: Probing LLMs for Knowledge Boundaries and Referent Specificity

    Dananjay Srinivas, Saksham Khatwani, Maria Pacheco

    cs.CL · cs.AI

    When asked about entities outside their knowledge boundary, LLMs routinely fabricate plausible-sounding details rather than backing off to safer, more general claims. We frame this failure through a Gricean lens: a cooperative speaker who is uncertain about a referent retreats up the specificity hierarchy, trading informativeness for truthfulness. We ask whether LLMs have the ingredients to perform this retreat. Using a T-REx-based benchmark...

    arxiv.org/abs/2608.13484 · PDF

  4. 04

    Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity

    Irina Proskurina, Mayank Kumar, Oyindolapo O. Komolafe

    cs.CL · cs.AI

    Instruction-tuned language models achieve strong performance across a range of generation tasks, but have also recently been shown to exhibit verbalized overconfidence. In question answering, verbalized model overconfidence may be associated with the consistency of the generated supporting rationales. In this paper, we study whether corresponding changes in the lexical diversity of generated answer rationales accompany changes in model...

    arxiv.org/abs/2608.13430 · PDF

  5. 05

    It's How You Ask: Gender-Associated Linguistic Bias in LLMs

    Katherine Van Koevering, Anjalie Field

    cs.CL · cs.AI

    Professional communication is increasingly mediated by LLMs - but do these models serve all users equally? We show that when prompts contain linguistic features more commonly used by women (hedges, tag questions, collective reference), they systematically elicit shorter, less sophisticated, and less formal responses across three document types and four models. These effects persist after controlling for prompt complexity and feature...

    arxiv.org/abs/2608.13328 · PDF

  6. 06

    Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model

    Mohammed Sabry, Sean Augenstein, Keith Rush, Lucio Dery

    cs.CL · cs.AI

    We ask whether language-model pre-training can be decomposed into smaller, independently trainable jobs that can later be recomposed into a coherent larger model. We introduce Mixture of Training (MoT), a scaffolded modular pre-training procedure that partitions a target Transformer into contiguous layer blocks, trains each block inside a frozen pretrained aligner scaffold, and then recomposes the trained blocks with an optional short...

    arxiv.org/abs/2608.13277 · PDF

  7. 07

    How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures

    Paul Osemudiame Oamen, Owusu-Banahene Osei, Ananya Mukherjee, Christian Greisinger, Steffen Eger, Pius Onobhayedo, Wei Zhao

    cs.CL · cs.AI · cs.CV · cs.LG

    Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attention to behavioral reliability under uncertainty (how they behave when visual evidence is missing or misleading). We introduce SciFigBench, a diagnostic VLM benchmark for scientific figure understanding that jointly evaluates perception, reasoning, and behavioral...

    arxiv.org/abs/2608.13267 · PDF

  8. 08

    Self-Referential Induction Increases Response Instability Relative to Unresolvable and Verifiable Questions in Large Language Models

    Paras Balani, Subhrakanta Panda

    cs.CL · cs.AI

    Self-referential prompting has been shown to reliably induce large language models to produce first-person reports resembling subjective experience, but no prior work measures how consistent these reports are across repeated, independent trials, or how that consistency compares to the model's behavior on other kinds of open-ended questions. We measure response instability, defined as one minus the mean pairwise cosine similarity of sentence...

    arxiv.org/abs/2608.13258 · PDF

  9. 09

    GEM: A Generative Embedding Model Bridging Reasoning and Retrieval

    Zhili Shen, Craig Macdonald

    cs.CL · cs.AI · cs.IR

    Modern LLMs excel at reasoning and instruction following, enabling users to express complex and diverse information needs. However, conventional retrievers largely rely on surface-level matching between queries and documents, resulting in a growing gap between how users express their needs and how retrievers interpret them. In this paper, we present GEM, a generative embedding model that augments retrieval through its own knowledge by...

    arxiv.org/abs/2608.13200 · PDF

  10. 10

    Better Decomposition, Free Aggregation: A Synthesizer-Folding Framework for Multilingual Multi-Hop Question Answering

    Yilin Wang, Yuchun Fan, Weidong Bao, Zili Wei, Shi Feng, Tong Xiao, Zhengtao Yu, Jingbo Zhu

    cs.CL · cs.AI

    Multilingual retrieval-augmented generation (mRAG) equips large language models with access to globally distributed external knowledge for complex multilingual question answering. Recent approaches either translate retrieved documents into English or the query language to bridge the cross-lingual semantic gap, or decompose a complex query into sub-questions and aggregate the intermediate reasoning process. However, both lines of work suffer...

    arxiv.org/abs/2608.13160 · PDF

  11. 11

    LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation

    Chenrun Wang, Mingxuan Zhu, Tiancheng Huang, Wenjie Li, Yujie Zhang, Zichen Zhu, Zhiying Zou, Kai Yu, Lu Chen

    cs.CL · cs.AI · cs.DB · cs.MA

    With the rapid advancement of large language models (LLMs), research idea generation has attracted increasing attention. Existing approaches enable LLMs to retrieve relevant literature and propose novel ideas for research areas. However, current evaluation practices for idea generation remain fragmented and lack objective standards, often relying on direct LLM scoring, which limits their ability to provide unified and reliable assessments...

    arxiv.org/abs/2608.13136 · PDF

  12. 12

    Unifying Depth and Width Pruning for LLMs via Binary Knapsack Optimization

    Palaash Goel, Ayan Sengupta, Akshay Nambi, Tanmoy Chakraborty

    cs.CL · cs.LG

    Structured pruning is a promising approach for compressing large language models (LLMs), yet existing methods rely heavily on greedy heuristics that produce myopic decisions, and often fail to precisely meet target compression budgets. We present SNIPER, a two-stage structured pruning framework that solves a knapsack optimization over coarse-granularity components to yield conditionally optimal parameter allocations with respect to fixed...

    arxiv.org/abs/2608.12953 · PDF

  13. 13

    Falsehood and Impossibility Are Different Directions in an AI's Representation of Language

    Yoon Pyo Lee

    cs.CL · cs.AI

    Language can describe states of affairs that are false and states of affairs that could not be the case at all. Whether an AI model internally distinguishes these failures remains unclear. I report an exploratory activation study of the multimodal open-weight model Gemma 3 4B IT using 85 prompts from 17 philosophical families and a topic-matched modality set of 15 topics, each expressed as a truth, contingent falsehood, improbable claim,...

    arxiv.org/abs/2608.12852 · PDF

  14. 14

    AQuA: Recursively Self-Improving Quantitative Trading Research Agents

    Jiacheng Guo, Suozhi Huang, Yunlong Gao, Zihao Li, Jian Ge, Xu Kuang, Mengdi Wang

    cs.CL · cs.AI

    We study recursive self-improvement at the level of quantitative-investment research: whether an autonomous system can use evidence from earlier experiments to improve the hypotheses and candidates proposed in later iterations. We present AQuA, which comprises two separate language-model-driven research systems: one for symbolic factor discovery and one for trainable model development. The two systems do not share agents, memories, candidate...

    arxiv.org/abs/2608.12841 · PDF

  15. 15

    From Atomic Evidence to Logical Composition: Structured Compositional Reasoning over Compound Answer Options

    Obed Junias, Maria Leonor Pacheco

    cs.CL · cs.AI

    Large language models often fail when answer options require combining atomic judgments under explicit logical operators, even when they judge the individual atoms correctly. We study compound options connected by AND, OR, and NEITHER/NOR, introducing a framework that decomposes each option into atomic answers and scores contrastive hypotheses about each one, so the model never sees a compound option. An operator-constrained integer linear...

    arxiv.org/abs/2608.12836 · PDF

  16. 16

    CRAFT: LLM-Based Iterative Refinement for Temporal Reasoning over Clinical Narratives

    Chengyang He, Tahreem Arif, Marko Zivkovic, Lijing Wang, Yue Ning, Ping Wang

    cs.CL · cs.AI · cs.IR

    Understanding the temporal progression of symptoms in clinical narratives is critical for disease monitoring, safety surveillance, and causality assessment. Clinical narratives, however, rarely provide explicit temporal anchors. Current approaches to temporal information reasoning focus predominantly on pairwise relation classification across multi-visit and timestamp-rich records, leaving the reconstruction of structured symptom trajectories...

    arxiv.org/abs/2608.12779 · PDF

  17. 17

    ReconSpan: Reconstruction-Guided Adaptive Latent Tokenization

    Lixing Li

    cs.CL · cs.LG

    Adaptive latent tokenization maps a fine-grained input to a shorter sequence of continuous representations associated with input-dependent spans. We introduce ReconSpan, which divides text into chunks that a backward decoder can reconstruct from a single contextual prefix code and retains one such code as the latent token for each chunk. The reconstruction criterion is applied when chunks are formed, allowing one trained autoencoder to...

    arxiv.org/abs/2608.12756 · PDF

This edition is part of The Daily Abstract — cs.CL archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.