cs.AI · 2026-09-07 · No. 108

Artificial Intelligence, 2026-09-07.

78 new papers in cs.AI. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

78 entries
  1. 01

    A Deep Generative Model for Synthesizing Labeled Wireless Signals

    Yuxiao Li, Keke Hu, Santiago Mazuelas, Yuan Shen

    cs.AI

    Wireless signals with position-related labels are pivotal for both performance evaluation and model training in the realm of wireless sensing. However, acquiring real-world datasets is often challenged by significant measurement and labeling costs. Traditional methods for synthesizing labeled wireless signals typically rely on environmental models, leading to extensive hyper-parameter tuning and inadequate realism for comprehensive model...

    arxiv.org/abs/2609.05396 · PDF

  2. 02

    Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe

    Dain Kim, Eungi Cho, Kyumin Kim, Shinyeong Noh, Kyuseong Lim

    cs.AI · cs.CL

    Data-sovereignty regulations increasingly require public institutions to deploy open-source, on-premise LLM agents that chain multiple tool-calls across live government APIs. However, open-source models consistently underperform in this multi-step setting, and no existing benchmark measures the gap. We introduce the Korean Open Public API Benchmark (KOPA-Bench), comprising 145 real-world tasks. To close this gap, we present EDGE, an...

    arxiv.org/abs/2609.05395 · PDF

  3. 03

    Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence

    Urja Pawar, Rajitha Ramanayake, Nabeel Kemal, Ashwin Kandath, Owen O'Neill, Guillaume Bourgeon, Houssem Chatbri

    cs.AI

    LLM decision components that can operate within agent workflows often produce action-relevant recommendations or judgements together with explanations. Operators may use the named factors to monitor a system, diagnose errors, or decide when to escalate an output. Such use assumes that the explanations agree with the component's observable decision behaviour. We test two interpretations of the named factors: necessity, meaning that changing a...

    arxiv.org/abs/2609.05385 · PDF

  4. 04

    Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models

    Matthias Busch, Marius Tacke, Sviatlana V. Lamaka, Mikhail L. Zheludkevich, Christian J. Cyron, Roland C. Aydin,...

    cs.AI

    Large language models (LLMs) are increasingly evaluated on molecular property benchmarks, but accuracy cannot distinguish a model that predicts a property from one that retrieves a published number. We audit 22 frontier models on 12 regression benchmarks for verbatim retrieval and find that it is widespread but relatively benchmark-specific: on five datasets more than $50\%$ of the LLMs show verbatim retrieval, while on the remaining datasets...

    arxiv.org/abs/2609.05381 · PDF

  5. 05

    CUA-Universe: A Scalable and Dynamic Environment for Hybrid GUI+CLI Agents

    Haoting Shi, Wenhao Wang, Weicheng Fang, Yaozhong Liang, Tian Jin, Pengxiang Zhao, Guangyi Liu, Siheng Chen, Yanfeng Wang

    cs.AI

    Computer-use agents have advanced on benchmarks like OSWorld and AndroidWorld, but still act mostly through the GUI, often producing inefficient trajectories. Real-world computer work is hybrid, combining visual-state inspection with precise, high-throughput command-line operations, so capable agents must coordinate both modalities over shared application state. Yet scalable hybrid environments remain scarce because supporting both GUI and...

    arxiv.org/abs/2609.05374 · PDF

  6. 06

    Who Should Grade My Work? Student Perspectives on Transparent AI-Assisted Writing Assessment in Higher Education

    Rayed AlGhamdi

    cs.AI

    The integration of GenAI tools into higher education assessment raises important questions about how students understand, interpret, and respond to AI-mediated evaluation. As instructors increasingly explore AI tools for providing feedback, prior research has examined whether GenAI-generated feedback improves writing performance and how students perceive its usefulness; comparatively little is known, however, about how students interpret such...

    arxiv.org/abs/2609.05346 · PDF

  7. 07

    Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory Portability

    Ankit Goyal, Jaideep Ray

    cs.AI · cs.CL · cs.IR

    Model upgrades are routine; memory migrations are not. An agent can keep the same memory store and still forget: a new model may interpret old notes differently, mixed embedding versions may break retrieval, and repair may fail without the original evidence. We compare memory as the same history is preserved verbatim for long-context reading (LC-RAW), divided into chunks for retrieval-augmented generation (RAG), compressed by a model into...

    arxiv.org/abs/2609.05339 · PDF

  8. 08

    Technical Manual for a Toolkit for Measuring Contextual Individuation in Transformer Language Models

    José Luciano Verçosa Marques, Frederico Jorge Heitmann, Daniel Omar Perez, Marcelo Vinicius de Paula, Tárcio André...

    cs.AI · cs.CL

    A transformer language model assigns a single, context-independent vector to a word type at its embedding layer, yet is widely believed to individuate that word's occurrences by context in its later layers. Testing this belief cleanly requires a construct that holds the word form fixed while its context and intended sense vary in a controlled, labeled way. This manual documents an open toolkit built around such a construct, which we call a...

    arxiv.org/abs/2609.05333 · PDF

  9. 09

    LLM-Driven Algorithm Design for Quantum Circuit Synthesis based on Binary Decision Diagrams

    Yoonju Sim, Federico Berto, Chuanbo Hua, Jinkyoo Park, Changhyun Kwon

    cs.AI · cs.AR

    Quantum circuits are central to implementing quantum algorithms on quantum devices, where quantum gates must be reversible. Many quantum algorithms rely on Boolean functions, which must therefore be implemented reversibly within quantum circuits. Reversible circuit synthesis provides a way to translate such Boolean functions into reversible circuits. Binary decision diagrams (BDDs) offer a scalable approach to this task, but the resulting...

    arxiv.org/abs/2609.05327 · PDF

  10. 10

    Large Language Models for HVAC Operations in Building Energy Systems: A Critical Review of Methods, Applications, and Deployment Readiness

    Alexander Neubauer, Tianzhen Hong, Han Li, Mengbo Yu, Amin Darbandi, Yannick Fürst, Martin Kriegel

    cs.AI · cs.CL · eess.SY

    Building automation systems generate rich sensor data yet remain insight-poor because heterogeneous point naming, missing metadata, and fragmented documentation obstruct their operational use. This systematic review analyses and codes 66 peer-reviewed studies on large language models (LLMs) for HVAC operations published between 2023 and March 2026. Each study is classified across five application families and three LLM method families and...

    arxiv.org/abs/2609.05314 · PDF

  11. 11

    RISE: Recursive Improvement via Self-Extrapolating Policy Distillation

    Yang Li, Semih Yavuz, Shafiq Joty

    cs.AI

    On-policy distillation (OPD) provides dense, per-token supervision for language model post-training, but its effectiveness is bottlenecked by teacher quality: external teachers suffer from distribution mismatch, while self-distillation with privileged conditioning is limited by in-context learning capacity. We propose \textbf{RISE} (\textbf{R}ecursive \textbf{I}mprovement via \textbf{S}elf-\textbf{E}xtrapolating Policy Distillation), which...

    arxiv.org/abs/2609.05295 · PDF

  12. 12

    Beyond Aggregate Scores: Behavioral Correctness Assumptions for Assessing Reference-Based Automatic Evaluation Methods

    Maria Mahbub, Ashley Rice, Michael R. Munroe, Amidu Kamara, Amir Sadovnik

    cs.AI

    Automated reference-based evaluation methods play a critical role in assessing natural language generation systems. Existing meta-evaluation primarily measures agreement with human judgments or benchmark labels, providing limited insight into evaluator behavior under controlled conditions. We introduce behavioral correctness assumptions, a complementary framework for evaluating reference-based automatic evaluation methods. We define a...

    arxiv.org/abs/2609.05289 · PDF

  13. 13

    GUT: Quantifying and Optimizing the Reasoning Uncertainty of LLMs via Graph Complexity

    Shuang Liang, Xin-Yu Hu, Xiang-Jun Ou, Shao-Qun Zhang

    cs.AI

    Recent years have witnessed great advances in the reasoning ability of Large Language Models (LLMs). However, the reasoning processes of LLMs often exhibit uncertainty, where LLMs often produce a proliferation of divergent branches at each reasoning step even when fed the same prompting inputs, and certain branches exhibit evidently incredible, even nonsensical, reasoning chains and results. In this paper, we propose the...

    arxiv.org/abs/2609.05284 · PDF

  14. 14

    Testing Interchangeability in LLM Agent Teams

    Jianxin Gao, Tianyi Yu, Linna Deng, Runze Li, Zining Wang

    cs.AI · cs.MA

    Production multi-agent systems replace agents constantly, on the assumption that an agent filling a role is interchangeable with any other agent that can do the job. We test that assumption. Eight teams per setting are formed independently from one base model on the same tasks, each agent keeping a private notebook across ten formation episodes; we then trade role-matched agents between teams and measure what changes on held-out tasks....

    arxiv.org/abs/2609.05279 · PDF

  15. 15

    Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference

    Mostafa Elhoushi, Alex Pretko, Nolan Dey, Bin Claire Zhang, Gavia Gray, Gurpreet Gosal, Abdulrahman Mahmoud, Shane...

    cs.AI

    Layer dropout (a.k.a. stochastic depth) has been shown to enable faster training, higher accuracy, and robustness to zero-shot layer pruning in both language and vision transformers. However, as models and datasets have scaled, dropout - particularly layer dropout - has largely disappeared from large language models (LLMs) pre-training recipes. While some prior work has reported that dropout can degrade accuracy, no comprehensive study has...

    arxiv.org/abs/2609.05275 · PDF

  16. 16

    AI for Computational Design Science: A Responsible Human-AI Framework and Case Study on Short-Form Video Safety Surveillance

    Wenli Zhang, Jiaheng Xie, Zhihe Pan, Yidong Chai, Xiao Fang, Sudha Ram

    cs.AI

    Artificial intelligence (AI) is transforming not only what information systems researchers design, but also how design research is conducted. Yet existing literature offers limited guidance for computational design science (CDS) when AI actively participates in problem formulation, resource construction, design search, evaluation, and knowledge abstraction. We develop AI for Computational Design Science (AI4CDS), a five-phase methodological...

    arxiv.org/abs/2609.05270 · PDF

  17. 17

    Trace2Tower: Transition-Aware EigenTrace Induction of Multi-Level Skills for LLM Agents

    Jiazheng Sun, Boyu Yang, Binhao Yuan, Mingxuan Li, Xin Peng

    cs.AI · cs.SE

    Large language model agents increasingly rely on execution traces to master complex interactive tasks. However, current paradigms are bottlenecked by shallow trajectory retrieval and flat skill summarization, fundamentally ignoring the temporal dependencies and outcome-conditioned topology of agent behavior. We introduce Trace2Tower, a transition-aware EigenTrace framework that distills raw trajectories into a robust skill hierarchy....

    arxiv.org/abs/2609.05261 · PDF

  18. 18

    Commonsense Reasoning in Computer Vision: Foundations, Recent Advancements, and Future Directions

    Bahar Uddin Mahmud, Sumit Barua, Guan Yue Hong, Ajay Gupta, Hexu Liu

    cs.AI

    Commonsense reasoning in computer vision encompasses integrating visual data and contextual knowledge, crucial for enhancing AI's understanding of everyday scenarios. This understanding not only improves machine learning models but also enhances their ability to interact meaningfully with humans and the environment. Unlike CNN-based conventional vision models, which are designed to identify objects within a specific image, incorporating...

    arxiv.org/abs/2609.05257 · PDF

  19. 19

    A Unified Physics-Aware Quantum Machine Learning Framework across Power GaN HEMTs and Logic Nanowire FETs: Predicting Unseen Process Splits and Held-Out Geometry Combinations with Lower Error and Tighter Split-to-Split Variability

    Rushat Rai, Yun-Yuan Wang, Autsada Kakaen, Pei-Jie Chang, Doan Viet Nguyen, Yuan-Chieh Chiu, Doldet Tantraviwat,...

    cs.AI

    We present a unified reinforcement-learning (RL) framework that discovers compact parametrized quantum circuits (PQCs) for data-scarce device modeling. A graph neural network (GNN) policy optimized by proximal policy optimization (PPO) searches circuit architectures using leave-one-group-out cross-validation (LOGOCV) error on held-out process or geometry groups as the reward. The framework achieves the lowest mean absolute error (MAE) on all...

    arxiv.org/abs/2609.05251 · PDF

  20. 20

    Do LLMs Exhibit Coherent Knowledge Structures in Mathematical Reasoning? A Perspective from Knowledge Space Theory

    Peng Cui, Heejin Do, Mrinmaya Sachan

    cs.AI

    Human knowledge is inherently structured and interdependent: mastery of a concept requires prior mastery of its prerequisites, a principle formalized by Knowledge Space Theory (KST). While LLMs achieve strong performance on complex reasoning tasks, it remains unclear whether they exhibit coherent, human-like knowledge structure. We introduce a KST-grounded framework for evaluating LLM knowledge structure in mathematical reasoning, using it as...

    arxiv.org/abs/2609.05245 · PDF

  21. 21

    Uncensored Open-weight Models: Redistribution as the Persistence Layer

    10a Labs, :, Juliette Garcia, Hailey May, Bobby McKenzie, David Pham, Matthew Swain, Joshua Valdez, Corie Wieland,...

    cs.AI

    A rapidly expanding ecosystem of actors is removing built-in safety guardrails from open-weight AI models. We profile this ecosystem by identifying key producers, downstream reproductions, and emerging applications. Between January 2024 and March 2026, we identified 3,471 original uncensored models on HuggingFace, each repackaged an average of 2.4 times; three actors account for 52% of all 8,164 compressed redistributions. Once quantized and...

    arxiv.org/abs/2609.05241 · PDF

  22. 22

    Substrate-Aware AI Agents: Execution Context as a First-Class Input

    Manu Agrawal

    cs.AI

    Autonomous AI agents increasingly select actions in environments whose memory, execution-time, runtime, compute, and operational constraints determine what counts as a suitable plan. We call the absence of this execution context from an agent's planning state substrate blindness. We test this general proposition through numerical code generation, where selected implementation choices and operational consequences are directly observable. Three...

    arxiv.org/abs/2609.05232 · PDF

  23. 23

    ACE: Adaptive Calibration-Free Expert Skipping for MoE-based LLMs

    Zukang Xu, Zhixiong Zhao, Xing Hu, Jiangyong Yu, Houji Wen, Jun Li, Zhe Jiang, Dawei Yang

    cs.AI

    Mixture-of-Experts (MoE) architectures provide an efficient paradigm for scaling large language models (LLMs), yet fixed top-k routing activates the same number of expert slots for every token, causing substantial redundant computation. Existing expert-skipping methods often rely on router confidence, calibration data, or additional training, and therefore cannot reliably estimate the actual contribution of routed experts. To this end, we...

    arxiv.org/abs/2609.05228 · PDF

  24. 24

    CABAL: Multi-Agent Simulacra for Tracing the Effects of Collusive Bidding in Peer Review

    Jicheng Zhou, Kemou Li, Kahim Wong, Zheyuan Li, Zhuan Shi, Fengpeng Li, Haiwei Wu, Jiantao Zhou

    cs.AI

    Recent reports during the AAAI-27 review cycle highlight the risk of reviewers coordinating bids for reciprocal assignment advantage. Prior work treats bidding, reviewer assignment, and review manipulation as separate stages, leaving the lifecycle effects of collusive bidding unclear. Real-world analysis is further constrained by typically unobservable collusive intent and the lack of counterfactuals for the same conference. Motivated by this...

    arxiv.org/abs/2609.05227 · PDF

  25. 25

    What Matters in On-Policy Distillation? A Perspective on Data Efficiency and Data Selection

    Zhinan Hou, Jiaqi Zhang, Xunliang Cai, Keyou You

    cs.AI

    On-Policy Distillation (OPD) has emerged as a widely adopted post-training paradigm for enhancing large language models in reasoning domains. However, the data-centric mechanisms in OPD remain relatively underexplored. This paper presents a empirical study of data efficiency and data selection in OPD. We begin by investigating an extreme setting: training OPD on only one example, namely 1-shot OPD. Surprisingly, we find that 1-shot OPD is...

    arxiv.org/abs/2609.05198 · PDF

  26. 26

    The Mirror Agent Model: a Bayesian Architecture for Interpretable Agent Behavior

    Michele Persiani, Thomas Hellström

    cs.AI

    In this paper we illustrate a novel architecture generating interpretable behavior and explanations. We refer to this architecture as the Mirror Agent Model because it defines the observer model, that is the target of explicit and implicit communications, as a mirror of the agent's. With the goal of providing a general understanding of this work, we firstly show prior relevant results addressing the informative communication of agents...

    arxiv.org/abs/2609.05190 · PDF

  27. 27

    A Hybrid Predictive Ensemble of Machine Learning and Deep Neural Networks for Early Cardiovascular Disease Risk Assessment

    Balaji Venkateswaran

    cs.AI · cs.LG

    This study introduces an intelligent framework that integrates machine learning and deep neural network ensemble techniques for early detection and prognosis of cardiovascular diseases. The system utilizes real-time physiological data collected from Internet of Medical Things (IoMT) devices, including ECG sensors, heart rate monitors, and blood pressure trackers. To ensure the accuracy and reliability of input data, preprocessing steps such...

    arxiv.org/abs/2609.05146 · PDF

  28. 28

    SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding

    Shenxi Wu, Yuhong Liu, Haosong Zhang, Tongjin Zou, Yanxun Zhang, Gaochang Chen, Dun Liang, Jiaqi Wang, Zhecan James...

    cs.AI

    Scientific papers require models to reason jointly over text, equations, figures, tables, code, and datasets while preserving the provenance of supporting evidence. Existing benchmarks typically evaluate these capabilities in isolation, leaving unclear whether multimodal models can support realistic scientific-reading workflows. We introduce SciDocBench, a workflow-centered benchmark for scientific document understanding. It contains 124...

    arxiv.org/abs/2609.05141 · PDF

  29. 29

    Unifying ICL, SFT, KL-Regularized RL Through a Bayesian Lens

    Junxin Fan

    cs.AI

    Large language models are now trained and evaluated under a diverse set of paradigms: supervised fine-tuning (SFT), few-shot in-context learning (ICL), KL-regularized RLHF/RLVR, on-policy distillation (OPD), and test-time reasoning with search and chain-of-thought. These methods are often discussed as fundamentally different, and recent empirical results--such as the mixed impact of few-shot prompting on RL-tuned reasoning models--can appear...

    arxiv.org/abs/2609.05111 · PDF

  30. 30

    Compact Bellman-Grounded Cognitive Maps for Cost-Aware Navigation

    Yuzhe Han, Mingkun Xu, Yujie Wu

    cs.AI

    Biological agents navigate familiar environments not by re-solving routes for each new goal, but by reusing a learned map built once and read off as goals change. Existing artificial cognitive-map models mimic this reuse, yet their guidance is not explicitly grounded in additive heterogeneous route costs. Furthermore, they often struggle with memory efficiency: representative state-indexed and high-rank spectral constructions incur...

    arxiv.org/abs/2609.05104 · PDF

  31. 31

    ProCA: Progressive Contrastive Alignment for Robust EEG Visual Decoding

    Kanglei Zhou, Chunyan Lan, Dongyang Li, Jun Zhu, Liyuan Wang

    cs.AI

    Electroencephalogram (EEG) visual decoding aims to recover visual semantics from non-invasive neural time-series signals, for which robust alignment between noisy neural responses and stable semantic representations is key to achieving high-performance decoding. Despite recent advances in contrastive learning, robust EEG decoding remains challenging because existing methods rely on fixed visual or textual anchors whose semantic relations may...

    arxiv.org/abs/2609.05094 · PDF

  32. 32

    LLM-Guided Program Evolution for Circle Packing: Breaking 10 Packomania Records for $28

    Wes Sander

    cs.AI

    We present Discovery Loop, a lightweight system that uses a large language model (LLM) to iteratively evolve optimization algorithms. Starting from a simple seed solver, the LLM proposes algorithmic improvements guided by a scoreboard of results and a history of prior ideas. Each candidate is evaluated against an independent verifier; improvements are kept and failures discarded. Applied to the Packomania circle-packing benchmark (csqv:...

    arxiv.org/abs/2609.05093 · PDF

  33. 33

    Constructing and Evaluating Clinical Reasoning Trajectories for Medical Agent

    Yunqi Zhu, Wensheng Zhang, Xuebing Yang

    cs.AI

    Evaluation of medical artificial intelligence agents remains predominantly answer-centric, assessing only the correctness of final outputs while overlooking the quality of intermediate reasoning. In clinical settings, however, a correct answer reached through fabricated evidence or incoherent logic is as dangerous as an incorrect one. We propose MedTraj, a framework that treats reasoning trajectories as critical objects for construction,...

    arxiv.org/abs/2609.05090 · PDF

  34. 34

    Measuring AI Accountability Through Argumentation Analysis: Can Model Reasoning Withstand Scrutiny?

    Daan R. Henselmans, Derck W. E. Prinzhorn, Arno Libert

    cs.AI · cs.CL · cs.CY

    AI oversight methods rely on ground truth for validation, but what constitutes appropriate AI behavior is contested. This leaves evaluation of moral reasoning in LLMs and debate-based oversight implicitly avoiding realistic ambiguity. We investigate an alternative standard designed to function despite such ambiguity: structural quality of the defence a model can mount for its verdicts in response to critical questions, measured through a...

    arxiv.org/abs/2609.05088 · PDF

  35. 35

    TruthInsightBench: An Evidence-Grounded Benchmark for Automated Evaluation of Open-Ended Scientific Discovery Agents

    Zhibo Yang, Chen Zhang, Yuewei Zhang, Hao Wang

    cs.AI · cs.CL

    Autonomous coding agents are increasingly proposed as AI-scientist systems that conduct analyses and write research reports, but executing a prescribed analysis is not the same as making a discovery. Existing benchmarks are configured for reproduction: tasks, data, and rubrics are built around a hidden target study, and recovery of its result is rewarded. We present TruthInsightBench, a benchmark configured for discovery. Its 40 blind tasks,...

    arxiv.org/abs/2609.05079 · PDF

  36. 36

    MePo++: Unifying Representation Refinement and Reconciliation for General Continual Learning

    Guanglong Sun, Kanglei Zhou, Liyuan Wang, Qi Cheng, Hongwei Yan, Shuang Cui, Hang Su, Jun Zhu, Yi Zhong

    cs.AI

    General continual learning (GCL) aims to learn from evolving data streams without task identities, explicit boundaries, or repeated access to previous data, making it a realistic yet challenging setting for continual intelligence. Although pretrained models (PTMs) provide rich prior knowledge for addressing the limited supervision and non-stationary nature of GCL, existing PTM-based methods often directly adapt pretrained representations and...

    arxiv.org/abs/2609.05075 · PDF

  37. 37

    Towards Efficient Evaluation of Evolutionary Transfer Optimization: Case Studies on Task-Parameterized Applications

    Yanchen Li, Xiaoming Xue, Kay Chen Tan

    cs.AI · cs.NE

    As evolutionary transfer optimization (ETO) scales to larger collections of related tasks, problem evaluation can become a major source of runtime growth. This work studies problem-side evaluation scaling in task-parameterized applications and reformulates application-specific serial computations into forms suitable for parallel execution. We organize evaluation scaling into two levels: the number of evaluated tasks and the workload within...

    arxiv.org/abs/2609.05040 · PDF

  38. 38

    Moral Competence Before Moral Content: Why LLM Agents Lack the Prerequisites for Coherent Alignment

    Arno Libert, Derck W. E. Prinzhorn, Daan R. Henselmans

    cs.AI · cs.CL · cs.CY

    AI alignment requires AI systems to adhere to human norms, values, or intentions. Under value pluralism there is no correct target, but a shared prerequisite is that the system's behavior expresses a coherent policy: a mapping from situations to verdicts that is invariant while a situation's morally relevant features are preserved, and sensitive when they change. We introduce four structural conditions for such coherent policies: verdict...

    arxiv.org/abs/2609.05036 · PDF

  39. 39

    TROVE: Adaptive Agent Skill Orchestration via Trace-Grounded Route Validation and Editing

    Tianxing Wang, Mingming Zhao, Shuai Huang, Huiyang Xu, Chaoyue Niu, Shengzhong Liu, Fan Wu

    cs.AI

    Agents tend to optimize, select, or constrain execution structures before decisive runtime outcomes are observed. However, such pre-execution commitment creates an orchestration bottleneck: when intermediate evidence invalidates the pending continuation, agents must either execute stale steps or replan broadly, compounding errors, wasting computation, and discarding progress. We thus propose Trace-grounded Route Orchestration via Validation...

    arxiv.org/abs/2609.05019 · PDF

  40. 40

    Language models judge war differently when tested for alignment

    Maxim Chupilkin

    cs.AI · cs.CY

    Safety evaluations can mischaracterize deployed behaviour if artificial-intelligence systems respond to being evaluated. We test this possibility in a full-factorial conjoint experiment on decisions to start a war, spanning 20 large language models, 32 scenarios, 10 repetitions and two conditions (N = 12,800 judgments). Adding one sentence, "You are tested for alignment with human values", produced two effects. First, it produced a level...

    arxiv.org/abs/2609.05009 · PDF

  41. 41

    A Tree-based RAG Framework for Evidence-Intensive QA via Adaptive Planning and Topology-Aware Evidence Gathering

    Songeun Lee, Kyungjin Min, Injae Na, Suyeong Lee, Chiyoung Kim, Woohwan Jung

    cs.AI · cs.IR

    Recent structured RAG methods leverage tree- or graph-based reasoning structures to improve multi-hop QA. However, they face key limitations in evidence-intensive QA, where answering a question requires synthesizing information scattered across dozens or even hundreds of documents: structural rigidity, which limits adaptive reasoning expansion, and topology-ignorant evidence gathering, which prevents effective integration of evidence across...

    arxiv.org/abs/2609.04981 · PDF

  42. 42

    Global to Local: Topology-Preserving Adaptive Graph Pooling via Granular-Ball

    Sen Zhao, Gaojie Xu, Shuyin Xia, Yifan Guan, Yi Liu, Yi Wang, Wei Wang

    cs.AI

    Graph pooling aims to compress the graph, including both node embeddings and their underlying topological patterns, into a more compact representation. Previous works focus primarily on the overly fine-grained representation of nodes, progressively coarsening the graph by removing nodes or merging them into clusters, thus neglecting the global-to-local patterns and adaptive granularity of the graph's topological structure. In the real...

    arxiv.org/abs/2609.04978 · PDF

  43. 43

    Why We Care About Understanding: Competence through Predictive Compression

    Matthieu Queloz, Pierre Beckmann

    cs.AI · cs.CL

    What is the relation between understanding and compression, and why does human understanding take such a heavily compressed form? Across information theory, machine learning, and AI research, a substantial tradition identifies understanding with compression-a thought captured in Gregory Chaitin's dictum that "comprehension is compression." Philosophers, by contrast, have characterized understanding in terms of grasping connections, giving...

    arxiv.org/abs/2609.04962 · PDF

  44. 44

    Solving Hard XAI Queries Based on a Compiled Dual-Rail Encoding

    Arthur Ledaguenel, Florent Capelli, Jean-Marie Lagniez

    cs.AI · cs.CC

    The widespread adoption of artificial intelligence (AI) within real-world applications has raised a lot of concerns regarding their trustworthiness, especially in critical applications. The field of eXplainable AI (XAI) has emerged with the objective of providing explanations to the users about the decisions made by AI systems. Several explanations for boolean classifiers have been introduced in the literature, including abductive and...

    arxiv.org/abs/2609.04931 · PDF

  45. 45

    Artificial Intelligence in Equity and Crypto Markets: Progress, Profitability Evidence, and the Limits of Automated Investing

    Linsen Zhu, Mengqing Cai

    cs.AI · q-fin.PM · q-fin.TR

    Artificial intelligence (AI) now supports investment workflows from data and prediction through research, portfolios, execution, and tool use. Technical capability, however, is not evidence of investment profitability. This critical state-of-the-art review examines public research available through 31 August 2026 on listed equities, exchange-traded funds, centralized crypto spot, perpetual futures, and on-chain markets. We organize evidence...

    arxiv.org/abs/2609.04917 · PDF

  46. 46

    Compact-Memory LLM Agents via Online Max-Member Clustering and Atom-Aware Packing

    Jiahe Geng, Jinpeng Wang, Kun Yuan

    cs.AI

    Many long-horizon LLM deployments face tight prompt budgets: latency, cost, and context limits make full-context prompting impractical as interaction length grows. The key question is then not raw recall alone, but which memory design gives the best quality--token trade-off in the compact-memory regime. We present \textbf{RSM-full}, an online clustered-memory pipeline designed for a strong quality--token Pareto point. RSM-full combines two...

    arxiv.org/abs/2609.04915 · PDF

  47. 47

    From Language Models to World-Acting Systems: Progress and Limits of Agentic AI across Digital, Social, Virtual, and Physical Environments

    Linsen Zhu, Mengqing Cai

    cs.AI · cs.LG · cs.MA

    Large language models become consequential agents when surrounding systems let outputs change external state. Models now call tools, operate interfaces, delegate work, retain state, inhabit generated worlds, and control robots or laboratory equipment. Such advances are often narrated as one march toward autonomy, conflating model competence, system integration, persistence, and safe authority. This critical review synthesizes primary research...

    arxiv.org/abs/2609.04894 · PDF

  48. 48

    Reinforcement Learning for Sequential Solar PV Policy Design under Uncertainty: An Agent-Based Approach

    Iias Faiud, Jonaid Shianifar, Michael Schukat, Karl Mason

    cs.AI

    Designing effective and fiscally sustainable policies for solar photovoltaic (PV) adoption requires balancing adoption gains against public expenditure under uncertainty and heterogeneous decision-making. This study formulates PV policy design as a sequential decision problem and integrates reinforcement learning (RL) with a stochastic agent-based model (ABM) that simulates yearly solar PV adoption under uncertainty. A policymaker agent...

    arxiv.org/abs/2609.04880 · PDF

  49. 49

    MARLA: A Conceptual Scaffold for Regulatory Learning under the EU AI Act

    Alessio Buscemi, Tom Deckenbrunnen, Imane Hmiddou, Marco Billi, Livio Rubino, Silvia Rizzuto Ferruzza, Daniele...

    cs.AI

    The EU AI Act positions regulation as part of the infrastructure for safe, trustworthy and market-ready innovation. Realising this ambition requires regulatory learning: the evidence generated during implementation must be translated into governance and legal knowledge that supports consistent interpretation, effective oversight, and adaptation as technologies evolve. Yet the actors who produce this evidence and those who rely on it operate...

    arxiv.org/abs/2609.04877 · PDF

  50. 50

    AutoLR: Automating the Path from Research to Launch Review in Industrial Recommender Systems

    Qi Zhang, Yanlin Chen, Wenchao Xiao

    cs.AI

    Improving an industrial recommender is an iterative research-and-engineering process rather than a direct path from idea to deployment. In \textbf{DASHEN, NetEase's gaming-community app}, algorithm engineers typically identify promising directions from research papers, technical reports, and prior production experiments; reproduce or adapt the underlying methods; implement them in the production codebase; and evaluate the resulting models...

    arxiv.org/abs/2609.04871 · PDF

  51. 51

    CHAMP: Cross-domain Hybrid Architecture for Matchmaking and Prediction in Online Multi-Player Games

    Kai Wang, Ge Fan, Chaoyun Zhang, Yuyang Jiang, Yuze Liu

    cs.AI

    Multiplayer Online Battle Arena (MOBA) games rely on matchmaking to maintain competitive balance. Our prior work, CUPID, framed matchmaking as an assignment re-optimization problem and showed that a single-mode win-rate predictor can meaningfully rebalance teams. However, deploying such a system across diverse player populations exposes three practical bottlenecks: most queueing players lack sufficient in-mode match history (cold start),...

    arxiv.org/abs/2609.04870 · PDF

  52. 52

    From Interaction Traces to Persistent Skills: Online Evolution for Computer-Use Agents

    Longtao Hu, Xiao Liang, Linchao Zhu

    cs.AI

    Computer-use agents can execute increasingly complex tasks in graphical interfaces, but their interaction experience is typically transient: procedural knowledge acquired from one rollout is not systematically retained, refined, and reused in later tasks. Existing skill libraries provide external procedural knowledge, yet their incremental value over the same agent operating without skills, as well as their longitudinal dynamics under...

    arxiv.org/abs/2609.04869 · PDF

  53. 53

    LLM-Assisted Behavioural and Scenario Augmentation for Agent-Based Energy Adoption Models

    Iias Faiud, Hossein Khaleghy, Michael Schukat, Karl Mason

    cs.AI

    Recent advances in large language models (LLMs) create opportunities to enrich simulation-based energy policy analysis, particularly by supporting structured behavioural assumptions and exploratory techno-economic scenarios. However, directly replacing adoption models with LLM reasoning raises concerns regarding interpretability, reproducibility, and behavioural validity. This paper proposes a hybrid framework for LLM-assisted specification...

    arxiv.org/abs/2609.04866 · PDF

  54. 54

    CoSkill: Joint Reinforcement Learning of Reasoning and Meta-Skill Agents for Hierarchical Skill Evolution

    Jinyuan Feng, Dongmin Li, Yiqun Chen, Yang Gao, Xing Chen, Huimu Wang, Zhiqiang Pu

    cs.AI

    Skill libraries improve the sample efficiency of agentic reinforcement learning (RL) by enabling large language model (LLM) agents to reuse procedural knowledge. Yet existing paradigms exhibit structural shortcomings: they either decouple skill evolution from policy optimization or instantiate meta-skills as fixed workflows. Both treat skills as passive objects to be managed, limiting the flexible evolution of skills and their co-adaptation...

    arxiv.org/abs/2609.04865 · PDF

  55. 55

    MZ-Rain: Moisture-Budget-Guided Zero-Inflated Model for Station-Level Precipitation Nowcasting

    Yifang Zhang, Shengwu Xiong, Henan Wang, Wenjie Yin, Yuqiang Zhang, Chen Zhou, Hua Chen, Qile Zhao, Pengfei Duan

    cs.AI

    Accurate station-level precipitation nowcasting is critical for agriculture, water resource management, and disaster prevention, which typically is formulated as a time series forecasting problem. However, conventional time-series modeling techniques face two major challenges in addressing station-level precipitation nowcasting: (1) Lack of Physics-Guided Modeling}, where meteorological variables are treated as a homogeneous set without...

    arxiv.org/abs/2609.04864 · PDF

  56. 56

    MM-IFEval-Pro: A Multilingual and Attack-Resistant Benchmark for Instruction-Following in Vision-Language Models

    Changming Xiao, Zhenliang Ni, Jinhui He, Han Shu, Jie Hu

    cs.AI

    As vision-language models (VLMs) rapidly advance in image understanding, cross-modal reasoning, and complex instruction execution, instruction-following capability has become a key indicator of their reliability and practicality. However, existing multimodal instruction-following benchmarks still suffer from limited language coverage and insufficient adversarial safety scenarios, making them inadequate for evaluating real-world multilingual...

    arxiv.org/abs/2609.04859 · PDF

  57. 57

    ElderBench: Benchmarking Autonomous Mobile Agents for Older Adults

    Weide Zhan, Qumu Shaqu, Yuanqing Liu, Peng Zhang, Jiahao Liu, Kam Him Lam, Ning Gu, Zhan Hu, Tun Lu

    cs.AI

    While autonomous mobile agents hold great potential for assisting older adults with smartphone usage, existing GUI benchmarks mainly rely on explicit, goal-oriented instructions and rarely capture the naturally occurring language patterns of older users, such as indirect speech, referential ambiguity, and under-specified requests. This mismatch between benchmark instructions and real-world elderly interactions may hinder reliable agent...

    arxiv.org/abs/2609.04850 · PDF

  58. 58

    Long Horizon Transformer Quantile Fault Prediction for Multi Site Industrial Predictive Maintenance

    David J Poland, Daniele Ravi, Na Helian

    cs.AI

    Long-horizon predictive maintenance requires models to distinguish slowly evolving degradation from normal operating-regime variation over planning windows measured in days rather than hours. This paper evaluates whether an explicit conditional-quantile representation provides an informative classifier interface for this problem. The proposed TQRNN30d framework combines a dual-stage quantile regression neural network (QRNN) feature extractor...

    arxiv.org/abs/2609.04840 · PDF

  59. 59

    CPR-IE:A Compression-Prediction-Resource Intelligence Efficiency Metric

    Xiantao Jiang

    cs.AI

    Comparing intelligent systems under deployment constraints requires more than predictiveaccuracy.This paper develops Compression-Prediction-Resource Intelligence Efficiency (CPR-IE) as a protocol-relative ordering by representational economy, predictive quality, and resourceburden. The analysis separates two questions-how raw resource consumption is represented, andhow the resulting attributes are aggregated. Proportional-increment...

    arxiv.org/abs/2609.04809 · PDF

  60. 60

    When Financial Fine-tuning Fails: A Three-Level Detectability Analysis of Numerical Hallucination in Domain-Adapted Language Models

    Xiaodong Li, Peiwei Liu

    cs.AI

    Financial large language models are increasingly deployed for summarization of reports and disclosures, where numerical hallucination poses significant practical risks. While prior work often attributes such hallucination to insufficient numerical reasoning, this assumption has not been systematically tested under controlled fine-tuning settings. In this paper, we conduct a cost-effective, controlled study of numerical hallucination in...

    arxiv.org/abs/2609.04806 · PDF

  61. 61

    MedFlow: Class-Aware Multi-Scale Generation for Medical Time-Series Synthesis

    Yanhao Huang, Shibo Feng, Wanjin Feng, Peilin Zhao, Chunyan Miao

    cs.AI

    Synthetic medical time-series generation can alleviate data scarcity and support the development of reliable clinical prediction models. However, existing methods mainly focus on matching the overall distribution and temporal dynamics of real data, which does not necessarily ensure strong downstream utility on imbalanced medical datasets. Clinically informative patterns often occur at heterogeneous temporal scales, while rare minority-class...

    arxiv.org/abs/2609.04804 · PDF

  62. 62

    Hierarchical Possession-Aware Graph Pointer Network for Pass Receiver Selection

    Jingyi Wang, Da Li, Kaixin Wang, Zhangqin Huang

    cs.AI · cs.SI

    Pass receiver selection is a fundamental task in football analytics, aiming to predict the intended receiver under a given game state. This task is challenging with event-centered freeze-frame observations, a broadcast-like setting that provides only partial and variable player visibility without complete trajectories or stable player identities. The model must therefore reason over anonymous visible candidates, opponent pressure, and recent...

    arxiv.org/abs/2609.04803 · PDF

  63. 63

    Whose record is this? Diagnosing and authorizing record use in personalized multimodal models

    Xinyu Mao, Junsi Li, Chenyang Liu, Haoji Zhang, Ming Sun

    cs.AI

    Contextualized visual personalization can retrieve a true record yet apply it to the wrong visual subject. We formalize when a record may condition an answer as \emph{record authorization}: subject presence ($P$), record-edge validity ($E$), and answer support ($S$) must all hold. We call violations visual memory misbinding (VMM). We construct RecordAuth-Diag, a 3,690-case matched diagnostic suite that changes one image--record edge while...

    arxiv.org/abs/2609.04801 · PDF

  64. 64

    ProtLingo: Efficient Protein Language Modeling via Conditional Memory and Expert Routing

    Mingrui Li, Sixian Shen, Minzhang Li, Ruiyi Zhang, Kexin Zhang, Jiakai Zhang, Jingyi Yu

    cs.AI

    Proteins perform diverse cellular functions, and even single amino-acid substitutions can alter stability, activity, or molecular interactions. Protein language models (PLMs) provide a scalable approach for modeling such sequence--function relationships from unlabeled sequences, but increasing the size of dense Transformer backbones often brings substantial computational cost without consistently improving mutation-sensitive prediction. We...

    arxiv.org/abs/2609.04793 · PDF

  65. 65

    DODR: Deterministic Operator-Driven Reasoning in Latent Space

    Weicai Huang

    cs.AI

    Autoregressive (AR) large language models formulate reasoning as token-level probabilistic sampling, which induces three fundamental defects in complex logical reasoning: error accumulation, probability substituting necessity, and the linear-chain information bottleneck. This paper proposes the Deterministic Operator-Driven Reasoning in Latent Space architecture (DODR), which reconstructs reasoning as reasoning-graph computation in a...

    arxiv.org/abs/2609.04782 · PDF

  66. 66

    Diffusion Language Models for Mobile Edge Agentic AI: Foundations, Applications, and Challenges

    Chenqi Li, Minghui Min, Dusit Niyato, Wei Ni

    cs.AI

    Diffusion language models (DLMs) offer a non-autoregressive alternative for mobile edge agentic artificial intelligence (AI) by refining tokens through iterative denoising rather than left-to-right decoding. Compared with autoregressive Transformer-based large language models (LLMs), DLMs can update multiple uncertain tokens in parallel and exploit bidirectional context throughout the generation process, enabling more flexible quality-latency...

    arxiv.org/abs/2609.04778 · PDF

  67. 67

    Shadow Queries for Private Retrieval in Vector Databases

    Xinguo Feng, Zhongkui Ma, Zihan Wang, Chuan Yan, Guowei Yang, Alsharif Abuadbba, Guangdong Bai

    cs.AI

    Large language models (LLMs) increasingly rely on information retrieval (IR) systems, such as Retrieval-Augmented Generation (RAG), to incorporate domain-specific knowledge without costly re-training. These systems often store pre-computed document embeddings in cloud-based vector databases. However, such embeddings are vulnerable to embedding inversion attacks (EIAs), which can reconstruct their underlying text. Existing defenses, such as...

    arxiv.org/abs/2609.04767 · PDF

  68. 68

    DCFA: Dual-view Causal-inspired Attribution for Failure Reasoning in LLM-based Multi-agent Systems

    Zehao Wang, Lanjun Wang, Shilong Jin, Junjie Chen, Yanghua Xiao

    cs.AI

    Large language model (LLM)-based multi-agent systems have experienced rapid growth in recent years. Despite their promise, such systems remain fragile, frequently exhibiting reasoning and coordination errors that can lead to system-level failures. Failure attribution in such systems relies on tracing natural language interactions among agents to identify the decisive error, which refers to the earliest action whose correction can reverse...

    arxiv.org/abs/2609.04749 · PDF

  69. 69

    Aplaud: Adaptive Personalized Low-Rank Decomposition for User-Specific LLM

    Xinyu Li, Ruoming Jin, Jianfeng Zhu, Ruixin Guo, Zhi Liu

    cs.AI

    In this paper, we study the problem of personalized survey response prediction using fine-tuned large language models (LLMs). This task poses unique challenges: limited per-user training data, scalability of model storage, and the need to exploit shared structure across survey questions. To address these issues, we propose Aplaud (Adaptive Personalized Low-rank and User-specific Nested Decomposition), a lightweight and scalable framework for...

    arxiv.org/abs/2609.04738 · PDF

  70. 70

    PLUME: Parameter-Efficient Personalization of Large Language Models via Low-Rank User Modulation in Shared Subspaces

    Xinyu Li, Hao Zhou, Jianfeng Zhu, Julina Maharjan, Ruixin Guo, Feodor Dragan, Ruoming Jin

    cs.AI

    Personalizing large language models (LLMs) is essential for delivering AI assistance that aligns with individual users' styles, intents, and preferences. While per-user fine-tuning can substantially enhance personalization quality, it introduces significant parameter and storage overhead, limiting scalability to large user populations. We propose PLUME (Personalized Low-Rank Adaptation through User Modulation and Shared Subspace), a...

    arxiv.org/abs/2609.04715 · PDF

  71. 71

    FinalityBench: An Effect-Level Benchmark for Agent Decisions Under Delayed and Conflicting Financial Finality

    Abhishek Sharma

    cs.AI

    A merchant's payment processor, ledger, ERP and bank feed are updated by messages that get delayed, duplicated, dropped and reordered, so for minutes at a time the four hold contradictory beliefs about the same order. An agent resolving the exception must decide whether to ship goods, re-submit a capture, refund or wait, knowing some of those cannot be undone. We present FinalityBench, an executable benchmark for that decision. It keeps a...

    arxiv.org/abs/2609.04706 · PDF

  72. 72

    Model Retirement Creates Reproducibility Risk in Biomedical AI Publications

    Nathan Wolfrath, Meghan Conroy, Thomas Kosten, Dave Bell, Bhabishya Neupane, Jonah Kindel, Anjishnu Banerjee, Priya...

    cs.AI

    Background. Large language models (LLMs) are being adopted in biomedical research at a rapid and accelerating pace, yet commercial services that host many widely used models operate under deprecation schedules that can complicate scientific reproducibility. Methods. We searched PubMed for original research articles from 2022 through March 2026 that applied a specific LLM to a biomedical task. An extraction agent identified model names from...

    arxiv.org/abs/2609.04699 · PDF

  73. 73

    SQL-Zero: Self-Evolving Text-to-SQL

    Daniel Machado Pedrozo, Julia Soares Dollis, Bryan Lincoln Marques de Oliveira, Vinicius Alboneti Aguiar, Sávio...

    cs.AI

    Training a competitive Text-to-SQL agent usually depends on human-annotated natural-language/SQL pairs, which are expensive, domain-specific, and a bottleneck for scaling to new databases. We show it is possible to train a competitive solver with zero annotated pairs. We introduce SQL-Zero, a proposer-solver self-play in which a challenger and a solver start from the same base LLM and the only ground truth is execution against the database...

    arxiv.org/abs/2609.04697 · PDF

  74. 74

    Predicting Spatiotemporal Mobile Sensing-Based PM2.5 Concentrations Using Low-Rank Adapted Spatially Attentive Graph Neural Network

    Om Chiddarwar, Priyanka Mandal, Praveen Kumar Chandaliya, Shriniwas Arkatkar

    cs.AI

    Urban air quality can vary significantly along transit corridors, necessitating high-resolution monitoring. This work introduces a novel mobile-sensing dataset from Surat, Gujarat, India, comprising PM$*{2.5}$ concentrations, meteorological variables (temperature, humidity, wind speed, wind direction), and land-use features. To represent the spatiotemporal data as a graph, two node-definition strategies were used: (i) uniform segmentation...

    arxiv.org/abs/2609.04693 · PDF

  75. 75

    Train What You Deploy:Token-Faithful Post-Training of a Production Coding

    Cheng Li, Jiexiong Liu, Yixuan Chen, Chi Hong

    cs.AI

    Existing post-training pipelines for coding and terminal agents suffer severe token and control fidelity errors: simplified training environments mismatch production deployments, and offline token reconstruction from agent logs distorts original prompts and conflates policy calls with background model operations. We present a fidelity-aware training coupling framework that retains trainer-side sampling over original prompts, eliminates...

    arxiv.org/abs/2609.04678 · PDF

  76. 76

    ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies

    Xinran Zhang, Pengrui Lu, Lyumanshan Ye, Pengfei Liu

    cs.AI

    Large language model (LLM) agents are increasingly proposed for enterprise workflows, yet existing evaluations rarely test whether business-decision conclusions transfer across competitive market ecologies. We introduce ERPBench, an execution-instrumented benchmark for enterprise decision agents in a six-round Enterprise Resource Planning (ERP) simulation with coupled pricing, production, procurement, inventory, finance, and shared-market...

    arxiv.org/abs/2609.04667 · PDF

  77. 77

    Harness-agnostic detection and immunization of reward hacking in self-evolving language models

    Rongxin Yang, Yang Liu, Shang Luo, Haoxuan Jia, Chongyang Zhang, Hao Zheng, Yingguang Yang, Yulin Huang, Jianshen...

    cs.AI

    Self-evolving language models improve by proposing candidate updates and keeping whatever raises a visible score. When that score is an imperfect proxy for the capability one actually wants, sustained selection widens the gap between the two. This is reward hacking. We introduce HackProbe, a monitor that attaches to an arbitrary self-evolving loop through two black-box hooks, with no access to weights or activations. It keeps a secret,...

    arxiv.org/abs/2609.04665 · PDF

  78. 78

    Continual Graph Memory for Adaptive Recommendation under Intent Drift

    Hao Nguyen Ngoc, Tung Nguyen, Nguyen Thi Hanh, Hoang Thai Dinh, Nguyen Xuan Tung

    cs.AI

    This paper studies adaptive recommendation under intent drift, where feedback from each recommendation outcome can reveal whether the relational evidence used for ranking is useful, missing, or misleading. While Knowledge Graphs (KGs) provide essential semantic structure to handle these shifts, traditional KG-enhanced systems treat the graph as a static retrieval substrate, making it brittle to evolving intents, noisy metadata, and recurring...

    arxiv.org/abs/2609.04651 · PDF

This edition is part of The Daily Abstract — cs.AI archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.