cs.SE · 2026-09-24 · No. 123

Software Engineering, 2026-09-24.

7 new papers in cs.SE. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

7 entries
  1. 01

    Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark

    Hamed Taherkhani, Mohammad Abdollahi, Melika Sepidband, Hridya Dhulipala, Tien N. Nguyen, Hadi Hemmati

    cs.SE · cs.AI · cs.CL

    Large language models (LLMs) are increasingly used in coding tasks, but their ability to reason about code execution remains unclear. Existing repository-level QA benchmarks mainly evaluate static code understanding and often rely on LLM-based evaluation, while execution-reasoning benchmarks are mostly limited to snippets or functions. We introduce SWE-Flux, a repository-level benchmark for dynamic execution reasoning containing 480...

    arxiv.org/abs/2609.28449 · PDF

  2. 02

    From Agent Output to Authorized Transition

    Christopher Koch

    cs.SE · cs.AI

    Agentic engineering systems can edit repositories, run tools and tests, build firmware, synthesize schematics, and prepare deployable or manufacturable artifacts. The assurance problem is therefore shifting from whether an agent can produce an output to whether an engineering lifecycle is justified in acting on claims about that output. Current products and standards provide sandboxes, approvals, hooks, traces, policy enforcement,...

    arxiv.org/abs/2609.28216 · PDF

  3. 03

    FDE-Bench: Evaluating LLM Agents for Deployment Environment Configuration

    Weihang Ding, Junfei Zhan, Yueting Li, Qirong Guo

    cs.SE · cs.AI

    Deployment requires an agent to turn application code into a running system whose services connect, become ready, and remain observable. FDE-Bench evaluates this capability with 136 deployment-configuration tasks spanning Docker images, multi-service Compose stacks, and Kubernetes, in greenfield and diagnose-and-repair modes. Agents submit declarative artifacts that are collected, rebuilt, and redeployed in a pristine environment. Four gated...

    arxiv.org/abs/2609.27571 · PDF

  4. 04

    Constraint-Driven Context Engineering: Designing Domain Interfaces for AI Systems

    Xiwei Xu, Chen Wang, Mengmeng Yang, Yipeng Zhang, Jacky Jiang, Suyu Ma, Youyang Qu, Ming Ding, Liming Zhu

    cs.SE · cs.AI

    Generative AI systems are increasingly deployed to address domain problems. These systems operate under technical, regulatory, institutional, and normative constraints that define acceptable AI behaviour and outcomes within their domains. We observe a recurring pattern in our industry engagement: partners often arrive with a functioning but relatively generic AI solution. The challenge is no longer to build an AI system from scratch, but to...

    arxiv.org/abs/2609.27354 · PDF

  5. 05

    FairTest: Search-Based Fairness Testing for Multi-Agent Reinforcement Learning Systems

    Xiaotong Wang, Xuan Xie

    cs.SE · cs.LG

    Multi-agent Reinforcement Learning (MARL) trains a team of agents that share one environment and learn their policies together. Training maximizes the team return, and a high return does not imply that the rewards are shared fairly among the agents in every episode. Testing is an established way to discover the failures of deep reinforcement learning, yet few methods address the fairness of MARL. In this work, we propose FairTest, a...

    arxiv.org/abs/2609.27309 · PDF

  6. 06

    Teach-to-Crash: A Closed-Loop Student-Teacher LLM Framework for Collision-Inducing Test Scenario Generation

    Zaid Ghazal, Khouloud Gaaloul, Bruce Maxim

    cs.SE · cs.AI · cs.MA · eess.SY

    Validating Autonomous Driving Systems (ADS) in simulation requires testing architectures that can discover rare, safety-critical failures while generating scenarios that are executable, diverse, and useful for downstream failure analysis. We introduce Teach-to-Crash, a closed-loop testing framework that combines a constrained ego-centric scenario representation, stagnation-aware search control, and a dual-LLM architecture for adaptive failure...

    arxiv.org/abs/2609.27296 · PDF

  7. 07

    From PyTorch to the NPU: LLM-Agent-Driven Model Conversion Across Heterogeneous Inference Runtimes

    Jianhao Su, Zhanwei Wu, Chia-Heng Tu, ShengTing Huang

    cs.SE · cs.DC

    Edge AI model deployment is a multi-stage engineering process involving model conversion, operator compatibility handling, runtime integration, and precision verification. While prior work has demonstrated agent-based automation for Qualcomm AI Runtime, the broader edge inference runtime ecosystem, including Intel OpenVINO, Rockchip RKNN, NVIDIA TensorRT, and ONNX Runtime, presents distinct toolchains and optimization strategies. This paper...

    arxiv.org/abs/2609.27249 · PDF

This edition is part of The Daily Abstract — cs.SE archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.