cs.SE · 2026-07-02 · No. 41

Software Engineering, 2026-07-02.

7 new papers in cs.SE. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

7 entries
  1. 01

    Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?

    Zhi Chen, Zhensu Sun, Yuling Shi, David Lo, Lingxiao Jiang

    cs.SE · cs.AI

    Repository-level performance-optimization benchmarks such as GSO, SWE-Perf and SWE-fficiency evaluate coding agents by applying patches to real repositories and comparing runtime against unoptimized baselines and official reference patches. Their leaderboard scores are increasingly used as evidence of coding-agent progress, but those scores can conflate runtime instability, benchmark-specific scoring rules, and how many tasks are already...

    arxiv.org/abs/2607.01211 · PDF

  2. 02

    Skills Are Not Islands: Measuring Dependency and Risk in Agent Skill Supply Chains

    Changguo Jia, Tianqi Zhao, Runzhi He, Minghui Zhou

    cs.SE · cs.AI

    Agent skills package reusable operational knowledge for Large Language Model (LLM) agents, yet as they grow in scope, they become dependency-bearing artifacts whose identities, versions, and provenance remain implicit. This opacity already causes duplicated dependencies and inconsistent installations, exposing a gap that dependency management has yet to close. We introduce Agent Skill Supply Chains (ASSCs) to characterize mixed...

    arxiv.org/abs/2607.01136 · PDF

  3. 03

    Cheap Code, Costly Judgment: A Case Study on Governable Agentic Software Engineering

    James C. Davis, Paschal C. Amusuo, Tanmay Singla, Berk Çakar, Kirsten A. Davis

    cs.SE · cs.AI

    Generative AI is shifting software engineering from a practice organized around scarce implementation effort toward one organized around abundant, low-cost code production. This shift changes the central engineering problem: not whether AI can generate useful code, but how engineers organize architectures, tools, evidence, and feedback loops so that AI-mediated development remains inspectable, correctable, and maintainable. We study this...

    arxiv.org/abs/2607.01087 · PDF

  4. 04

    SWE-Doctor: Guiding Software Engineering Agents with Runtime Diagnosis from Multi-Faceted Bug Reproduction Tests

    Yaoqi Guo, Yang Liu, Jie M. Zhang, Yun Ma, Yiling Lou, Zhenpeng Chen

    cs.SE · cs.AI

    Large language model (LLM)-based software engineering agents are increasingly developed to resolve software issues by generating patches from issue reports and code repositories. Bug reproduction tests (BRTs) are an important building block for such agents and have been shown useful for patch validation. However, it remains unclear whether BRTs can also help the more central stage of patch generation. We first conduct a preliminary study and...

    arxiv.org/abs/2607.00990 · PDF

  5. 05

    Stochastic Connectivity as the Foundation of a Runtime Model for Microservice Availability Analysis

    Anatoly A. Krasnovsky, Anna Maslovskaya

    cs.SE · cs.DC · cs.PF

    Microservice availability is commonly assessed by fault injection and chaos experiments, but such experiments are costly, operationally risky, and difficult to repeat for every architectural change. Distributed tracing and deployment metadata provide cheaper evidence, yet they usually remain descriptive: they show which services interacted, not what endpoint-level availability property follows. This paper proposes a formal runtime...

    arxiv.org/abs/2607.00740 · PDF

  6. 06

    LLVM-Bench: Benchmarking and Advancing Large Language Models for LLVM Compiler Issue Resolution

    Zhao Tian, Yingquan Zhao, Chenyao Suo, Meng Wang, Junjie Chen

    cs.SE · cs.AI · cs.PL

    LLVM is a widely used compiler infrastructure whose scale and complexity make issue resolution labor-intensive and challenging. Although large language models (LLMs) have recently achieved remarkable success in issue resolution, their effectiveness on complex system-level LLVM compiler remains largely unexplored. To address this gap, we introduce LLVM-Bench, the first large-scale benchmark for LLVM issue resolution, containing 423 real-world,...

    arxiv.org/abs/2607.00700 · PDF

  7. 07

    A Methodology for Investigating AI Patterns Prevalence in Software Repositories

    Srinath Perera, Hasinthaka Piyumal, Frank Leymann, Rania Khalaf

    cs.SE · cs.AI

    As Artificial Intelligence(AI)-based applications take off, a clear understanding of AI patterns can uplift the quality of AI applications. Many AI patterns have been proposed in the literature; however, their prevalence in real-life code has not yet been validated. Understanding the actual use of those patterns in practice can clarify our understanding both of the significance of these patterns and their utility. In this paper, we present a...

    arxiv.org/abs/2607.00558 · PDF

This edition is part of The Daily Abstract — cs.SE archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.