cs.SE · 2026-08-19 · No. 89

Software Engineering, 2026-08-19.

5 new papers in cs.SE. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

5 entries
  1. 01

    What Aggregate Scores Miss: Measuring Item-Level Regressions in Commercial LLM API Migrations

    Xiaonan Xu, Wenjing Wu

    cs.SE · cs.AI · cs.CL

    Context: Software systems that depend on commercial large language model APIs must migrate to successor versions when vendors deprecate older models. Migration decisions typically rely on aggregate benchmark scores, which compress heterogeneous item-level behaviour into a single net figure. Objective: We measure what that compression conceals. Method: On three pairwise upgrades in the GPT-5.4 to GPT-5.6 Sol product sequence, we query 900...

    arxiv.org/abs/2608.17719 · PDF

  2. 02

    GADR: Gathering Architecture Decision Records from Meeting Transcriptions

    Lucas Daniel Costa da Silva, Kiev Gama

    cs.SE · cs.AI

    Existing LLM-based approaches to Architecture Decision Record (ADR) generation share a critical and largely unexamined assumption: that input is already reasonably structured. In practice, architectural decisions emerge from informal, noisy meetings where choices are implicit, fragmented, and entangled with off-topic dialogue, precisely the conditions under which single-pass prompting degrades. This paper presents GADR, a multi-agent,...

    arxiv.org/abs/2608.17694 · PDF

  3. 03

    Benchmarking Automated Security Patch Backporting: How Far Are We?

    Jincheng Yang, Yulong Fu, Chengwei Liu, Lyuye Zhang, Fangyuan Zhang, Bingyang Ren, Yang Liu, Hui Li

    cs.SE · cs.AI · cs.CR

    Automated security patch backporting is critical for mitigating N-day vulnerabilities. Recent tools report success rates above 80% on their respective datasets. However, these evaluations are often confined to homogeneous environments, such as one repository or specific project versions. Consequently, it remains unclear how well these tools generalize beyond their originally targeted scenarios. We present Porting Benchmark, a curated dataset...

    arxiv.org/abs/2608.17671 · PDF

  4. 04

    Beyond FLOPs: Energy-Aware Knowledge Distillation for Sustainable LLMs on Code-Related Task

    Enrique Barba Roque, Luís Cruz, Annibale Panichella

    cs.SE · cs.AI

    Background: Large Language Models (LLMs) are increasingly being applied to Software Engineering (SE) tasks, achieving high accuracy across problems such as clone detection, vulnerability prediction, and code summarization. However, their high computational demands and energy consumption raise sustainability concerns and hinder their use on consumer hardware and resource-constrained platforms. A common way to report the computational cost of...

    arxiv.org/abs/2608.17515 · PDF

  5. 05

    Graphectory Viewer: A Tool for Process-Centric Analysis of Agentic Software Trajectories

    Charlie Jyu, Shuyang Liu, Reyhaneh Jabbarvand

    cs.SE · cs.AI

    We present Graphectory Viewer, a web-based tool for interactive, process-centric analysis of software-agent trajectories. Building on the Graphectory representation introduced in our previous work, Graphectory Viewer transforms heterogeneous raw trajectories into phase-aware graphs that connect low-level execution details with higher-level behavioral structures. The tool supports trajectories from multiple agent frameworks and provides...

    arxiv.org/abs/2608.17195 · PDF

This edition is part of The Daily Abstract — cs.SE archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.