cs.SE · 2026-07-14 · No. 53

Software Engineering, 2026-07-14.

7 new papers in cs.SE. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

7 entries
  1. 01

    Evaluating RE Practices for Explainability: Synthesizing Insights from Daimler Truck into an Explainable RE Framework Proposal

    Umm-e- Habiba, Lucas Mauser, Jonas Fritzsch, Justus Bogner, Stefan Wagner

    cs.SE · cs.AI

    Explainability has emerged as a critical requirement for AI-based systems, particularly in safety-critical and regulated domains. Although prior research has proposed frameworks, patterns, and user-centered approaches to support explainability, there is limited empirical understanding of how existing Requirements Engineering (RE) practices support explainability requirements across the RE lifecycle, especially in an industrial context. This...

    arxiv.org/abs/2607.11771 · PDF

  2. 02

    Understanding the Impact of AI Code Assistants on Security API Usage: An Empirical Study

    Zahra Mousavi, Chadni Islam, M. Ali Babar, Alsharif Abuadbba, Kristen Moore

    cs.SE · cs.AI

    AI code assistants are transforming software development, but their implications for software security remain a major concern, particularly in the context of security APIs. These APIs are critical for safeguarding software systems, yet their complexity often leads to incorrect use and serious vulnerabilities. Developing an evidence-based understanding of how AI assistants influence developers' use of these APIs is therefore essential for...

    arxiv.org/abs/2607.11348 · PDF

  3. 03

    Fail-Aware and Explainable Test Oracle Prediction

    Yue Zhao, Binish Tanveer, Jelena Zdravkovic

    cs.SE · cs.AI

    Despite their central role in fault detection, test oracles remain challenging to construct effectively. Recent learning based methods address this challenge by automatically generating test assertions, yet even if syntactically correct, they are often ineffective in revealing bugs. Rather than generating assertions, this study explores a different approach by training a model to directly predict whether a given test prefix passes or fails....

    arxiv.org/abs/2607.11342 · PDF

  4. 04

    An Empirical Study for GUI Test Migration from Android to OpenHarmony System

    Yakun Zhang, Xinjia Chen, Yiyun Chen, Yuxia Zhang, Mingyi Zhou, Xiang Gao, Shaokun Zhang, Li Li, Yunming Ye

    cs.SE · cs.AI

    To reduce the substantial engineering effort required to test the corresponding applications from Android to OpenHarmony, migrating existing GUI test cases has become a critical problem. However, current research neither proposes solutions tailored for OpenHarmony nor provides a systematic evaluation of migration approaches on this system, leaving developers with limited empirical guidance in practice. In this paper, we present the first...

    arxiv.org/abs/2607.11245 · PDF

  5. 05

    RepTran: Search-Based Repair of Transformer Models

    Yuta Ishimoto, Paolo Arcaini, Fuyuki Ishikawa, Masanari Kondo, Naoyasu Ubayashi, Yasutaka Kamei

    cs.SE · cs.AI

    To ensure the overall quality of AI-enabled software, not only traditional software components but also AI components need to be tested and repaired. Among AI components, Transformer models are increasingly integrated into software systems, which makes their misbehaviors critical. Although prior work in the software engineering community has proposed deep neural network (DNN) repair methods, most overlook Transformer-specific structures. We...

    arxiv.org/abs/2607.11193 · PDF

  6. 06

    AgentCheck: A Reproduce-Intervene-Mitigate Workbench for LLM Agents over MCP

    Aritra Mazumder, Nusrat jahan Lia

    cs.SE · cs.AI · cs.CL

    Tool-using LLM agents are mostly evaluated assuming all tools work. When a tool times out, returns a week-stale value, or has its description poisoned in deployment, the developer needs a controlled way to reproduce the failure, test a fix, and confirm the fix worked before deployment. We present AgentCheck, an open-source web workbench that turns an MCP server into an intervention surface. AgentCheck runs an agent against its real tools and...

    arxiv.org/abs/2607.11098 · PDF

  7. 07

    BackendForge: Benchmarking Agentic End-to-End Code Generation with Backend Services

    Yuzhe Guo, Mengzhou Wu, Yuan Cao, Jialei Wei, Dezhi Ran, Wei Yang, Tao Xie

    cs.SE · cs.AI

    Large language models (LLMs) are increasingly used in agentic coding settings, where they can inspect files, execute commands, run tests, observe failures, and iteratively revise code. This shift raises a central evaluation question: can an agentic LLM generate an end-to-end software artifact that is both deployable and behaviorally correct under execution? Backend services provide a controlled but realistic substrate for this evaluation....

    arxiv.org/abs/2607.11042 · PDF

This edition is part of The Daily Abstract — cs.SE archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.