cs.SE · 2026-09-07 · No. 108

Software Engineering, 2026-09-07.

5 new papers in cs.SE. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

5 entries
  1. 01

    ARIA - An Agentic Framework for Autonomous Testing of Infotainment Systems

    António Azevedo, Bruno Lima, João Pascoal Faria

    cs.SE · cs.AI

    Automotive infotainment validation still relies on manual testing, slow, costly, and incompatible with agile releases and OTA updates. Scripted automation only partly helps: it couples test logic to implementation, yielding brittle, high-maintenance suites. Existing LLM-driven frameworks mostly target web/mobile apps, using single- or dual-agent setups that overload one or two models with perception, planning, action selection, and validation...

    arxiv.org/abs/2609.04913 · PDF

  2. 02

    Better Understanding, Better Fixes? A Study of Hallucination in LLM-based Automated Program Repair

    Xuemeng Cai, Jiakun Liu, Linhan Yang, Wei Ma, Lingxiao Jiang

    cs.SE · cs.AI

    Large language models (LLMs) have significantly advanced automated program repair (APR), yet existing evaluations remain largely result-centric and provide limited insight into hallucination during repair. In APR, hallucination may arise not only in final patches but also in the intermediate artifacts that guide patch generation. To address this gap, we perform a multi-layered analysis of hallucination throughout the APR process....

    arxiv.org/abs/2609.04909 · PDF

  3. 03

    Same Request, Different Answer: Quantization Amplifies Cache-Induced Divergence in LLM Serving

    Aditi Patodiya

    cs.SE · cs.DC · cs.LG

    Prefix caching, in which a serving engine reuses the key and value tensors of a shared prompt prefix across requests, is enabled by default in the major open-source stacks and treated as a transparent optimization. We measure what it costs in reproducibility, and find that the cost rises sharply with weight quantization. Holding the model, decoding parameters, seed, and request order fixed, and issuing every request serially at batch size...

    arxiv.org/abs/2609.04748 · PDF

  4. 04

    Building a research-software catalog with a coding agent: from hackathon prototype to public deployment

    Kazuyoshi Yoshimi, Satoshi Terasaki, Gotai Yamada

    cs.SE · cs.AI · cs.CY · physics.ed-ph

    Generative AI and coding agents can accelerate research software development, but they also increase the need for efficient software discovery and maintenance. We developed a repository catalog during a three-day hackathon and subsequently examined the engineering required to make it suitable for public deployment, including adversarial review, data-quality checks, browser-level validation, and publication safeguards. We then explored whether...

    arxiv.org/abs/2609.04711 · PDF

  5. 05

    Beyond Code Generation: Reliability, Verification, and Cost Economics in the Agentic Software Development Lifecycle

    Happy Bhati

    cs.SE · cs.AI

    AI coding systems are moving from autocomplete and chat toward agents that can inspect repositories, edit multiple files, run tools, write tests, open pull requests, and work for long periods with limited supervision. This capability changes the bottleneck in software delivery. Recent field studies show meaningful gains in coding activity, but newer evidence also shows that those gains attenuate sharply between writing code and shipping...

    arxiv.org/abs/2609.04681 · PDF

This edition is part of The Daily Abstract — cs.SE archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.