cs.SE · 2026-09-15 · No. 114

Software Engineering, 2026-09-15.

4 new papers in cs.SE. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

4 entries
  1. 01

    IWC-Bench: Evaluating Web Application Generation from a Software Testing Perspective

    Chenxu Liu, Zilu Zou, Peizhong Gao, Jiawen Tao, Zhexin Zhang, Guang Chen, Haowei Lin, Ying Zhou, Tianyi Bai, Dolly...

    cs.SE · cs.AI

    Human evaluation provides a direct measure of the quality of LLM-generated web applications. However, fitting human judgments through automated evaluation remains challenging. Static benchmarks can credit functionality that exists in source code but is unreachable at runtime. Interactive benchmarks exercise the application, yet incomplete exploration can cause them to miss implemented functionality and confound application defects with agent...

    arxiv.org/abs/2609.15387 · PDF

  2. 02

    Failure-Guided Co-Evolution of Prompts and Training Data

    Tianyu Yuan, Zhuzhong Qian

    cs.SE · cs.AI

    Automatic prompt optimization (APO) improves language-model programs by revising prompts from task feedback, yet it typically holds its training data fixed. Repeatedly optimizing against the same instances confines feedback to weaknesses already represented in those data, leaving related failure conditions unexplored. We therefore view each failure as a dual signal: it indicates both how the prompt should be revised and what new training...

    arxiv.org/abs/2609.15209 · PDF

  3. 03

    woma: a real-time foundation model and its fine-tuned models for endoscopy

    Thang Tran, Lan Dang

    cs.SE · cs.CV · cs.LG

    woma is a real-time foundation model for gastrointestinal endoscopy: a network trained without labels on about a million endoscopy frames, from which task models are fine-tuned. We contribute a systematic design for production. Requirements and pass marks were fixed before any run, eight candidates screened under pre-registered rules, self-supervised training taken to a stopping rule, then fine-tuning and deployment optimisation, all on one...

    arxiv.org/abs/2609.15130 · PDF

  4. 04

    DepthBenchCAD: When Does Deeper Auditing Yield More Reliable Conclusions?

    Hongye Yang, Zhihao Xie, Shengjun Xiong, Boxiao Huang

    cs.SE · cs.AI

    Generative CAD models are expected to remain behaviorally correct after parameter edits, so increasing the number of edit checks is often treated as a direct route to more reliable evaluation. Under a fixed budget, however, auditing each program more thoroughly reduces the number of tasks and independent generations that can be evaluated, which can ultimately make model-level estimates less accurate. We study this phenomenon and the...

    arxiv.org/abs/2609.15122 · PDF

This edition is part of The Daily Abstract — cs.SE archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.