cs.SE · 2026-07-15 · No. 54

Software Engineering, 2026-07-15.

9 new papers in cs.SE. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

9 entries
  1. 01

    Form, Not Content? A Preregistered, Placebo-Controlled Evaluation of Learned Error-Conditioned Self-Repair Through Prompts and Weights in Frozen Small Code Models

    Mehmet Iscan

    cs.SE · cs.AI · cs.LG

    Frozen small code LLMs are deployed locally, yet the information guiding a retry after a failed attempt is still measured without placebo controls in the self-repair literature. We treat a failed program as a conjecture and an execution counterexample as an oracle-relative refutation, and introduce PoPE (Popperian Placebo-controlled Evaluation): a methodology for measuring whether evidence that falsifies LLM-generated code can be used...

    arxiv.org/abs/2607.12962 · PDF

  2. 02

    Deep4ge: DNN Training Trajectories for Fault Detection and Diagnosis

    Sigma Jahan

    cs.SE · cs.LG

    Deep learning systems often fail due to subtle implementation faults that alter training behavior. Recent work has studied how to detect and diagnose such failures from changes observed across training epochs. However, the software engineering community still lacks a public dataset of per-epoch training runs with documented fault history, feature extraction details, and clear reuse support for fault detection and diagnosis tasks. We present...

    arxiv.org/abs/2607.12868 · PDF

  3. 03

    Toward Localizing and Repairing Bias in Transformer Attention Heads

    Sigma Jahan

    cs.SE · cs.LG

    Transformer language models are increasingly used as software components, yet biased outputs remain difficult to localize and repair inside the model. Existing fairness testing and repair methods largely operate at the input-output or retraining level, while recent work suggests that bias-related behavior can concentrate in a small set of attention heads. This paper studies whether attention heads can be localized and repaired through a...

    arxiv.org/abs/2607.12863 · PDF

  4. 04

    Line-Anchored Feedback Cuts Token Costs and Improves Correctness in AI Code Editing

    William Franz Lamberti

    cs.SE · cs.AI

    Generated tokens are a direct driver of the cost, latency, and energy of generative AI (GAI) code editing. We show the format of feedback is a lever on all three. We compare two deliveries of the same requested changes: a holistic prompt (control) versus the structured, line-anchored export of FileMark (treatment). FileMark is a VSCodium extension for inline comments on any file. In a paired experiment line anchoring cut generated tokens by...

    arxiv.org/abs/2607.12713 · PDF

  5. 05

    Multi-Perspective Agentic Program Repair via Code Property Graphs and Temporal Execution Graphs

    Zhili Huang, Ling Xu, Hongyu Zhang

    cs.SE · cs.AI

    Large language models (LLMs) have improved automated program repair (APR), but two limitations remain. First, raw execution traces are often too large and repetitive to serve as effective model context. Second, repeated patch sampling may produce different implementations without yielding distinct root-cause hypotheses or repair strategies. We present CT-Repair, an agentic APR framework representing static and dynamic evidence as queryable...

    arxiv.org/abs/2607.12605 · PDF

  6. 06

    Agent-Safety Evaluations as Load-Bearing Evidence: A Vendor-Neutral, Cross-Harness Reconstructability Metric

    Oleg Solozobov

    cs.SE · cs.AI

    Many agent-safety evaluation results are not yet load-bearing evidence: identical nominal outcomes (task success, attack success, monitor scores) may sit atop materially different evidence regimes. No vendor-neutral, runnable instrument scores reconstructability as an evaluation-validity metric: whether captured evidence can reconstruct the decision a claim depends on. This paper introduces a property-level reconstructability metric over...

    arxiv.org/abs/2607.12469 · PDF

  7. 07

    Code-MUE: Measuring Code LLMs' Uncertainty through Execution-based Semantic Interaction Graphs

    Xiaoning Ren, Yinxing Xue, Lei Ma, Yuheng Huang

    cs.SE · cs.AI · cs.CL

    As Code Large Language Models (LLMs) become central to modern software engineering, their inherent stochasticity poses significant real-world risks, where even minor errors can lead to severe functional, security, or safety consequences. Reliable automation, therefore, demands the ability to distinguish between confident, well-supported predictions and stochastic guessing. However, existing uncertainty estimation methods face a critical gap:...

    arxiv.org/abs/2607.12273 · PDF

  8. 08

    TraceSynth: Generating Production-Quality Kernel Traces with Constraint-Guided Diffusion Models

    Yuvraj Sehgal, Sneh Patel, Mahsa Panahandeh, Naser Ezzati-Jivan, Francois Tetreault

    cs.SE · cs.LG

    Machine learning models for system diagnostics rely on kernel execution traces to capture fine-grained system behavior, but collecting production traces in industrial systems is costly due to runtime overhead, storage demands, and privacy constraints. We present TraceSynth, a diffusion-based framework for generating synthetic kernel traces that augment limited real data for downstream ML tasks. TraceSynth models traces as multi-channel...

    arxiv.org/abs/2607.12104 · PDF

  9. 09

    AutoTrace: From Patches to Triggers via Agentic Interprocedural Exploration

    Arastoo Zibaeirad, Marco Vieira, Thomas Zimmermann

    cs.SE · cs.AI · cs.CR

    Given a vulnerability-fixing commit, trigger localization asks which specific statement turns the vulnerable program state into a concrete unsafe operation. This question is harder than binary vulnerability detection because the answer demands interprocedural, causal reasoning: in a substantial fraction of real-world CVEs the triggering statement lies several call layers outside the patched function, beyond the reach of static rule sets and...

    arxiv.org/abs/2607.12058 · PDF

This edition is part of The Daily Abstract — cs.SE archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.