eess.AS · 2026-09-25 · No. 124

Audio and Speech Processing, 2026-09-25.

5 new papers in eess.AS. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

5 entries
  1. 01

    Does per-frame early exit pay? A compute-matched study of dynamic depth for on-device speech enhancement

    Clément Laroche, Riccardo Miccini

    eess.AS · cs.LG · cs.SD

    Deep learning-based speech enhancement is increasingly deployed on-device in hearing aids, headsets, and earbuds. Most of these devices, however, can only accelerate static int8 graphs, so a depth-varying network must be implemented as several graphs, orchestrated by a policy. In this paper, we supervise every intermediate depth of one causal model, then we fine-tune its output heads to guarantee that deeper outputs are never worse than...

    arxiv.org/abs/2609.29867 · PDF

  2. 02

    Beyond Model Size: Redesigning LiSenNet for embedded speech enhancement

    Clément Laroche, Rasmus Kongsgaard Olsson

    eess.AS · cs.LG · cs.SD

    Deploying real-time speech enhancement on resource-constrained devices requires meeting strict latency, memory, and energy constraints. Microcontroller NPUs can accelerate neural inference under these constraints, but only through a restricted set of operators in static, integer-quantized graphs. Recent speech-enhancement networks have reduced parameter counts and MACs to levels nominally suitable for microcontrollers, but their operators and...

    arxiv.org/abs/2609.29866 · PDF

  3. 03

    Anatomy-aware cross-speaker adaptation of complete vocal-tract acoustic-to-articulatory inversion

    Nhat-Nam Nguyen, Pierre-Andre Vuissoz, Yves Laprie

    eess.AS · cs.AI

    Cross-speaker acoustic-to-articulatory inversion requires accounting for anatomical differences between speakers. We propose a geometric adaptation framework that uses anatomical landmarks, primarily on vertebrae and dental structures,to transfer predictions from a fixed inversion model to unseen speakers. An affine transformation followed by thin-plate spline (TPS) deformation maps the predicted contours of 10 vocal-tract structures into...

    arxiv.org/abs/2609.29766 · PDF

  4. 04

    Transcript-Supervised Post-Training of Generative Speech Enhancement on Real Recordings via Reinforce Adjoint Matching

    Julius Richter, Christoph Boeddeker, Yoshiki Masuyama, Kohei Saijo, Dominik Klement, Gordon Wichern, Jonathan Le Roux

    eess.AS · cs.LG

    We adapt Reinforce Adjoint Matching (RAM), a reward-based post-training method, to generative speech enhancement (SE). Starting from a pretrained SE model, RAM tilts the model's conditional distribution toward outputs with higher reward. During training, the current model generates enhanced speech on-policy, evaluates each generated endpoint with a potentially non-differentiable reward, and analytically re-noises the endpoint to construct...

    arxiv.org/abs/2609.29405 · PDF

  5. 05

    WST-Graph: Topology-Preserving Wavelet Scattering Front-End for Speech Deepfake Detection

    Kwok-Ho Ng, Tingting Song, Bingwen Feng, Zhihua Xia

    eess.AS · cs.AI · cs.SD

    The acoustic front-end determines which forensic cues a speech deepfake detector can exploit. The wavelet scattering transform (WST) provides stable multiscale coefficients with explicit coordinates, yet direct flattening obscures the parent relation between paths. We introduce WST-Graph, reconstructing these paths as a sparse modulation-carrier grid for an AASIST graph backend. Modulation-level normalization and length-aware adaptive local...

    arxiv.org/abs/2609.29372 · PDF

This edition is part of The Daily Abstract — eess.AS archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.