eess.AS · 2026-09-25 · No. 124
Audio and Speech Processing, 2026-09-25.
5 new papers in eess.AS. Titles, authors,
abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →
01 — The papers
5 entries-
01
Does per-frame early exit pay? A compute-matched study of dynamic depth for on-device speech enhancement
Clément Laroche, Riccardo Miccini
eess.AS · cs.LG · cs.SD
Deep learning-based speech enhancement is increasingly deployed on-device in hearing aids, headsets, and earbuds. Most of these devices, however, can only accelerate static int8 graphs, so a depth-varying network must be implemented as several graphs, orchestrated by a policy. In this paper, we supervise every intermediate depth of one causal model, then we fine-tune its output heads to guarantee that deeper outputs are never worse than...
-
02
Beyond Model Size: Redesigning LiSenNet for embedded speech enhancement
Clément Laroche, Rasmus Kongsgaard Olsson
eess.AS · cs.LG · cs.SD
Deploying real-time speech enhancement on resource-constrained devices requires meeting strict latency, memory, and energy constraints. Microcontroller NPUs can accelerate neural inference under these constraints, but only through a restricted set of operators in static, integer-quantized graphs. Recent speech-enhancement networks have reduced parameter counts and MACs to levels nominally suitable for microcontrollers, but their operators and...
-
03
Anatomy-aware cross-speaker adaptation of complete vocal-tract acoustic-to-articulatory inversion
Nhat-Nam Nguyen, Pierre-Andre Vuissoz, Yves Laprie
eess.AS · cs.AI
Cross-speaker acoustic-to-articulatory inversion requires accounting for anatomical differences between speakers. We propose a geometric adaptation framework that uses anatomical landmarks, primarily on vertebrae and dental structures,to transfer predictions from a fixed inversion model to unseen speakers. An affine transformation followed by thin-plate spline (TPS) deformation maps the predicted contours of 10 vocal-tract structures into...
-
04
Transcript-Supervised Post-Training of Generative Speech Enhancement on Real Recordings via Reinforce Adjoint Matching
Julius Richter, Christoph Boeddeker, Yoshiki Masuyama, Kohei Saijo, Dominik Klement, Gordon Wichern, Jonathan Le Roux
eess.AS · cs.LG
We adapt Reinforce Adjoint Matching (RAM), a reward-based post-training method, to generative speech enhancement (SE). Starting from a pretrained SE model, RAM tilts the model's conditional distribution toward outputs with higher reward. During training, the current model generates enhanced speech on-policy, evaluates each generated endpoint with a potentially non-differentiable reward, and analytically re-noises the endpoint to construct...
-
05
WST-Graph: Topology-Preserving Wavelet Scattering Front-End for Speech Deepfake Detection
Kwok-Ho Ng, Tingting Song, Bingwen Feng, Zhihua Xia
eess.AS · cs.AI · cs.SD
The acoustic front-end determines which forensic cues a speech deepfake detector can exploit. The wavelet scattering transform (WST) provides stable multiscale coefficients with explicit coordinates, yet direct flattening obscures the parent relation between paths. We introduce WST-Graph, reconstructing these paths as a sparse modulation-carrier grid for an AASIST graph backend. Modulation-level normalization and length-aware adaptive local...
This edition is part of The Daily Abstract — eess.AS archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.
#D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.