eess.AS · 2026-07-29 · No. 68

Audio and Speech Processing, 2026-07-29.

3 new papers in eess.AS. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

3 entries
  1. 01

    Depression Markers in Speech: An Approach based on Tract Variables Dynamics

    Sahar Altalhi, Tanaya Guha, Alessandro Vinciarelli

    eess.AS · cs.AI

    This study identifies new depression biomarkers based on the dynamical properties of tract variables, which represent geometric features describing the configuration of the speech articulators. A key advantage of this approach lies in its ability to quantify aspects of the articulatory process that have not been previously explored in the context of depression, namely predictability, complexity, and randomness. These properties are...

    arxiv.org/abs/2607.25888 · PDF

  2. 02

    Device Invariance using Domain Adaptation on Acoustic Scene Classification

    Abhishek dileep, Shubham Sharma, Padmanabhan Rajan

    eess.AS · cs.AI · cs.SD

    This paper explores the effectiveness of domain adaptation techniques when using convolutional neural network (CNN)-based and transformer-based feature representations for acoustic scene classification. Two well-known domain adaptation techniques, namely domain adversarial neural network (also called DANN) and conditional domain adversarial network (also called CDAN) are evaluated under various domain shifts. Our study indicates that DANN...

    arxiv.org/abs/2607.25887 · PDF

  3. 03

    VAD to the Bone: Ultra-Tiny Speech Activity Detection for Edge Deployment

    Stephen Bauer, Sheila Seidel, Shanza Iftikhar, Scott Veidenheimer, Gorkem Ulkar

    eess.AS · cs.LG

    Voice activity detection (VAD) triggers downstream speech processing in always-on systems under strict memory, latency, and compute constraints. Recent compact models report strong accuracy but rely on components that are not widely supported: learnable filterbanks, recurrent layers, or non-causal post-processing. We propose kiloVAD, designed for embedded inference using standard Mel features, CNN-only layers, and tunable context/spectral...

    arxiv.org/abs/2607.25870 · PDF

This edition is part of The Daily Abstract — eess.AS archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.