eess.AS · 2026-09-12 · No. 113
Audio and Speech Processing, 2026-09-12.
4 new papers in eess.AS. Titles, authors,
abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →
01 — The papers
4 entries-
01
RetroThinker: Enabling Retrospective Thinking in Speech LLMs
Yi-Jen Shih, Puyuan Peng, Abdelrahman Mohamed, David Harwath
eess.AS · cs.AI · cs.CL
Speech large language models (SpeechLLMs) offer reduced latency and retain paralinguistic nuances that are typically lost in cascaded automatic speech recognition (ASR) and text-based LM architectures. However, they continue to lag behind text-only LLMs on complex reasoning tasks, while real-time spoken interaction imposes strict latency constraints. Although prior works employ Chain-of-Thought (CoT) and concurrent reasoning to enhance...
-
02
Investigating catastrophic forgetting in sound event classification
Riccardo Casciotti, Annamaria Mesaros
eess.AS · cs.AI · cs.SD
This work investigates a number of approaches to prevent catastrophic forgetting in class incremental learning scenarios for sound event classification tasks. We analyze the problem using architectural and regularization approaches, using FSD50K and AudioSet datasets. We design incremental stages and solutions that selectively protect the kernels of the network from weight updates to prevent catastrophic forgetting, and a dynamic head...
-
03
Exploring Second-Order Pattern Recognition in Speaker Recognition
Yanze Xu, Wenwu Wang, Mark D. Plumbley
eess.AS · cs.AI
In classical pattern recognition tasks, neural networks are trained to recognise human-defined patterns for model inputs. Some Explainable AI (XAI) methods can explain other latent patterns that underlie the network's recognition of inputs as human-defined patterns; in this work, we call these latent patterns second-order patterns, and we propose to discover them. To this end, we apply a hierarchical clustering algorithm to analyse whether...
-
04
Less can be More: What Aspects of Speech Drive End-of-Turn Detection
Rini Sharon, Manickavela A, Kadri Hacioglu, Andreas Stolcke
eess.AS · cs.AI · cs.SD
In conversational AI, detecting when a speaker has finished talking is crucial for natural turn taking. While recent work incorporates semantics, the relative contribution of different modalities remains unclear. We present a controlled ablation of acoustic, prosodic, and semantic signals for streaming end of turn detection using a lightweight trimodal classifier. Under identical training conditions, the acoustic prosodic combination achieves...
This edition is part of The Daily Abstract — eess.AS archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.
#D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.