cs.SD · 2026-08-24 · No. 94

Sound, 2026-08-24.

3 new papers in cs.SD. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

3 entries
  1. 01

    DAMOS: Learning Distortion-Aware Speech Quality Assessment through Explicit Distortion Localization

    Naiyuan Li, Li Dong, Diqun Yan

    cs.SD · cs.AI

    Automatic speech quality assessment aims to predict Mean Opinion Scores (MOS) consistent with human subjective perception and is essential for evaluating speech generation, enhancement, and communication systems. For speech signals, especially synthetic speech, distortions often occur locally, and overall perceptual quality is usually dominated by a small number of perceptually salient distortion regions. However, most existing methods are...

    arxiv.org/abs/2608.21176 · PDF

  2. 02

    AudioWorldSim: Realistic Binaural Audio Datasets For World Models

    Luis Vitor Zerkowski, Luiz Velho

    cs.SD · cs.LG

    This technical report presents AudioWorldSim, an open-source platform designed to generate realistic binaural audio datasets and advance research in audio-based machine learning, particularly world models. Built as a custom extension of Meta's SoundSpaces 2.0 platform, AudioWorldSim leverages their comprehensive acoustics framework, but focuses on the automatic rollout of random agent navigations, as well as implements crucial fixes to how...

    arxiv.org/abs/2608.21075 · PDF

  3. 03

    Do SpeechLMs Hear Their Own Opinions? Diagnosing and Mitigating Previous-Belief Contamination in Streaming Emotion Understanding

    Haoyue Liu, Zhichao Wang, Ye Chen, Haonan Deng, Xiaoying Tang

    cs.SD · cs.AI

    Streaming emotion understanding uses historical state while continuously interpreting current audio, often feeding the model's previous prediction back as context. We show that this history conditioning can distort current perception. On a balanced CREMA-D-Stream counterfactual diagnostic, changing only the injected previous emotion label while holding the audio fixed reduces current-audio accuracy from 72.50% to 30.42% and flips 65.69% of...

    arxiv.org/abs/2608.20769 · PDF

This edition is part of The Daily Abstract — cs.SD archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.