cs.SD · 2026-09-18 · No. 117

Sound, 2026-09-18.

3 new papers in cs.SD. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

3 entries
  1. 01

    Multi-Dimensional Prosody Judgment For Live Streaming Speech Synthesis

    Zifan Guan, Longyu Lu, Junan Zhang, Zhizheng Wu, Meiguang Jin, Junfeng Ma

    cs.SD · cs.AI

    Evaluating live streaming speech synthesis (TTS) requires assessing fine-grained, highly expressive prosody such as emotion, intonation, and energy which traditional MOS predictors fail to capture. While proprietary Large Language Models (LLMs) like Gemini can evaluate these aspects, they are too costly for massive inference and reinforcement learning feedback. To address this, we first introduce Live-ProsodyJudge (LPJ), a cost-effective...

    arxiv.org/abs/2609.20124 · PDF

  2. 02

    Robust Workflow Generation via Adversarial Learning for Audio Deepfake Detection

    Xiang Li, Pin-Yu Chen, Wenqi Wei

    cs.SD · cs.AI

    The rapid advancement of speech synthesis and voice conversion technologies has made audio deepfakes increasingly realistic, posing serious security risks in practical applications. While existing detection methods achieve strong performance under controlled conditions, they often fail to generalize under real-world perturbations and corruptions. In this paper, we propose ROGUE, a framework that dynamically constructs robust detection...

    arxiv.org/abs/2609.20063 · PDF

  3. 03

    CoRELoop: Parameter-Efficient Controlled Recurrent Refinement for Audio Deepfake Detection

    Kunyu Feng, Yuxiang Wang, Li Wang, Wan Lin, Zhizheng Wu

    cs.SD · cs.AI

    Generalizing to unseen attacks remains challenging for audio deepfake detectors, and collecting training data covering all potential attacks is impractical. We explore recurrent refinement in an already-trained SSL-based detector without additional data or changes to its original parameters. However, directly recycling encoder outputs as inputs degrades detection in our diagnostic. We propose CoReLoop, which makes this reuse effective by...

    arxiv.org/abs/2609.19818 · PDF

This edition is part of The Daily Abstract — cs.SD archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.