cs.SD · 2026-09-17 · No. 116

Sound, 2026-09-17.

5 new papers in cs.SD. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

5 entries
  1. 01

    Beyond EER: Multi-Dimensional Evaluation of Information Leakage in Speaker De-Identification

    Seungmin Seo, Oleg Aulov, P. Jonathon Phillips, Kevin Mangold, Jonathan Eskin

    cs.SD · cs.AI

    Speaker de-identification (SDID) aims to preserve privacy by concealing speaker identity while maintaining speech utility. However, current evaluations often reduce privacy to a single dimension - biometric verification performance - typically measured by Equal Error Rate (EER). This narrow focus ignores critical leakage channels, such as soft biometric inference, embedding-level re-identification, and structural template similarity, which...

    arxiv.org/abs/2609.18673 · PDF

  2. 02

    TTM-Bench: A Framework for Text-to-Music System Performance Benchmarking

    Giorgia Adorni, Michela Papandrea, Battista Rimoldi, Tiziano Leidi

    cs.SD · cs.LG · cs.MM · eess.AS

    Text-to-music (TTM) systems are increasingly used to generate musical audio from natural-language descriptions. Robust evaluation is therefore essential, yet reliable performance comparison remains challenging. This difficulty stems from differences in system architecture, supported conditioning information, and access mode, as well as heterogeneous and fragmented metrics that cannot be applied uniformly across systems. To address these...

    arxiv.org/abs/2609.18585 · PDF

  3. 03

    VoiceTrace: A Benchmark and Retrieval Framework for Who-Said-What Speech Retrieval

    Aaron Yee, Fengjie Lu, Jiarui Hai, Chenang Jiang, Helin Wang, Siwei Tu, Weitao You, Lingyun Sun

    cs.SD · cs.AI · eess.AS

    Speech retrieval has become increasingly important as spoken content continues to grow across meetings, lectures, podcasts, and videos. Existing benchmarks and models have advanced semantic search over spoken content, but largely focus on \emph{what} is said while overlooking \emph{who} says it. In many real-world scenarios, however, users need to retrieve speech based jointly on semantic content and a target speaker, where the speaker may be...

    arxiv.org/abs/2609.18521 · PDF

  4. 04

    Look Less, Hear Better: Jointly Rewarded GRPO for Streaming ASR

    Xiuwen Zheng

    cs.SD · cs.AI

    Streaming automatic speech recognition (ASR) must be judged jointly on what it transcribes and on how quickly it commits each word. Delayed streams modeling (DSM) has become the dominant paradigm for streaming large audio-language models, exposing a structural delay $τ$ that bounds the decoder's lookahead. We show that $τ$ is a poor proxy for user-perceived latency, and that the alignment-based supervision of DSM leaves latency on the table:...

    arxiv.org/abs/2609.18333 · PDF

  5. 05

    CPR: Combining global composing, local performing and full-sequence refining in piano rendering with continuous autoregressive modelling

    Chong Jing, Junan Zhang, Zhizheng Wu

    cs.SD · cs.AI

    Prompt-conditioned piano MIDI-to-Music rendering aims to faithfully render target notes while reproducing the timbre of a reference recording. Existing approaches primarily follow two paradigms: autoregressive (AR) modeling and flow matching (or diffusion). Discrete-codec AR models provide causal temporal modeling, but quantization can discard acoustic detail. Flow matching better preserves acoustic structure in the cost of full-sequence...

    arxiv.org/abs/2609.18216 · PDF

This edition is part of The Daily Abstract — cs.SD archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.