cs.SD · 2026-10-05 · No. 134

Sound, 2026-10-05.

3 new papers in cs.SD. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

3 entries
  1. 01

    Revisiting Input Time-frequency Representations in Multi-pitch Estimation for Vocal Ensembles

    Junyoung Koh, Hao-Wen Dong

    cs.SD · cs.AI

    Multi-pitch estimation in vocal ensembles is challenging because singers occupy overlapping pitch ranges and often sing at closely spaced fundamental frequencies, causing their harmonics to overlap in time-frequency representations. Existing models commonly use harmonic constant-Q transform (HCQT)-based representations to provide frequency-adaptive resolution, at the cost of expensive feature extraction when training mixtures are generated on...

    arxiv.org/abs/2610.03656 · PDF

  2. 02

    DriftTTS: Few-Step Text-to-Speech Without Distillation via Distribution-Matching Drift

    Mohammad Nur Hossain Khan, Subrata Biswas, Bashima Islam

    cs.SD · cs.AI

    Few-step neural text-to-speech models often rely on short- ened diffusion or flow-matching schedules, or on distillation from pretrained multi-step teachers. To avoid these depen- dencies, we present DriftTTS, a few-step mel-spectrogram generator trained without a generative teacher, distillation, or adversarial discrimination. DriftTTS uses a distribution- matching drift objective in a mel-domain feature space defined by raw mels and a...

    arxiv.org/abs/2610.03390 · PDF

  3. 03

    ParaGeo: Decomposing Paralinguistic Variation into a Shared Latent Geometry

    Yuhan Liu, Yuxuan Ou, Ruoxi Su, Mohamed Ahmed Zaki, Yunbo Long

    cs.SD · cs.LG

    Speech delivery varies with both the requested paralinguistic attribute and the linguistic content. We introduce ParaGeo, a matched-content decomposition of paralinguistic variation in a frozen speech language model. Synthesized audio tokens are replayed with a fixed listening prompt; pooled key/value (K/V) representations are centered and projected into a shared low-dimensional space. Our GLM-4-Voice probe spans 80 requested controls from 12...

    arxiv.org/abs/2610.03125 · PDF

This edition is part of The Daily Abstract — cs.SD archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.