eess.AS · 2026-06-19 · No. 28

Audio and Speech Processing, 2026-06-19.

4 new papers in eess.AS. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

4 entries
  1. 01

    Repurposing a Speech Classifier for Guided Diffusion-Based Speech Generation

    Rostislav Makarov, Timo Gerkmann

    eess.AS · cs.AI · cs.LG

    Classifier guidance is a way to control diffusion generation by using a noise-conditioned classifier to steer the sampling process toward a target class. One drawback of classifier guidance is that it requires two separately trained models: a classifier and a diffusion model. We therefore study a more compact alternative in which a conventionally trained speech classifier is repurposed as the backbone for diffusion generation. Starting from a...

    arxiv.org/abs/2606.20457 · PDF

  2. 02

    PASQA: Pitch-Accent-Focused Speech Quality Assessment Model Trained on Synthetic Speech with Accent Errors

    Masaya Kawamura, Yuma Shirahata, Kentaro Mitsui, Reo Shimizu

    eess.AS · cs.CL · cs.LG · cs.SD

    Existing mean opinion score (MOS) prediction models typically predict utterance-level naturalness MOS and can be insensitive to localized pitch-accent errors. We propose Pitch-Accent-focused Speech Quality Assessment (PASQA), which explicitly targets pitch-accent correctness. To train our model, we construct a controlled Japanese accent-error dataset by changing accent patterns using an accent-controllable text-to-speech system, and compute a...

    arxiv.org/abs/2606.20137 · PDF

  3. 03

    Investigating Human-Model Discrepancies in Speech Quality Assessment via Acoustic and Prosodic Perturbations

    Masato Takagi, Masaya Kawamura, Reo Shimizu, Yuma Shirahata

    eess.AS · cs.CL · cs.LG · cs.SD

    Mean opinion score (MOS) prediction models are widely used as proxy metrics in text-to-speech (TTS) research, yet their ability to capture quality differences beyond acoustic fidelity remains unclear. We investigate this via controlled perturbations on speech: acoustic degradation, prosodic errors, and manipulation of speaker-specific characteristics such as pitch and speaking rate. We obtained MOS predictions for these speech samples from...

    arxiv.org/abs/2606.19951 · PDF

  4. 04

    Low-Burden Data Augmentation for Dysarthric ASR via Zero-Shot Voice Cloning

    Satwinder Singh, Qianli Wang, Zihan Zhong, Clarion Mendes, Hasegawa-Johnson, Waleed Abdulla, Seyed Reza Shahamiri

    eess.AS · cs.LG

    Automatic speech recognition remains unreliable for dysarthric speech due to data scarcity and high inter-speaker variability. While synthetic data can address these gaps, traditional methods often require extensive speaker-specific data, reintroducing the collection bottleneck. We investigate zero-shot voice cloning as a low-burden augmentation strategy, using Higgs Audio V2 to clone speakers in the TORGO dataset. We fine-tune (FT)...

    arxiv.org/abs/2606.19823 · PDF

This edition is part of The Daily Abstract — eess.AS archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.