cs.SD · 2026-06-13 · No. 22

Sound, 2026-06-13.

3 new papers in cs.SD. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

3 entries
  1. 01

    Generative Modeling of Bach-Style Symbolic Music: A Comparative Study of Autoregressive, Latent-Variable, and Adversarial Approaches

    Kyuil Lee, Dezhi Yu, Yongkang Huang

    cs.SD · cs.LG

    We study generative modeling of Bach-style symbolic piano music using a shared MIDI corpus and three model families: autoregressive LSTMs with attention, latent-variable models including recurrent VAEs and vector-quantized VAEs, and generative adversarial networks. We compare their ability to model polyphonic note sequences, learn useful latent representations, and generate stylistically coherent compositions. Our experiments show that the...

    arxiv.org/abs/2606.13626 · PDF

  2. 02

    Towards Personalized Federated Learning for Dysarthric Speech Recognition

    Tao Zhong, Mengzhe Geng, Jiajun Deng, Shujie Hu, Xunying Liu

    cs.SD · cs.AI

    Speech recognition is challenging for dysarthric speakers. While federated learning (FL)-based ASR can be an effective tool for protecting privacy, it suffers from heterogeneity issues caused by speaker variability. Forcing all speakers to share the same model components can be suboptimal under such heterogeneity, making personalization a promising direction; however, related research on dysarthric speech remains limited. To this end, this...

    arxiv.org/abs/2606.13253 · PDF

  3. 03

    Self-Guidance: Enhancing Neural Codecs via Decoder Manifold Alignment

    Xiang Li, Yixuan Zhou, Jingran Xie, Zhiyong Wu, Hui Wang

    cs.SD · cs.LG

    Neural speech codecs based on Vector-Quantized VAEs (VQ-VAEs) are core audio tokenizers for speech LLMs, yet their reconstruction fidelity is bottlenecked by quantization error. Modifying the quantizer or increasing model capacity are common fixes, but they complicate downstream language modeling. Our core idea is to align the decoder's internal feature manifolds when processing both the quantized tokens and their original continuous...

    arxiv.org/abs/2606.12940 · PDF

This edition is part of The Daily Abstract — cs.SD archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.