cs.SD · 2026-07-13 · No. 52

Sound, 2026-07-13.

3 new papers in cs.SD. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

3 entries
  1. 01

    ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models

    Sang-Hoon Lee, Ha-Yeong Choi

    cs.SD · cs.AI · eess.AS · eess.SP

    Representation alignment (REPA) has been investigated to accelerate diffusion training, but we observe that regularizing intermediate representations in diffusion Transformers (DiT) may implicitly entangle latents and limit generative capacity. To address this issue, we propose ReGen, a hierarchical multi-prompt representation generation framework that jointly estimates multiple vector fields for both representations and data within a single...

    arxiv.org/abs/2607.09134 · PDF

  2. 02

    Clean2FX: Label-conditioned modeling for clean-to-effect guitar audio transformations

    Oliverio Bombicci Pontelli, Iran R. Roman

    cs.SD · cs.LG

    We present Clean2FX, a study and demo of label-conditioned clean-to-effect transformation for electric guitar audio. Given a clean guitar input and a target effect label, the task is to synthesize the corresponding effected signal while preserving the musical content. Training and evaluation pairs are constructed from EGFxSet real, single tone recordings by assembling matched clean/effected chords, melodies, and mixed timelines. This allows...

    arxiv.org/abs/2607.08863 · PDF

  3. 03

    MulTTiPop: A Multitrack Transcription Dataset for Pop Music

    Nathan Pruyne, Benjamin Stoler, William Chen, Chien-yu Huang, Shinji Watanabe, Chris Donahue

    cs.SD · cs.LG

    We present MulTTiPop, a dataset of pop music segments and their associated multitrack MIDI recordings for the evaluation of automatic music transcription models. MulTTiPop contains 572 segments of popular music totaling 3.5 hours of audio, and contains songs from diverse genres and decades from the 1930s to 2000s. To collect this dataset, we perform metadata-based matching on song segments from the Lakh MIDI and TheoryTab datasets, manually...

    arxiv.org/abs/2607.08756 · PDF

This edition is part of The Daily Abstract — cs.SD archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.