cs.SD · 2026-08-05 · No. 75

Sound, 2026-08-05.

3 new papers in cs.SD. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

3 entries
  1. 01

    Equivariant Music Transformer

    Zixun Guo, Simon Dixon

    cs.SD · cs.AI

    Humans recognize a musical passage even when it is shifted in time or transposed in pitch, indicating a notion of equivariance in the representation space. Our analysis, however, shows that standard music transformers map such time-shifted or pitch-transposed inputs onto uncorrelated representations: these models become progressively less equivariant as they scale in size or train longer. This suggests that in standard music transformers,...

    arxiv.org/abs/2608.03920 · PDF

  2. 02

    AI-Based Sound Effect Generation: A Narrative Review of Generative Models Across Input Modalities

    Sandy Abdo, Bill Kapralos, Priyamvada Tripathi, KC Collins, Adam Dubrowski

    cs.SD · cs.AI

    Sound effects play a crucial role in conveying actions, events, and environmental cues across digital applications, often requiring a high degree of variation and contextual adaptability. Artificial intelligence (AI)-driven audio generative models are rapidly growing in popularity and have the potential to transform the way sound is synthesized and used across various applications. In response to this growing momentum, this chapter reviews...

    arxiv.org/abs/2608.03742 · PDF

  3. 03

    Multi-Task Multi-Frame Visual Piano Transcription

    Yonghyun Kim, Hoyeol Sohn, Juhan Nam, Alexander Lerch

    cs.SD · cs.AI · cs.CV · cs.MM · eess.IV

    Audio-based piano transcription performs well on onset, pitch, and velocity, but the sustain pedal lets sound persist long after key release, so audio systems predict pedal-extended offsets rather than physical key release. Yet existing Visual Piano Transcription (VPT) systems focus on onset detection from short video windows, offset accuracy lags onset by a wide margin, and note-level velocity has not been reported. To address these gaps, we...

    arxiv.org/abs/2608.03419 · PDF

This edition is part of The Daily Abstract — cs.SD archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.