cs.SD · 2026-08-03 · No. 73

Sound, 2026-08-03.

3 new papers in cs.SD. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

3 entries
  1. 01

    DoubleHelix: Structured Cross-Modal Fusion for Audio-Visual Speech Recognition with LLMs

    Ziwei Cheng, Zhenhua Tan, Zhuomin Zhu

    cs.SD · cs.AI

    Audio-visual speech recognition (AVSR) relies on effective fusion of audio and visual modalities, yet existing approaches treat cross-modal interaction as a single-step operation without structured iterative refinement. We present DoubleHelix, a multimodal fusion framework that reformulates fusion as an iterative cross-modal interaction process with adaptive degradation-aware enhancement. The framework comprises three components including...

    arxiv.org/abs/2607.29112 · PDF

  2. 02

    TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models

    Aryan Vijay Bhosale, Harshit Rajgarhia, Abhishek Mukherji, Dinesh Manocha

    cs.SD · cs.AI · cs.CL

    Unified audio models capable of audio understanding, audio generation and, increasingly, audio editing are proliferating rapidly. Yet a basic question about them remains unanswered: do the two heads of a unified model agree about the same audio? Current practice evaluates each capability in isolation on specialized benchmarks, and never asks whether a model can make sense of its own generations. We present TORUS, the first self-coherence test...

    arxiv.org/abs/2607.28896 · PDF

  3. 03

    Learning to Predict Performance-induced Emotion Differences in Classical Piano Music

    Joann Ching, Gerhard Widmer

    cs.SD · cs.LG · cs.MM

    Music is often used as a medium for communicating emotion, with performers shaping perceived affect through interpretation. This study addresses the challenge of identifying and predicting subtle changes in perceived emotion that are exclusively due to differences in performance. We focus on classical solo piano music, using a set of 6 commercial recordings of Bach's Well-Tempered Clavier Book I, annotated in terms of valence and arousal. By...

    arxiv.org/abs/2607.28876 · PDF

This edition is part of The Daily Abstract — cs.SD archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.