cs.SD · 2026-09-03 · No. 104

Sound, 2026-09-03.

2 new papers in cs.SD. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

2 entries
  1. 01

    Scalable Direction-Following TTS via Voice Impression-Guided Pseudo Triplet Construction

    Kenichi Fujita, Yusuke Ijima

    cs.SD · cs.CL · cs.LG

    Voice actors often re-read the same script while modifying their delivery in response to performance directions. We study this setting as direction-following TTS, where a system generates a new utterance that reflects a given direction relative to a reference utterance while preserving speaker identity and linguistic content. A key challenge is the lack of training data capturing such relative modifications. To address this, we propose a...

    arxiv.org/abs/2609.02623 · PDF

  2. 02

    Auditory Illusion Benchmark for Large Audio Language Models

    Hayoon Kim, Eunice Hong, Kyogu Lee

    cs.SD · cs.AI

    Perceptual illusions have long served as crucial probes into human cognition, revealing biases and limitations of perception. In the auditory domain, such illusions provide a unique lens for testing whether Large Audio Language Models (LALMs) replicate human perceptual tendencies. Despite their importance, most benchmarks focus on visual illusions or general audio tasks, leaving auditory illusions underexplored. To this end, we present AIB,...

    arxiv.org/abs/2609.02277 · PDF

This edition is part of The Daily Abstract — cs.SD archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.