cs.SD · 2026-09-12 · No. 113

Sound, 2026-09-12.

3 new papers in cs.SD. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

3 entries
  1. 01

    Continuous-Time Acoustic Modelling with Neural Controlled Differential Equations

    Mattias Cross, Minghui Zhao, Anton Ragni

    cs.SD · cs.AI

    Text-to-speech (TTS) models commonly address text--speech alignment by expanding phone-level encoder states to frame-level decoder inputs using predicted durations. While this length-regulation step resolves alignment structurally, this use of duration typically changes only where and how often latent states appear, not the values of the states themselves. This paper proposes a continuous-time mechanism for duration-aware acoustic modelling...

    arxiv.org/abs/2609.11725 · PDF

  2. 02

    ZipCodec: Ultra-Low-Frame-Rate Streaming Speech Coding

    Luca Della Libera, Cem Subakan, Mirco Ravanelli

    cs.SD · cs.AI · cs.LG

    Neural audio codecs are a fundamental component of modern speech generation systems. While recent codecs achieve increasingly low bitrates, reducing frame rate remains challenging, as each token must preserve more information while maintaining reconstruction quality. We present ZipCodec, a streaming neural speech codec operating at 6.25 Hz and 0.80 kbps with a theoretical latency of 160 ms. Our approach combines large-scale WavLM distillation...

    arxiv.org/abs/2609.11642 · PDF

  3. 03

    X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation

    Haojun Zhang, Yi Zou, Min Chen, Qize Yu, Lianrui Fan, Xini Ding, Hao Li, Shuchang Zhou, Xianming Liu, Shiyu Huang

    cs.SD · cs.AI

    Reducing audio-encoder depth lowers the inference cost of speech large language models, but removing complete blocks perturbs the embeddings consumed by the decoder and can cause deletion and premature end-of-sequence errors. We introduce X-AuT, a progressive framework that selects layer combinations through short behavioral probes and restores the pruned model through representation alignment, cross-scale distillation, scheduled...

    arxiv.org/abs/2609.11412 · PDF

This edition is part of The Daily Abstract — cs.SD archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.