cs.SD · 2026-06-10 · No. 19

Sound, 2026-06-10.

3 new papers in cs.SD. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

3 entries
  1. 01

    What Do Deepfake Speech Detectors Actually Hear?

    Vojtěch Staněk, Veronika Jirmusová, Anton Firc, Kamil Malinka, Jakub Reš, Martin Perešíni

    cs.SD · cs.AI · cs.CR · cs.LG

    Deepfake speech detectors often output a single score without explaining why an audio sample is flagged, where in the signal the evidence lies, or what cues drive the decision. We propose an audio-native explainability pipeline using Integrated Gradients on time-aligned self-supervised representations to localize decision evidence over time. We apply the proposed method to three WavLM-based detectors (AASIST, CA-MHFA, SLS) on ASVspoof 5 and...

    arxiv.org/abs/2606.10912 · PDF

  2. 02

    Ethical and Technical Limits of Deepfake Speech Datasets

    Vojtěch Staněk, Eva Trnovská, Kamil Malinka, Anton Firc

    cs.SD · cs.AI · cs.CR · cs.LG

    Claims about the robustness and fairness of deepfake speech detectors are only as credible as the datasets used to train and evaluate those systems. We present a dataset-level audit of the deepfake speech landscape. We compile and analyze 39 deepfake speech datasets, examining key attributes including accessibility, documentation, demographic and language coverage, dataset scale, and the underlying bona fide speech sources. Our audit reveals...

    arxiv.org/abs/2606.10911 · PDF

  3. 03

    RAT: Reference-Augmented Training for ASV Anti-Spoofing

    Vojtěch Staněk, Anton Firc, Jakub Reš, Kamil Malinka

    cs.SD · cs.AI · cs.CR · cs.LG

    We introduce a spoofing countermeasure architecture conditioned on speaker-reference recordings, but observe that it converges to a solution that effectively ignores the reference during inference. Surprisingly, training with a reference channel induces invariance that improves deepfake detection, even when the reference is absent or mismatched during inference. Based on this observation, we propose a Reference-Augmented Training (RAT)...

    arxiv.org/abs/2606.10908 · PDF

This edition is part of The Daily Abstract — cs.SD archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.