eess.AS · 2026-09-21 · No. 120
Audio and Speech Processing, 2026-09-21.
2 new papers in eess.AS. Titles, authors,
abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →
01 — The papers
2 entries-
01
Samsone: A Family of Open Small Audio Language Models for On-Device Inference
Piotr Masztalski, Michał K. Grzeszczyk, Olaf Sikorski
eess.AS · cs.AI
The success of Large Audio Language Models has driven the development of massive multimodal networks exceeding billions of parameters. However, the demand for privacy-preserving, low-latency processing has shifted focus toward Small Audio Language Models (SALMs) capable of on-device execution. In this paper, we introduce Samsone, a family of SALMs designed for edge computing. Our core model, Samsone-134M, establishes a new state-of-the-art...
-
02
OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue
Haolin He, Yunfei Chu, Qi Chen, Wen Huang, Yuan Feng, Muzhi Zhu, Zheqi Dai, Haoning Xu, Dongchao Yang, Chunyat Wu,...
eess.AS · cs.AI · eess.IV
We define OmniVChat (Omni Video Chat) as the task of native audio-visual dialogue between a user and an omni model. In OmniVChat, omni models directly and simultaneously receive audio and video from a user and return text. The user's query is embedded in the audio and video, without a separate text question, external captioning, or speech recognition. Direct audio-visual input reduces external latency and computation while preserving...
This edition is part of The Daily Abstract — eess.AS archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.
#D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.