cs.MM · 2026-05-30 · No. 12
Multimedia, 2026-05-30.
1 new papers in cs.MM. Titles, authors,
abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →
01 — The papers
1 entries-
01
Unveiling the Visual Counting Bottleneck in Vision-Language Models
Xingzhou Pang, Yifan Hou, Junling Wang, Mrinmaya Sachan
cs.MM · cs.CV · cs.LG
While Large Vision-Language Models (VLMs) excel at interpolation, they suffer catastrophic failures in systematic generalization, most notably in visual counting. In this work, we investigate this extrapolation bottleneck by deconstructing visual counting into three cognitive stages: visual individuation, magnitude awareness, and symbolic mapping. Using synthetic Go boards and linear probes, we demonstrate that visual backbones maintain...
This edition is part of The Daily Abstract — cs.MM archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.
#D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.