cs.CV · 2026-09-17 · No. 116

Computer Vision and Pattern Recognition, 2026-09-17.

13 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

13 entries
  1. 01

    Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection

    Girish A. Koushik, Diptesh Kanojia, Helen Treharne

    cs.CV · cs.AI · cs.CL · cs.LG

    When a large vision-language model misclassifies a harmful meme, the failure may reflect missing internal evidence or an inability to route represented evidence to its output. We distinguish these cases in Gemma-3 and Qwen3.5 using sparse autoencoders, role-conditioned probes, causal interventions, and recovery experiments across six harmful content benchmarks, with additional Spanish and Hindi-English code-mixed evaluations. Sparse readouts...

    arxiv.org/abs/2609.18860 · PDF

  2. 02

    Using OCR Heads to Verbalize Image Semantics

    Sheridan Feucht, Benno Krojer, Sarah Wang, Henry Abrahamsen, Byron C. Wallace, David Bau

    cs.CV · cs.AI · cs.CL

    How do VLMs map from pixels to semantics? To understand this general question, we focus on a narrow one: studying how VLMs perform optical character recognition (OCR). Across four models, we identify attention heads causally necessary for OCR, and discover that these are in fact general-purpose heads that output interpretable semantic features across all image tokens. For example, pointing these heads at an image token containing the word...

    arxiv.org/abs/2609.18823 · PDF

  3. 03

    Generalist-Specialist Mixture-of-Experts for Rare Pathology Detection in Multimodal Imaging

    Johannes Kaiser, Florian Braunmiller, Daniel Rückert, Georgios Kaissis

    cs.CV · cs.AI

    AI models for multimodal medical imaging must balance modality-specific specialization with cross-modal shared representations, a trade-off that pure Mixture-of-Experts (MoE) architectures currently fail to satisfy. Expert-based routing improves in-domain learning but may sacrifice cross-modal signals, which appear particularly important for rare (low-prevalence) pathologies in our experiments. To resolve this, we introduce...

    arxiv.org/abs/2609.18688 · PDF

  4. 04

    GenStream: Semantic Streaming Framework for Generative Reconstruction of Human-centric Media

    Emanuele Artioli, Daniele Lorenzi, Shivi Vats, Farzad Tashtarian, Christian Timmerer

    cs.CV · cs.AI · cs.MM

    Video streaming dominates global internet traffic, yet conventional pipelines remain inefficient for structured, human-centric content such as sports, performance, or interactive media. Standard codecs re-encode entire frames, foreground and background alike, treating all pixels uniformly and ignoring the semantic structure of the scene. This leads to significant bandwidth waste, particularly in scenarios where backgrounds are static and...

    arxiv.org/abs/2609.18634 · PDF

  5. 05

    On-the-Fly Homographies Calibration for Multi-Camera Tracking

    David Voihanski, Mor Sinai, Ben Zion Bobrovsky

    cs.CV · cs.AI

    Precise multi-camera tracking traditionally relies on rigorous 3D site calibration, yet this requirement is often operationally impossible in large-scale deployments. Privacy regulations frequently prohibit recording video for offline calibration; limited bandwidth precludes synchronizing high-resolution streams from hundreds of cameras; and covering immense physical sites with calibration targets is logistically infeasible. We present a...

    arxiv.org/abs/2609.18582 · PDF

  6. 06

    Learning from Distributed Eyes: Leveraging Collaborative Perception for Automated Model Adaptation

    Yanan Ma, Yihang Tao, Zhengru Fang, Zihan Fang, Yiqin Deng, Xianhao Chen, Yuguang Fang

    cs.CV · cs.LG

    In autonomous driving, perception models often struggle to generalize to new environments due to domain shifts. While unsupervised model adaptation offers a feasible solution without labor-intensive manual labeling, existing methods that rely solely on the ego-vehicle's data often lead to inferior pseudo-labeling performance. To address this critical issue, we propose LDE, Learning from Distributed ``Eyes", a novel framework that transforms...

    arxiv.org/abs/2609.18511 · PDF

  7. 07

    Beyond Random Couplings: Contrastive Noise Alignment in Generative Flows

    Lennart Wittke, Vinicius Azevedo

    cs.CV · cs.LG

    Diffusion and flow-matching models are typically trained by corrupting data through independently sampled Gaussian noise. While simple and scalable, this forward process induces arbitrary data-noise couplings, forcing the network to learn high-curvature transports between unrelated endpoints. Existing optimal-transport methods reduce this burden by reassigning fixed noise samples to data, but the source noise distribution itself remains...

    arxiv.org/abs/2609.18488 · PDF

  8. 08

    CSWAM: Better Causal Semantic Representations for Out-of-Distribution Generalization in World Action Models

    Tianbin Liu, Jian Zhu, Taiyi Su, Jianjun Zhang, Chong Ma, Zitai Huang, Yi Xu

    cs.CV · cs.AI

    FastWAM-style world action models enable efficient action-only inference, but generalize poorly under visual distribution shifts. Their reconstruction-oriented representations emphasize appearance-specific details, limiting generalization to unseen scenes and objects. Without observation history, the model also lacks temporal evidence for robustly identifying task-relevant state changes and motion in unfamiliar visual conditions. To address...

    arxiv.org/abs/2609.18462 · PDF

  9. 09

    A Non-Linear Neuron Based Detection of Isolated Pixels in Binary and Grayscale Images using Contrast Sensitive Receptive Fields

    Nassir Mohammad

    cs.CV · cs.AI

    Identifying isolated points is important in image processing applications such as medical imaging, astronomy and quality control management. Other domains, such as cybersecurity, also present challenges that can be framed as image processing problems. One example of particular interest is the identification of anomalous single nodes in spatially organised networks where groups of nodes in different regions share similar feature values. This...

    arxiv.org/abs/2609.18399 · PDF

  10. 10

    CapMap-MS-TTA: 3rd Place Solution for the MUMU Track of the 8th LSVOS Challenge at ECCV 2026

    Chengfeng Qiu, Kaifeng Wei

    cs.CV · cs.AI

    The MUMU track of the 8th Large-scale Video Object Segmentation (LSVOS) Challenge requires a single unified multimodal model to jointly solve image tagging (Task A), open-vocabulary object detection (Task B), and English captioning (Task C) under strict resource constraints (<=0.5B parameters and <=8 GB peak GPU memory). We present CapMap-MS-TTA, a training-free submission built on Microsoft Florence-2-base (~231M parameters), combining...

    arxiv.org/abs/2609.18206 · PDF

  11. 11

    MCLC-NET: Multimodal Continual Learning for Leaf Counting

    Ruchi Bhatt, Pratibha Kumari, Shreya Bansal, Vedant Agnihotri, Dwarikanath Mahapatra, Mukesh Saini

    cs.CV · cs.LG

    Leaf counting is an important task in plant phenotyping for monitoring plant growth and estimating crop yield. Most existing methods rely on RGB images, but their performance is often affected by occlusion, lighting variations, and other real-world challenges. Additional modalities, such as depth and thermal images, can provide useful complementary information. However, multimodal leaf counting remains underexplored. Also, many existing...

    arxiv.org/abs/2609.18129 · PDF

  12. 12

    Mask 2D-3D: Adaptive Dual-Masked Autoencoder Network for Image-to-Point Cloud Registration

    Zhixin Cheng, Jiacheng Deng, Xiaotian Yin, Baoqun Yin, Richang Hong, Tianzhu Zhang

    cs.CV · cs.AI

    Detection-free methods for image-to-point cloud registration are prone to erroneous correspondences caused by domain and modality discrepancies, limited sensitivity of feature extractors, and the presence of non-overlapping regions. The Masked Autoencoder (MAE) has shown strong performance in visual representation for images and point clouds. It may be helpful to apply this approach to image-to-point cloud registration, a task that requires...

    arxiv.org/abs/2609.18088 · PDF

  13. 13

    vidax: A Unified JAX Framework for Video Generative Models on Accelerator Meshes

    Congyue Deng

    cs.CV · cs.DC · cs.LG

    Open-source video generative models ship almost exclusively as PyTorch/CUDA reference implementations. This leaves Cloud TPU pods without a production-ready inference path, despite offering large, cost-effective accelerator memory pools ideal for long-sequence spatiotemporal attention. We present vidax, an open-source JAX/Flax inference engine and zero-copy PyTorch-to-JAX weight translator for modern video generation architectures. vidax...

    arxiv.org/abs/2609.18077 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.