cs.CV · 2026-09-21 · No. 120

Computer Vision and Pattern Recognition, 2026-09-21.

13 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

13 entries
  1. 01

    Benchmarking the Explanatory Quality of Open-Weight Vision-Language Models in Face Recognition

    Laurent Colbois, Sébastien Marcel

    cs.CV · cs.AI

    Vision-Language Models (VLMs) have recently been proposed as promising tools for face recognition, as they can produce natural language explanations alongside similarity scores. This capability is considered appealing for face comparisons in forensic contexts, which require decisions to be transparent and auditable. However, existing evaluations of VLMs for that use case focus mostly on recognition accuracy, while the validity of generated...

    arxiv.org/abs/2609.21879 · PDF

  2. 02

    Chronosphere: Space-Time Tessellation of Local Climate Experts

    Daniel Cher, Eric Xing, Kexing Li, Brian Wei, Isaac Corley, Nathan Jacobs

    cs.CV · cs.LG

    We introduce Chronosphere, a spatio-temporal neural field that learns representations of climate. A central challenge in geographic representation learning is modeling environmental processes whose spatial and temporal complexity varies widely. Yet existing location encoders typically fix a single level of detail everywhere. Global bases such as spherical harmonics spread capacity uniformly across space and time. Localized bases resolve only...

    arxiv.org/abs/2609.21872 · PDF

  3. 03

    Beyond Benchmark Scores: Auditing Medical Vision-Language Models for Chest X-Ray Tuberculosis Screening

    Mushir Akhtar, M. Tanveer, Mohd. Arshad

    cs.CV · cs.LG

    A medical model's benchmark score does not establish that the same conclusion holds under a different evaluation. This study tests whether claims about model ranking, score reliability and screening performance survive changes in cohort, prompt, negative spectrum, specified prevalence and operating threshold. We audit three medical vision-language models (BioMedCLIP, CheXficient, and MedSigLIP) and a general-domain OpenCLIP comparator on...

    arxiv.org/abs/2609.21763 · PDF

  4. 04

    Balanced Prompt Adaptation against Entropy-Induced Collapse for Test-Time Binary Segmentation

    Zhengshan Wang, Joshua Charles Webster-Ford, Yifei Tian, Xinxin Wang, Long Chen, Weiping Ding

    cs.CV · cs.AI

    Entropy minimization is a standard objective for test-time adaptation (TTA), but it can fail in imbalanced binary segmentation. Unlike image classification, dense segmentation aggregates thousands of pixel predictions, allowing the larger predicted class to dominate the update, pull minority predictions toward itself, and produce a degenerate mask as predictions saturate and their entropy gradients vanish. We theoretically establish this...

    arxiv.org/abs/2609.21743 · PDF

  5. 05

    Configurable Multi-Stage Vision Pipeline for Crop Disease and Pest Diagnosis

    Naga Ganesh, Chandrashekar M S, Lakshmi Pedapudi, Aakash Singh, Vineet Singh

    cs.CV · cs.CL · cs.LG

    Farmer.Chat is Digital Green's farm advisory service for smallholder farmers. When something looks wrong with a crop, the farmer takes a photograph and sends it, and that photograph is the whole question: no symptom described, no crop named, often no text at all. The service has to determine whether the picture can be used, what crop it shows, and what is wrong with it, from images taken on cheap phones in a field, in poor light and with a...

    arxiv.org/abs/2609.21651 · PDF

  6. 06

    Detection is solved, delineation is not: what governs tooth segmentation on panoramic radiographs

    Muhammad Rehan, Moaz Amjad, Syed Danial Ahmed, Mariam Adnan, Haider Ali

    cs.CV · cs.LG

    Automatic tooth segmentation and FDI numbering on panoramic radiographs underpins computer-assisted dental diagnosis, yet which factors govern performance remains unclear. We assemble a corpus of 1,422 panoramic radiographs containing 42,142 expert-delineated tooth polygons across the 32-class FDI taxonomy, annotated by 30 dental practitioners and independently reviewed by two others, and use it to isolate input resolution, architecture and...

    arxiv.org/abs/2609.21628 · PDF

  7. 07

    Purification and Regulation: Comorbidity-Aware Multi-Label Few-Shot Learning for Medical Image Classification

    Ying-Chih Lin, Po-Chih Kuo, Yong-Sheng Chen

    cs.CV · cs.LG

    Multi-label few-shot learning (MLFSL) remains a significant challenge in medical image analysis (MIA). Current metric-based meta-learning methods face two critical limitations in MIA. First, conventional prototype generation often entangles irrelevant disease information, leading to contaminated prototypes and degraded performance. Second, prior studies typically enforce inter-class separability in embedding space, largely neglecting the...

    arxiv.org/abs/2609.21541 · PDF

  8. 08

    VidOmni-Bench: A Benchmark for Fine-Grained Video Understanding via Spatio-Temporal Event Verification across Complexity and Duration

    Changbeen Kim, Junwon Chang, Kipyo Kim, Risa Shinoda, Kuniaki Saito, Donghyun Kim

    cs.CV · cs.AI

    While Video Large Language Models (Video-LLMs) have recently demonstrated strong performance, reliably evaluating their fine-grained video understanding remains challenging. Existing benchmarks often rely on question answering or ground-truth caption matching, where models may succeed through superficial cues and incomplete annotations. To this end, we introduce VidOmni-Bench, a benchmark that requires models to verify whether each event in...

    arxiv.org/abs/2609.21521 · PDF

  9. 09

    Think Locally, Refine Globally for Memory-Efficient 3D Reconstruction

    Jingke Zhou, Chenhang Ma, Zhizhou Zhong, Mingkai Liu, Zhuang Zhou, Yicheng ji, Binghua Su, Bo Cai, Xianliang Huang

    cs.CV · cs.AI

    We propose LoG-VGGT, a memory-efficient framework for long-sequence 3D reconstruction that balances local temporal modeling with global camera consistency. Instead of relying on full global attention, our method introduces cross-window attention at a small subset of transformer blocks, enabling effective information propagation across adjacent temporal windows while keeping memory usage bounded. To mitigate long-term pose drift, we further...

    arxiv.org/abs/2609.21437 · PDF

  10. 10

    AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents

    Seoyeon An, Hyeonseo Jang, Minsu Kim, Chanho Lee, Younghan Park, Kangwook Lee

    cs.CV · cs.AI

    Comprehensive video understanding is crucial for advancing artificial intelligence toward the intricate dynamics of the physical world. While recent advances in Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in video understanding, existing benchmarks remain confined to simple scene-level queries or global summaries that require only single-step inference. Real-world video understanding involves more...

    arxiv.org/abs/2609.21386 · PDF

  11. 11

    Hiding in Plain Sight: A Diffusion-based Mitigation of Geolocation Privacy Leakage in Vision-Language Models

    Yining Wang, Xi Li, Mi Zhang, Xiaohan Zhang, Xiaoyu You, Zhenxing Qian, Mi Wen

    cs.CV · cs.LG

    Multimodal large reasoning models (MLRMs) have demonstrated remarkable capabilities in complex visual understanding. However, this very power introduces a critical yet underexplored privacy threat: adversaries can exploit MLRMs to precisely infer users' geographic locations from casually shared photographs, by performing structured reasoning over subtle visual cues such as architectural styles, vegetation, and lighting conditions. In this...

    arxiv.org/abs/2609.21363 · PDF

  12. 12

    Beyond Exact Match: Task-Aware GRPO for Cross-Domain PCBA Visual Question Answering

    Jia Li, Li Dai, Peng Jia, Zhenzhen Hu, Chee Seng Chan, Bingkun Bao, Richang Hong

    cs.CV · cs.AI

    In automated Printed Circuit Board Assembly (PCBA) inspection, standards-guided decisions require systems to jointly reason over fine-grained visual cues, component semantics, and manufacturing knowledge. Although large vision-language models (VLMs) provide a promising foundation, their deployment is hindered by the domain shift between standards-derived samples and real-world production-line imagery, together with heterogeneous output spaces...

    arxiv.org/abs/2609.21276 · PDF

  13. 13

    Multiclass Semantic Segmentation of Wildland Fire Images Using Context-Aware Centralized Copy-Paste Data Augmentation

    Joon Tai Kim, Nishanth Kunchala, Vishv Patel, Tianle Chen, Ziyu Dong, Daniel Ospina Acero, Roger Williams, Mrinal Kumar

    cs.CV · cs.LG

    Producing accurate annotations for deep learning based image segmentation is both costly and labor intensive. This challenge is especially evident in wildland fire applications, where accurately labeled datasets are scarce due to the difficulty of collecting and annotating dynamic fire scenes. To address this problem, our previous work introduced the Centralized Copy-Paste Data Augmentation (CCPDA) method for semantic segmentation of wildland...

    arxiv.org/abs/2609.21241 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.