cs.CV · 2026-09-15 · No. 114
Computer Vision and Pattern Recognition, 2026-09-15.
16 new papers in cs.CV. Titles, authors,
abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →
01 — The papers
16 entries-
01
Anatomical Grounding and Leakage-Aware Multimodal Contrastive Learning for Alzheimer's Disease Classification from Structural MRI
Paul-Gabriel Nicolae, Irina Georgiana Mocanu
cs.CV · cs.AI
Deep networks trained on structural MRI for Alzheimer's disease (AD) staging often reach reasonable accuracy while attending to anatomically irrelevant regions, and multimodal models that add clinical tables frequently rely on variables that were used to assign the diagnostic label in the first place. We study both issues with a deliberately lightweight slice-based encoder (ResNet18 with a one-layer Transformer over slices) on 1,075 baseline...
-
02
Don't Send What You Don't Need: Question-Guided Token Pruning as a Privacy Defense for Vision-Language Models
Md Khalid Syfullah, Alvi Ataur Khalil
cs.CV · cs.AI · cs.CR
Visual Question Answering (VQA) with Vision-Language Models (VLMs) is increasingly used in privacy-sensitive and bandwidth-constrained settings. Federated Learning (FL), Split Learning (SL), and U-Shaped Split Learning (USL) keep raw data local, but transmitting all visual tokens across a model partition remains costly and can expose private information. We propose QPriv-VL, a question-guided, privacy-aware token-pruning framework for FL, SL,...
-
03
Benchmarking Intra-Patient 3D Deformable Multimodal Image Registration
Matteo Barbieri, Giammarco La Barbera, Juan Pablo De La Plata, Sabine Sarnacki, Isabelle Bloch, Pietro Gori
cs.CV · cs.AI
Multimodal image registration is a key component of many clinical workflows, yet it remains challenging because corresponding anatomical structures often exhibit substantially different image intensities across modalities. In this work, we present a comprehensive benchmark of intra-patient 3D multimodal deformable registration methods across three datasets covering different anatomical regions and difficulty levels, including both synthetic...
-
04
Circuit-MLLM: Topological Logic-Guided Latent-Space Visual Reasoning for Circuit Schematic Understanding
Jinyuan Deng, Yuqi Jiang, Wenjing Huang, Xin Li, Qi Sun, Cheng Zhuo
cs.CV · cs.AI · cs.LG
Through pre-training on extensive text and image datasets, current multi-modal large language models (MLLMs) achieve strong performance on general tasks. However, circuit schematics present a unique challenge for MLLMs due to their dense component layouts and distinct topological logic, demanding fine-grained structural parsing to extract the electrical semantics. To address this, we propose Circuit-MLLM, a multimodal reasoning framework that...
-
05
ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMs
Sebastián Andrés Cajas Ordóñez, Maximin Lange, Quang Bui, Anqi Peter Li, Felipe Ocampo Osorio, Rafi Al Attrach,...
cs.CV · cs.AI
A radiology report can already answer a clinical question, so it is hard to tell whether a vision-language model also uses the image. ModaLens, a paired image-swap audit, measures how report availability changes image sensitivity: MedGemma-27B on 3,199 paired MIMIC-CXR cases from 293 patients, all 14 questions per case (13 finding-specific and one composite), each image replaced by one from another study, usually of the same patient, with...
-
06
VideoScout: Learning Agentic Active Exploration with Adaptive Reasoning Pacing for Long Video Understanding
Weixin Xu, Zhenyu Yang, Bing Wang, Shengsheng Qian, Changsheng Xu
cs.CV · cs.AI
Multimodal Large Language Models (MLLMs) have achieved remarkable progress on short video understanding yet remain limited on long videos due to the limited visual context window. Prevailing approaches rely on uniform frame sampling or recent coarse-to-fine agentic zooming, both of which struggle to localize sparse, decisive evidence in sufficiently long videos. We formulate long video understanding as a \textbf{Sequential Evidence...
-
07
A Unified Vision-Language Model for PSMA PET/CT Report Generation, Visual Question Answering, and Lesion Segmentation
Yang Xing, Jiong Wu, Savas Ozdemir, Yang Zhou, Boxiao Yu, Ying Zhang, Zheren Zhu, Chenyu You, Wei Shao, Yang Lu,...
cs.CV · cs.AI
Accurate PSMA PET/CT interpretation is central to prostate cancer management, yet existing PET/CT AI models typically address isolated tasks. We propose a unified PSMA PET/CT vision-language model for report generation, visual question answering, and lesion segmentation. The framework adopts an LLaVA-style architecture, comprising a PET/CT vision encoder, an MLP-Mixer projection module, a LoRA-tuned large language model, and a 3D segmentation...
-
08
PIVOT: Physics-Grounded Verification for AI-Generated Audio-Video Detection
Bo Zheng, Kangran Zhao, Xiaoyu Zhang, Weinan Guan, Zhiheng Li, Yize Chen, Haizhou Li, Qingshan Liu, Siwei Lyu, Baoyuan Wu
cs.CV · cs.AI · cs.MM
As generative models continue to advance, AI-generated content (AIGC) is becoming increasingly realistic, weakening the artifact cues commonly exploited by existing detectors. Nevertheless, faithfully reproducing the physical behavior of real-world events remains challenging for current generators. We therefore explore detecting AIGC by assessing whether the depicted event satisfies measurable constraints derived from physical laws. We...
-
09
A Conservative OCR-Enabled Workflow for R214 Sodium Screening of South African Packaged Foods
Mayimunah Nagayi, Alice Scaria Khan, Tamryn Frank, Rina Swart, Clement Nyirenda
cs.CV · cs.AI
Using food package images to monitor sodium and salt content against South Africa's R214 sodium limits is challenging when screening decisions require product identity, nutrition facts panel evidence, reporting basis, and category-specific thresholds. This study presents a conservative image-based workflow that combines region detection, optical character recognition (OCR), product identity and sodium evidence extraction, R214 category...
-
10
End-to-End Cell Detection via Instance-aware Graph Modeling
Ruochen Liu, Yalin Zheng, Jingxin Liu, Jianfeng Zhang, Shoujun Huang, Dexing Kong, Haofeng Li, Wei Lou
cs.CV · cs.AI
Accurate cell detection and classification are crucial for pathological analysis, directly affecting diagnostic accuracy and treatment planning. To capture complex cellular interactions beyond visual appearance within the tumor microenvironment, several approaches have employed graph neural networks to model spatial and relational patterns among cell nuclei, yielding promising results. However, these methods typically adopt a two-stage...
-
11
Concept-Grounded Reasoning with Prompt-Driven Localization for Interpretable Structured Report Generation
Xinyue Xu, Hongbin Lin, Juangui Xu, Hualiang Wang, Lehan Wang, Lijie Hu, Weiyang Liu, Adrian Weller, Xiaomeng Li
cs.CV · cs.AI · cs.LG
Medical imaging modalities such as ultrasound and X-ray are widely used in clinical practice, where diagnosis follows a structured, evidence-driven workflow aligned with standardized criteria. While multimodal large language models (MLLMs) show promise for automated medical report generation, most existing systems rely on end-to-end multimodal fusion without modeling clinically defined intermediate attributes, leading to limited grounding and...
-
12
Pre-PEFT Probing: Weight Statistics and Perturbation Robustness for Layer Selection in VLM Vision Encoders
Qingtao Xia, Jiahua Bao, Siyao Cheng, Jie Liu
cs.CV · cs.AI
We propose a pre-fine-tuning probing method for Parameter-Efficient Fine-Tuning (PEFT) layer selection, aiming to obtain more stable and higher gains with fewer trainable parameters when adapting large vision--language models (VLMs). Unlike the common practice of applying LoRA and other adapters to all layers at once---where layer selection often relies on heuristic rules---we focus on the vision encoder and directly evaluate the...
-
13
PACE: Progressive Angular-to-Norm Contrastive Embedding
Yanping Li, Wei Zhou, Yawen Liu, Yibo Wang, Ke Zhu, Guangda Huzhang, Qing-Guo Chen, Zhao Xu, Jun Zhang, Wei Wei
cs.CV · cs.AI · cs.CL · cs.LG
Multimodal embedding models encode heterogeneous inputs into a shared embedding space, enabling efficient similarity computation across modalities and tasks. Most existing methods optimize cosine-based contrastive objectives, which promote stable training but restrict semantic compatibility to angular geometry, precluding embedding norms from serving as an additional semantic signal. However, directly optimizing the more expressive...
-
14
SparseTalk - Sparsifying 3D Gaussian Language Fields for Efficient 3D Visual Question Answering
Davit Soselia, Joseph JaJa, Amitabh Varshney
cs.CV · cs.LG
3D Gaussian language fields provide an explicit, spatially grounded representation for 3D visual question answering (VQA), but their dense semantic features can require tens of thousands of embeddings per scene, resulting in substantial storage, memory, and inference costs. We investigate how much of this representation is actually necessary for downstream reasoning. Starting from a full embedding representation, we systematically sparsify...
-
15
AdaVSkip: Adaptive Visual Token Skipping Across Layers For Efficient MLLMs Inference
Yuyao Sun, Tao Deng, Shuang Li, Deqing Wang
cs.CV · cs.AI
Multimodal large language models (MLLMs) require substantial computation to process numerous visual tokens across all transformer layers. Most methods for efficient MLLM inference exploit horizontal redundancy by compressing visual tokens. Beyond token reduction, recent studies exploit vertical redundancy through early exit or fixed-layer skipping. However, we find that the extent and distribution of this redundancy vary across inputs and...
-
16
ChatGPT Images 2.5 in the Wild: A Launch-Period Dataset and Detector Evaluation
Dennis Ng, Xingyu Shen, Ankit Raj, Kidus Zewde, Tommy Duong, Yuchen Zhou, Yuxin Zhang, Neo Tiangratanakul, Simiao Ren
cs.CV · cs.AI
An image tool can change its underlying generator while retaining its public name, making version attribution from online posts ambiguous. We study this problem after the ChatGPT Images 2.5 launch. Our frozen collection contains 3,478 images from 2,440 posts across 8 sources. Recorded posting times fall within the first 51.1 hours after the announcement. It records three attribution tiers and retains standalone images after image-form...
This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.
#D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.