cs.CV · 2026-07-21 · No. 60

Computer Vision and Pattern Recognition, 2026-07-21.

27 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

27 entries
  1. 01

    The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric

    Sheng-Yu Wang, Yotam Nitzan, Aaron Hertzmann, Jun-Yan Zhu, Eli Shechtman, Alexei A. Efros, Richard Zhang

    cs.CV · cs.LG

    Human visual similarity judgments are context-dependent. For example, two images may be similar in shape but distinct in color. Existing perceptual similarity metrics, however, collapse these nuances into a single scalar value, offering no mechanism to condition on specific aspects. To bridge this gap, we introduce a large-scale dataset of human similarity judgments over image triplets, where each triplet is annotated across multiple,...

    arxiv.org/abs/2607.18237 · PDF

  2. 02

    Simple Domain Generalization for Strong Pixel-Level Image Tampering Detection in Modern VLMs

    Yi Tang, Xinyi Shang, Jiacheng Cui, Sondos Mahmoud Bsharat, Jiacheng Liu, Xiaohan Zhao, Tran Dinh Tien, Ahmed...

    cs.CV · cs.AI

    Modern vision-language models (VLMs) have significantly improved image generation and editing capabilities, making pixel-level image tampering detection increasingly important yet challenging under cross-model and out-of-distribution shifts. This work studies domain generalization for pixel-level image tampering detection in modern VLMs like ChatGPT, Gemini, Qwen-Image, etc., aiming to learn tampering localization models that remain robust...

    arxiv.org/abs/2607.18230 · PDF

  3. 03

    GigaPath-Flash and GigaTIME-Flash: Efficient Pathology Foundation Models for Whole-Slide and Tumor Microenvironment Analysis

    Naoto Usuyama, Jeya Maria Jose Valanarasu, Sicong Yao, Hanwen Xu, Jaspreet Bagga, Guanghui Qin, Robert E. Kramer,...

    cs.CV · cs.AI

    Foundation models have emerged as a driving force in computational pathology, with the potential to transform cancer diagnosis, prognosis, and treatment selection by learning transferable representations from large-scale histopathology data. A growing landscape of pathology foundation models now spans diverse data sources, architectures, and downstream applications. However, most pretrained models operate only at the image-tile level, use...

    arxiv.org/abs/2607.18218 · PDF

  4. 04

    Certified Training for Convolutional Perturbations

    Benedikt Brückner, Alessio Lomuscio

    cs.CV · cs.LG

    Vision models have been found to be susceptible to perturbations such as motion blur induced at runtime by a shaking camera. This impedes their deployment in critical applications since phenomena such as slightly blurred vision might lead to failures, for example an object detector missing objects. While methods such as data augmentation or Adversarial Training can improve empirical robustness, they lack formal safety guarantees, making it...

    arxiv.org/abs/2607.18195 · PDF

  5. 05

    O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning

    Mei Yuan, Qi Long, Qifeng Wu, Zhenyang Li, Yizhou Zhao, Lei Wang, Yang Liu, Min Xu

    cs.CV · cs.AI · cs.CL · cs.MA

    Industrial Video Anomaly Detection (IVAD) aims to identify anomalous objects and events in an industrial process, which is crucial for modern manufacturing and quality control systems. Existing VLM-based anomaly reasoning methods are capable of detecting open-ended anomalies in general domains. However, their performance declines in industrial settings characterized by intricate object transformations, strict physics, and procedural...

    arxiv.org/abs/2607.18142 · PDF

  6. 06

    SciForma: Structure-Faithful Generation of Scientific Diagrams

    Yuxuan Luo, Peng Zhang, Xinjie Zhang, Xun Guo, Zhouhui Lian, Yan Lu

    cs.CV · cs.GR · cs.LG

    Structural fidelity is essential to scientific methodology diagrams. To communicate research logic, these diagrams must faithfully render components, directional relations, and textual annotations. Since a single error, such as a reversed arrow or an unreadable equation, can invalidate the entire figure, structural fidelity is inherently conjunctive: correctness on one axis cannot compensate for failure on another. Current open-source models...

    arxiv.org/abs/2607.18091 · PDF

  7. 07

    Sparse Evidence Can Suffice: Agentic Evidence Seeking for Multimodal Video Misinformation Detection

    Haochen Zhao, Yongxiu Xu, Xinkui Lin, Dong Xie, Jiarui Lu, Yuqi Qian, Yubin Wang, Hongbo Xu, Gaopeng Gou

    cs.CV · cs.AI

    Multimodal video misinformation detection is commonly formulated as a holistic video-understanding task, where the entire video and its associated content are processed and judged in a single pass. However, real-world misinformation often exhibits a sparse and compositional evidence structure: a reliable decision may depend on only a few coupled clues, while most video content contributes limited additional information. Exhaustive multimodal...

    arxiv.org/abs/2607.18080 · PDF

  8. 08

    Anticipate Before Acting: Future-State-Conditioned Vision-Language Navigation

    Lingfeng Zhang, Zhanguang Zhang, Liheng Ma, Tongtong Cao, Yingxue Zhang

    cs.CV · cs.AI

    End-to-end vision-language navigation (VLN) with causal vision-language models can map instructions and egocentric observations directly to actions, but standard behavior cloning supervises only the next action and does not explicitly train the policy state to be predictive of future visual outcomes. We first ask a diagnostic question: if the policy is given an expert-trajectory future image as privileged input at training and testing time,...

    arxiv.org/abs/2607.18042 · PDF

  9. 09

    HAS: Highlight-guided Attention Steering for Multimodal LLM Video Summarization

    Rui Chu, Yingjie Lao

    cs.CV · cs.AI

    Video understanding has become more and more important with the growth of Artificial Intelligence (AI) for video generation. Recently, Multimodal Large Language Model(M-LLM) has shown its capability in video understanding. Video summarization, a specific domain of video understanding, has proven its importance for efficient navigation and retrieval. Both video understanding and video summarization require a good selection of key frames in a...

    arxiv.org/abs/2607.17994 · PDF

  10. 10

    CaT-GS: Efficient 3DGS Rendering for Large Scale Scenes via Inter-frame Caching and Tile Scheduling

    Tingjia Zhang, Bo Chen, Shengzhong Liu, Fan Wu, Guihai Chen

    cs.CV · cs.AI

    Recent breakthroughs in 3D Gaussian Splatting (3DGS) have advanced neural rendering with high fidelity and speed. However, its performance degrades significantly in large-scale scenes due to the computational burden of tile-based rasterization. Existing optimization efforts either require costly scene re-training or focus on narrow aspects of the pipeline, overlooking critical inefficiencies in real-world deployments. Through a comprehensive...

    arxiv.org/abs/2607.17842 · PDF

  11. 11

    Measuring and Improving Complex-Atomic Answer Consistency in Endoscopic VQA

    Yuhao Liu, Cheng Zhao, Guanghui Yue

    cs.CV · cs.AI

    Endoscopic visual question answering (VQA) increasingly asks complex questions that combine several endoscopic answer components rather than isolated factual queries. Such complex answers may be scored as correct even when the same model fails on associated atomic questions. We introduce EndoCA, a paired complex-atomic answer consistency benchmark for evaluating whether complex answers remain consistent with same-image atomic answers. EndoCA...

    arxiv.org/abs/2607.17834 · PDF

  12. 12

    Vis2Reg: Visibility-Aware Landmark-Free Geometric 3D--2D Registration for Liver Laparoscopy

    Jiaming Feng, Xukun Zhang, Shahid Farid, Sharib Ali

    cs.CV · cs.AI · cs.HC

    Accurate 3D--2D liver registration, which aligns preoperative 3D models to partial, view-dependent intraoperative surface observations, is critical for AR-guided laparoscopic surgery but remains challenging due to severe occlusion, limited visibility, and the lack of 3D ground-truth supervision. Existing landmark-free approaches perform partial-to-complete geometric alignment, yet robust self-supervision under extreme partial visibility...

    arxiv.org/abs/2607.17810 · PDF

  13. 13

    ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video

    Xiaozhong Lyu, Gen Li, Zhiyin Qian, Xucong Zhang, Marc Pollefeys, Siyu Tang

    cs.CV · cs.AI

    Egocentric devices, such as wearable front-facing cameras, provide a unique perspective for capturing the continuous interaction between a human viewer and the surrounding environment. A holistic and efficient multimodal model capable of reconstructing this 4D representation is therefore highly desirable. However, existing approaches often rely on auxiliary inputs such as pre-computed camera trajectories, treat scene perception and human...

    arxiv.org/abs/2607.17790 · PDF

  14. 14

    Medical Imaging Fusing Vision Transformer: Laryngeal Cancer Screening with Explanation

    Haiyang Wang, Luca Mainardi

    cs.CV · cs.AI

    Early and timely screening of laryngeal cancer is crucial for improving clinical outcomes. In recent years, NBI endoscopy has become a standard diagnostic tool for the detection of laryngeal lesions. However, its effective use requires well-trained clinicians and the procedure is time-consuming and subject to interobserver variability. In this context, the application of artificial intelligence (AI) offers a promising solution to support...

    arxiv.org/abs/2607.17789 · PDF

  15. 15

    BrainNext: A General-Purpose Self-Supervised Foundation Model for Brain MRI Analysis

    Moona Mazher, Abdul Qayyum, Steven A. Niederer, Daniel C. Alexander

    cs.CV · cs.AI

    Foundation models pretrained using self-supervised learning have transformed computer vision by learning transferable representations from large-scale unlabeled data. However, existing foundation models for neuroimaging remain limited by task-specific training, slice-based learning strategies, or relatively small pretraining datasets, restricting their generalizability across diverse brain MRI applications. In this work, we present BrainNext,...

    arxiv.org/abs/2607.17782 · PDF

  16. 16

    CDIS: Cross-Dimensional Class-Agnostic 3D Instance Segmentation via 2D Mask Tracking and 3D-2D Projection Merging

    Juno Kim, Hye-Jung Yoon, Yesol Park, Byoung-Tak Zhang

    cs.CV · cs.AI

    Class-agnostic 3D instance segmentation is critical for robotic systems operating in unknown environments, enabling perception of previously unseen objects for reliable manipulation and navigation. Existing approaches typically project per-frame 2D instance masks into 3D and merge them, which often breaks object identities across time and yields fragmented 3D instances. We introduce Cross-Dimensional Class-Agnostic 3D Instance Segmentation...

    arxiv.org/abs/2607.17778 · PDF

  17. 17

    Measuring Monosemanticity in Sparse Autoencoders via Latent Activation Coherence

    Katarzyna Filus, Sebastian Pokuciński

    cs.CV · cs.AI · cs.LG

    Within Explainable Artificial Intelligence, mechanistic interpretability uses Sparse Autoencoders (SAEs) to extract more interpretable features from neural representations. However, assessing their monosemanticity, and thus explanation quality, remains challenging. Existing metrics require external concept labels or depend on pretrained embedding models, making them sensitive to encoder's geometry. We introduce the Tversky Monosemanticity...

    arxiv.org/abs/2607.17770 · PDF

  18. 18

    DA-Fusion: Deformable Attention-Based RGB-D Fusion Transformer for Unseen Object Instance Segmentation

    Yesol Park, Hye-Jung Yoon, Juno Kim, Byoung-Tak Zhang

    cs.CV · cs.AI

    In logistics automation, precise segmentation of unseen objects is crucial for efficient robotic manipulation in cluttered environments. Tasks such as bin-picking and shelf-picking require robust perception to handle occlusions, varying object shapes, and complex spatial arrangements. Traditional RGB-based methods tend to over-segment objects due to their reliance on texture, while depth-based methods often under-segment by focusing primarily...

    arxiv.org/abs/2607.17754 · PDF

  19. 19

    Early Yield Prediction for Sugar Beet Fields using Satellite Data -- Learnings from Specialized Vision Transformers

    Philipp Vaeth, Bhumika Laxman Sadbhave, Denise Dejon, Gunther Schorcht, Magda Gregorova

    cs.CV · cs.LG

    Remote sensing has become an increasingly valuable tool for agricultural monitoring, particularly through the use of publicly available satellite imagery. However, effectively integrating domain knowledge into machine learning methods remains challenging. This study presents a real-world example of early sugar beet harvest yield forecasting from purely optical Sentinel-2 imagery, demonstrating how a tight integration of domain knowledge and...

    arxiv.org/abs/2607.17661 · PDF

  20. 20

    LFM: Leveraging Foundation Models for Source-Free Universal Domain Adaptation

    Jing Li, Pan Liu, Meng Zhao, Wanli Xue, Yanhong Yang, Xu Cheng, Fan Shi, Jianhua Zhang, Qinghua Hu, Shengyong Chen

    cs.CV · cs.LG · cs.MM

    Source-free universal domain adaptation (SF-UniDA) adapts a pre-trained source model to an unlabeled target domain under both covariate and label shifts, without access to source data. However, existing SF-UniDA methods rely on inefficient techniques such as threshold tuning and clustering. Foundation models (FMs), known for their generalization and zero-shot capabilities, remain underexplored in SF-UniDA. In this paper, we propose a...

    arxiv.org/abs/2607.17653 · PDF

  21. 21

    Brain-Aligned Multi-Stream Video Transformers with Sparse Self-Selection

    Amir Hosein Fadaei, Mahyar Maleki, Mohammad-Reza A. Dehaqani

    cs.CV · cs.LG

    Modern video transformers typically ignore principles from primate vision and are rarely evaluated against neural data, limiting their biological interpretability. We introduce a sparse winner-takes-all token selection module that replaces dense self-attention to improve efficiency and approximate competitive routing observed in biological visual circuits. We further propose a neuro-inspired split-and-fuse video transformer which uses two...

    arxiv.org/abs/2607.17625 · PDF

  22. 22

    Coarse-to-fine Framework for Generative MEF via Implicit Neural Representation

    Sangmin Han, Jinho Kim, Jinwoo Kim, Dongyoung Kim, Seon Joo Kim

    cs.CV · cs.AI

    Multi-exposure fusion (MEF) expands the luminance range beyond what a single exposure can capture. Combining images taken at different exposure levels requires handling geometric differences while naturally merging their complementary brightness information. It often demands generative completion where details are missing. Diffusion-based generative methods address these challenges, however, they are computationally expensive and struggle to...

    arxiv.org/abs/2607.17611 · PDF

  23. 23

    Semantic Color Naturalness Breaker: Preventing Illegitimate Colorization via Content-Aware Color Priors

    Yuki Nii, Futa Waseda, Ching-Chun Chang, Isao Echizen

    cs.CV · cs.LG

    Automatic image colorization enables large-scale and low-cost reuse of grayscale media (e.g., manga panels and archival photographs), facilitating unauthorized reuse and redistribution. Once released online, grayscale content can be readily turned into unauthorized colorized derivatives using off-the-shelf models, creating a practical need for proactive, content-side protection at publication time. Building on Uncolorable Examples (UE), which...

    arxiv.org/abs/2607.17610 · PDF

  24. 24

    Hierarchy-Aware and Anatomy-Guided Learning for Lung Ultrasound Video Classification

    Alya Almsouti, Lotfi Mecharbat, Noha Aboukhater, Yousef Alabrach, Siddiq Anwar, Andre Kumar, Ibrahim Almakky, Mohammad Yaqub

    cs.CV · cs.AI

    Lung ultrasound (LUS) is a bedside tool for assessing pulmonary edema in patients at risk due to heart failure or impaired kidney function. However, automated LUS analysis remains challenging because of speckle noise, imaging artifacts, and operator-dependent acquisition variability. In this work, we present a deep learning framework for multi-class LUS video classification that explores two components: hierarchy-aware training, and...

    arxiv.org/abs/2607.17551 · PDF

  25. 25

    Thinking in Video: Can Video Generators Really Reason About the Real World?

    Yongheng Zhang, Guang Yang, Ruihan Hou, Qiguang Chen, Ziang Liu, Xiaolong Liu, Manman Zhang, Yanchao Hao, Zheng Wei,...

    cs.CV · cs.AI · cs.CL

    Recent advances in world models and video generation have given rise to an emerging reasoning paradigm that leverages video generative models to simulate, predict, and reason about real-world dynamics. We redefine this paradigm as Thinking in Video, where video is not merely an output artifact but a medium for constructing, extending, and verifying causal thought. However, this promise remains unverified: convincing rollouts may reflect...

    arxiv.org/abs/2607.17523 · PDF

  26. 26

    DecoyFace: Beyond Obfuscation via Controllable and Imperceptible Identity Misdirection for Privacy-Preserving Face Recognition

    Zhihan Ren, Lijun He, Xinyao Wang, Xinzhu Fu, Fan Li

    cs.CV · cs.AI

    Split face recognition reduces client-side computation but exposes intermediate features to feature inversion attacks and unauthorized analysis by honest-but-curious (HBC) servers. Existing privacy-preserving face recognition methods mainly aim to resist unauthorized reconstruction, typically producing features whose inversion yields visibly degraded results, which may reveal the existence of protection and motivate adaptive attacks. To...

    arxiv.org/abs/2607.17504 · PDF

  27. 27

    DA-MergeLoRA: Hypernetwork-Based LoRA Merging for Few-Shot Test-Time Domain Adaptation

    Siobhan Reid, Zhixiang Chi, Li Gu, Omid Reza Heidari, Ziqiang Wang, Yang Wang

    cs.CV · cs.LG

    Few-shot Test-Time Domain Adaptation (FSTT-DA) seeks to adapt models to novel domains using only a handful of unlabeled target samples. This setting is more realistic than typical domain adaptation setups, which assume access to target data during source training. However, prior FSTT-DA approaches fail to effectively leverage source domain-specific knowledge, relying on shallow batch normalization updates, prompt-based methods that treat the...

    arxiv.org/abs/2607.17467 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.