cs.CV · 2026-06-30 · No. 39

Computer Vision and Pattern Recognition, 2026-06-30.

21 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

21 entries
  1. 01

    Beyond 2D Matching: A Unified Single-Stage Framework for Geometry-Aware Cross-View Object Geo-Localization

    Liyao Wang, Ruipu Wu, Haojun Xu, Lei Shi, Linjiang Huang, Si Liu

    cs.CV · cs.AI

    Cross-view object geo-localization (CVOGL) aims to locate a target object from a query view (e.g., ground or drone) within a geo-tagged reference image (e.g., satellite). Existing approaches heavily rely on 2D appearance matching and are constrained by limited datasets lacking geometric metadata, diverse prompts, and standard field-of-view imagery. To address these intertwined challenges, we first introduce \dataset, a large-scale,...

    arxiv.org/abs/2606.30576 · PDF

  2. 02

    $μ$Flow: Leveraging Average Images for Improving Generalisation of Deepfake Faces Detectors

    Orazio Pontorno, Mattia Litrico, Luca Guarnera, Mario Valerio Giuffrida, Sebastiano Battiato

    cs.CV · cs.LG

    Current generative models, including GANs and diffusion models, have reached an outstanding level of photorealism, posing significant risks to privacy and security. To ensure real-world applicability, deepfake detectors must generalise effectively to unseen generators. However, most existing approaches rely on supervised training with both real and fake images, which limits their generalisation especially across generators categories (e.g....

    arxiv.org/abs/2606.30528 · PDF

  3. 03

    On the Faithfulness of Post-Hoc Concept Bottleneck Models

    Laines Schmalwasser, Jan Blunk, Niklas Penzel, Julia Niebling, Joachim Denzler

    cs.CV · cs.AI

    Human decision-making interprets the world through high-level concepts, such as recognizing a bird by its belly color. To bridge the gap between opaque deep learning representations and human understanding, Post-Hoc Concept Bottleneck Models (post-hoc CBMs) project latent features onto interpretable concept spaces using auxiliary datasets or vision-language models. However, relying on target task accuracy as the primary measure of post-hoc...

    arxiv.org/abs/2606.30498 · PDF

  4. 04

    Beyond Point Estimates for Glaucoma Visual Field Forecasting with Diffusion Models

    Marta Colmenar Herrera, Pablo Márquez Neila, Şerife Seda Kucur Ergünay, Martin S. Zinkernagel, Raphael Sznitman

    cs.CV · cs.AI

    Forecasting visual fields (VFs) is critical for personalized monitoring and treatment planning in glaucoma. This is inherently uncertain due to heterogeneous disease progression and measurement variability, yet most existing methods produce single deterministic predictions that fail to represent this uncertainty. We formulate VF forecasting as a probabilistic prediction problem and the use of conditioned denoising diffusion models to generate...

    arxiv.org/abs/2606.30417 · PDF

  5. 05

    Set-Inclusive Uncertainty Modeling for Robust Brain Tumor Segmentation

    Seunghun Baek, Jihwan Park, Jaeyoon Sim, Hoseok Lee, Seungjoo Lee, Won Hwa Kim

    cs.CV · cs.AI · cs.LG

    Multimodal MRI is essential for accurate brain tumor segmentation. However, acquiring all modalities at inference is often challenging in practice, which causes intrinsic uncertainty due to unavoidable information loss. Without modeling this uncertainty, existing methods encode incomplete evidence into deterministic representations that appear plausible but lack reliability. In this regime, we propose a probabilistic representation framework...

    arxiv.org/abs/2606.30374 · PDF

  6. 06

    Residual-Guided Expert Specialization for Incomplete Multimodal Learning

    Seunghun Baek, Jihwan Park, Jaeyoon Sim, Minjae Jeong, Hoseok Lee, Won Hwa Kim

    cs.CV · cs.AI

    As real-world prediction systems often face missing modalities at inference, incomplete multimodal learning (IML) remains a practical challenge. While prior methods aim to learn representations robust to missing inputs, representations from incomplete modalities inevitably deviate from their full-modality counterparts due to missing evidence. To explicitly leverage these deviations, we propose MARS (Missingness-Aware Residual-guided...

    arxiv.org/abs/2606.30355 · PDF

  7. 07

    FFAvatar: Feed-Forward 4D Head Avatar Reconstruction from Sparse Portrait Images

    Jianjiang Yao, Ke Xian, Renxiang Dai, Robert Caiming Qiu

    cs.CV · cs.AI

    We present FFAvatar, a Transformer-based 3D Gaussian framework for fast construction of high-quality and animatable 4D head avatars from one or more reference portrait images. Unlike existing feed-forward approaches that require a fixed number of input views, FFAvatar supports incremental reconstruction, progressively refining the avatar representation as additional reference images become available. At the core of our method is an...

    arxiv.org/abs/2606.30347 · PDF

  8. 08

    Early Cue Precision Shapes Visual Shortcut Learning in Controlled Cue-Manipulation Benchmarks

    Chanho Park, Woochan Lee, Janyeong Oh, Geongho Gong, Minshu Kim, Yeachan Kwak, Seongim Choi

    cs.CV · cs.AI

    Visual classifiers can achieve high matched-distribution accuracy while relying on low-level cues that fail under conflict or suppression. We test whether this failure is shaped by early cue precision: the reliability with which a low-level cue predicts the label during early learning or downstream probe fitting. Across synthetic shape-texture tasks, sequential digit training, a 10-class frozen-representation audit, and a CIFAR-10...

    arxiv.org/abs/2606.30344 · PDF

  9. 09

    BrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and Language

    Haitao Wu, Qirui Zhang, Zhouheng Yao, Shangquan Sun, Qihao Zheng, Mianxin Liu, Chi Zhang, Wanli Ouyang, Chunfeng...

    cs.CV · cs.LG

    Modeling the bidirectional correspondence between external sensory stimuli and internal neural activity has emerged as a critical frontier in neuroscience. However, existing approaches predominantly treat brain encoding and decoding as isolated tasks, relying heavily on unimodal alignment and external priors while overlooking the brain's intrinsic nature as a multimodal integration system. To address these limitations, we propose BrainJanus,...

    arxiv.org/abs/2606.30319 · PDF

  10. 10

    TRACE: A Concept Bottleneck Model for Longitudinal 3D Glioblastoma Response Assessment

    Alia Tarek, Hamsa Saberr, Hamza Elghonemy, Youssef Afify, Tamer Basha, Omair Shahzad Bhatti, Abdulrahman M. Selim,...

    cs.CV · cs.LG

    Longitudinal glioblastoma response assessment requires comparing subtle tumor changes across MRI time points using structured clinical criteria such as RANO. However, most deep learning methods predict response labels directly from imaging features, which limits clinical inspection, verification, and correction. We introduce TRACE, a RANO 2.0-aligned concept bottleneck model for interpretable 4-class glioblastoma response classification on...

    arxiv.org/abs/2606.30313 · PDF

  11. 11

    Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation

    Shihao Zhang, Yuguang Yan, Junzhe Zhang, Wei Zhao, Bohan Wang, Hanwang Zhang

    cs.CV · cs.LG

    Recent text-to-video (T2V) diffusion models rely heavily on auxiliary reward signals (e.g., via reward models or DPO) to align generated content with human aesthetics and improve realism. These signals, however, incur substantial computational overhead, require costly human annotations, and often yield limited improvement in fine-grained local details. In this paper, we argue that your data manifold is secretly a reward model. By explicitly...

    arxiv.org/abs/2606.30248 · PDF

  12. 12

    Efficient RGB-T Object Detection via Sparse Cross-Modality Fusion

    Chao Tian, Zikun Zhou, Chao Yang, Guoqing Zhu, Zhenyu He

    cs.CV · cs.AI

    RGB-T detectors leverage the complementary strengths of visible and thermal infrared modalities, achieving robust performance under challenging conditions. Many of them resort to heavy dual backbones and exhaustive cross-modality fusion across the entire image, leading to impractically high computational costs. We observe that most image regions are smooth backgrounds (e.g., sky, ground) that can be easily handled by lightweight...

    arxiv.org/abs/2606.30215 · PDF

  13. 13

    A Multi Center Breast FNAC Whole-Slide Cytology Dataset for AI-Assisted Patch-Wise Classification Using C1 to C5 Reporting Categories

    Garima Jain, Abhijeet Patil, Surabhi Jain, Sanghamitra Pati, Amit Sethi, Sandeep Mathur, Pulkit Verma, Nishi...

    cs.CV · cs.AI

    We present a multi center breast fine needle aspiration cytology (FNAC) dataset designed for patch wise classification using C1 to C5 reporting labels. The prospective dataset includes 321 patients and 470 whole-slide images (WSIs) collected from participating tertiary medical centers in India between May 2023 and March 2026. Slides were stained using Papanicolaou (190 WSIs) or MayGrunwald Giemsa (280 WSIs), scanned on a Hamamatsu NanoZoomer...

    arxiv.org/abs/2606.30209 · PDF

  14. 14

    Few-Shot Domain Incremental Learning via Continual Vision-Language Consolidation

    Naeem Paeedeh, Mahardhika Pratama, Wolfgang Mayer, Mukesh Prasad, Weiping Ding, Yew-Soon Ong

    cs.CV · cs.AI · cs.LG

    Existing domain-incremental learning (DIL) strategies call for massive amounts of data to adapt to new domains and suffer from the overfitting problem in the case of data scarcity. This paper puts forward a relatively uncharted problem, namely, few-shot domain incremental learning (FSDIL), taking into account the problem of extreme data shortages in the realm of DIL. A novel algorithm, namely Continual Vision-Language Consolidation (CVLC), is...

    arxiv.org/abs/2606.30190 · PDF

  15. 15

    Hyper-Network Neural Functional Maps for Unsupervised Robust 3D Shape Matching

    Dongliang Cao, Florian Bernard

    cs.CV · cs.AI

    Functional maps are the cornerstone of recent non-rigid 3D shape matching methods due to their efficiency and performance. However, existing methods struggle with challenging scenarios, such as partiality, topological noise, and raw point clouds. A primary bottleneck is that significant intrinsic distortion prevents truncated spectral bases from being accurately aligned via linear transformations (i.e., functional maps). To address this, we...

    arxiv.org/abs/2606.30131 · PDF

  16. 16

    Bridging the Gap Between Image Restoration and Navigational Safety in Hazy Conditions: A New Visibility Estimation Metric for Maritime Surveillance

    Wentao Feng, Guobei Peng, Wengang Mao, Ryan Wen Liu

    cs.CV · cs.LG

    Visibility distance is critical to maritime navigational safety because it determines the effective observation range of shipborne and shore-based monitoring systems. Under hazy conditions, degraded visual information shortens observable distance and increases navigational risks and economic losses. Although numerous image dehazing methods have been developed, conventional image quality assessment metrics, such as PSNR, SSIM, FSIM, FADE, and...

    arxiv.org/abs/2606.30049 · PDF

  17. 17

    Consensus Clustering of Free-Viewing Gaze Data: New Insights into Human-Information Interaction

    Beryl Gnanaraj, Jaya Sreevalsan-Nair, Saqib Alam Ansari, Maanasa Rajaraman

    cs.CV · cs.HC · cs.LG

    Free-viewing gaze data provides a rich, task-free window into human visual attention. Conventional exploratory data analysis of the data provides user attention patterns through fixations and areas of interest. However, despite the richness of this gaze data, its human-information interaction (HII) patterns are understudied. We address this gap using consensus clustering of gaze data with respect to users and stimulus characteristics. We...

    arxiv.org/abs/2606.30035 · PDF

  18. 18

    MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs

    Yuxuan Fan, Gyusik Seo, Jing Hao, Jaemin Cho, Mohit Bansal, Jaehong Yoon

    cs.CV · cs.AI

    Audiovisual arts encompass diverse creative disciplines, including cinema, visual arts, stage performance, and game design, where artistic meaning arises from deliberate combinations of visual, auditory, and narrative elements (e.g., fear amplified through claustrophobic framing, or grief conveyed through silence and lingering close-ups). True artistic understanding extends beyond recognizing what is depicted to reasoning about why it is...

    arxiv.org/abs/2606.30026 · PDF

  19. 19

    IBRSteG: Learning a Generalizable Steganography Framework for 3D Gaussian Splatting

    Fanye Kong, Hongyu Xia, Yu Zheng, Boyang Gong, Jie Zhou, Jiwen Lu

    cs.CV · cs.AI

    Recent advances in deep learning have notably improved steganographic message hiding. However, designing a generalizable steganographic approach for 3D Gaussian Splatting (3DGS) that can embed meaningful 3D scene content remains challenging. In this paper, we propose IBRSteG, a generalizable framework for 3DGS steganography that enables undetectable concealment of secret scenes within a steganographic scene. Unlike existing approaches whose...

    arxiv.org/abs/2606.30024 · PDF

  20. 20

    Latent-CURE for Breast Cancer Diagnosis

    Weiyi Zhao, Xiaoyu Tan, Lu Gan, Liang Liu, Xihe Qiu

    cs.CV · cs.AI

    Multimodal Large Models have significantly advanced automated breast ultrasound diagnosis. However, most existing frameworks utilize opaque, end-to-end paradigms prioritizing global statistical correlations over structured clinical reasoning. Consequently, these models remain susceptible to shortcut learning amid extreme real-world epidemiological imbalances, often bypassing rare but decisive malignant indicators for dominant benign patterns....

    arxiv.org/abs/2606.29928 · PDF

  21. 21

    LLM-based Multimodal Personality Recognition via Facial Action Unit-Text Semantic Fusion

    Tianyi Zhang, Wei Shan, Yuan Zong, Tianhua Qi, Wenming Zheng

    cs.CV · cs.AI

    Personality recognition in asynchronous video interviews (AVIs) has become increasingly important due to their widespread adoption in modern recruitment. Existing approaches often rely on large language models (LLMs) to analyze textual responses of interviewees in AVI. However, unimodel methods often suffer from information loss (e.g., ignore facial cues). In contrast, multimodal methods that employ full-face images or sparsely sampled frames...

    arxiv.org/abs/2606.29900 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.