cs.CV · 2026-06-18 · No. 27

Computer Vision and Pattern Recognition, 2026-06-18.

12 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

12 entries
  1. 01

    Confidence is Not Reliability: Rethinking MC Dropout in Brain Tumour Segmentation

    Xin Ci Wong, Duygu Sarikaya, Kieran Zucker, Marc De Kamps, Nishant Ravikumar

    cs.CV · cs.LG

    Glioma segmentation in multiparametric MRI is a critical component of treatment planning. A segmentation model that fails silently on treatment-critical sub-regions represents a patient safety risk that overlap-based metrics such as Dice scores cannot expose. We ask whether voxel-level uncertainty estimation via Monte Carlo (MC) Dropout can reliably identify segmentation errors in clinically critical sub-regions, and whether calibration...

    arxiv.org/abs/2606.19300 · PDF

  2. 02

    A Multi-Domain Benchmark for Detecting AI-Generated Text-Rich Images from GPT-Image-2

    Yijin Wang, Shuyi Wang, Wenhan Zhang, Yuqi Ouyang

    cs.CV · cs.AI

    Text-rich images often contain privacy-sensitive, transactional, or decision-relevant information. As recent multimodal image generation models become increasingly capable of synthesizing realistic textual content and structured visual designs, detecting AI-generated text-rich images has become an important challenge for digital trust and content authenticity. Existing benchmarks, however, largely focus on object-centric images and provide...

    arxiv.org/abs/2606.19259 · PDF

  3. 03

    OneCanvas: 3D Scene Understanding via Panoramic Reprojection

    Bartłomiej Baranowski, Dave Zhenyu Chen, Matthias Nießner

    cs.CV · cs.AI · cs.LG · cs.RO

    Existing approaches to 3D scene understanding in Vision-Language Models (VLMs) either rely on complex, model-specific geometry encoders or large training budgets in pursuit of spatial reasoning. Instead, OneCanvas aggregates patch features from all views onto a single equirectangular panoramic canvas. Namely, each patch is unprojected to a 3D world coordinate using its depth and camera pose, then placed on the canvas at the continuous...

    arxiv.org/abs/2606.19253 · PDF

  4. 04

    Transformer Geometry Observatory TGO-I: Spectral Geometry Observatory

    Kaustubh Kapil, Kishor P. Upla

    cs.CV · cs.LG

    Despite the widespread adoption of Vision Transformers (ViTs) and their success across numerous computer vision applications, the fundamental understanding of their dimensional and representational geometry remains relatively underexplored. To address this gap, we introduce Transformer Geometry Observatory (TGO), a systematic framework of experiments and analysis pipelines designed to investigate the representational geometry and dynamics of...

    arxiv.org/abs/2606.19249 · PDF

  5. 05

    When AUC Misleads: Polarization-Aware Evaluation of Deepfake Detectors under Domain Shift

    Dat Nguyen, Cosmin Radoi, Romain Hermary, Marcella Astrid, Nesryne Mejri, Enjie Ghorbel, Djamila Aouada

    cs.CV · cs.LG

    Recent advances in generative AI, such as diffusion models and face-swapping tools, have enabled the creation of highly realistic deepfakes, leading to real-world harms including financial fraud and non-consensual explicit content. In response, deepfake detection has become an active research area, with recent methods increasingly focusing on improving generalization to unseen manipulations. This is typically evaluated using the Area Under...

    arxiv.org/abs/2606.19184 · PDF

  6. 06

    ProductConsistency: Improving Product Identity Preservation in Instruction-Based Image Editing via SFT and RL

    Mukund Khanna, Raj Singh Yadav, Kunal Singh

    cs.CV · cs.AI

    Recent advances in instruction-based image editing have enabled models to perform complex visual edits from natural language instructions. However, in product-centric scenarios where preserving product features, branding, and textual elements are critical, current open and closed source models often struggle to maintain this fine-grained object identity. This issue is further compounded by the lack of datasets for instruction-based product...

    arxiv.org/abs/2606.19103 · PDF

  7. 07

    Test-Time Adaptation in Optical Coherence Tomography Using Trajectory-Aligned Time-Independent Flow

    Veit Hucke, Thomas Pinetz, Gregor Reiter, Ursula Schmidt-Erfurth, Hrvoje Bogunović

    cs.CV · cs.LG

    Optical coherence tomography (OCT) is essential in ophthalmology, but inconsistent image quality especially in low-cost devices hinders automated analysis. To address this, we introduce a flow-matching-based test-time adaptation method that generates high-quality surrogate images from noisy inputs. Typically, domain gaps between test and training data cause pixel distribution mismatches during the denoising process. We overcome this by...

    arxiv.org/abs/2606.18876 · PDF

  8. 08

    URDF Synthesis from RGB-D Sequences via Differentiable Joint Inference and Energy-Consistent Verification

    Xinze Zhang

    cs.CV · cs.AI

    Reconstructing simulation-ready digital twins of articulated objects from sensor observations remains constrained by two persistent gaps: (i) part-level geometric reconstruction is decoupled from kinematic-parameter estimation, and (ii) the recovered models often violate basic dynamic invariants such as energy conservation, leading to drift when the URDF is replayed in physics simulators. We present KinemaForge, a constraint-driven pipeline...

    arxiv.org/abs/2606.18861 · PDF

  9. 09

    Quantification of Uncertainty with Adversarial Models in Medical Image Segmentation

    Hana Jebril, Thomas Pinetz, Günter Klambauer, Hrvoje Bogunović

    cs.CV · cs.LG

    Reliable pixel-level uncertainty quantification holds the potential to transform clinical workflows by enabling high-fidelity longitudinal monitoring and distinguishing true pathological changes from artifacts. Ideally, these models provide the stability required for critical treatment planning and surgical intervention. However, standard deep learning models often suffer from miscalibration, yielding overconfident predictions that mask...

    arxiv.org/abs/2606.18860 · PDF

  10. 10

    Where Will They Go? Modelling Multimodal Pedestrian Manoeuvres from Ego-centric Videos

    Yuxuan Xie, Nicolas Pugeault, Chongfeng Wei, Hubert P. H. Shum, Edmond S. L. Ho

    cs.CV · cs.LG

    Pedestrian trajectory prediction from an ego-centric camera is challenging since it depends on complex interactions with vehicles and scene context, as well as the intention of the pedestrian. By modelling correlation and intent from the historical and future trajectories of the pedestrian, it will usually result in a multimodal (i.e. multiple modes) distribution. Existing stochastic predictors often sample multiple futures from a single...

    arxiv.org/abs/2606.18824 · PDF

  11. 11

    Clinically Aligned Geometry Constraints for Robust IVUS Vessel Boundary Segmentation

    Yunshu Chen, Litao Yang, Giuseppe Di Giovanni, Jordan Tan, Deval Mehta, Andrew Lin, Derek Chew, Masasi Fujino, Julie...

    cs.CV · cs.LG

    Intravascular ultrasound (IVUS) lumen and external elastic membrane (EEM) segmentation is important for quantitative coronary plaque burden assessment. Errors in lumen or EEM delineation directly propagate to plaque area, plaque burden and geometric measurements. However, standard methods prioritising overlap scores often suffer from boundary drift and topology errors, leading to inaccurate clinical measurements. We present GeoCat, a...

    arxiv.org/abs/2606.18723 · PDF

  12. 12

    LandslideAgent with Multimodal LandslideBench: A Domain-Rule-Augmented Agent for Autonomous Landslide Identification and Analysis

    Chengfu Liu, Dongyang Hou, Junwu Xiang, Cheng Yang, Xuezhi Cui, Zeyuan Wang, Liangtian Liu, Zelang Miao

    cs.CV · cs.AI

    Intelligent landslide hazard interpretation is critical for disaster prevention, yet current paradigms struggle to simultaneously extract visual features and high-level geoscientific semantics, while general-purpose vision-language models (VLMs) suffer from perceptual limitations and domain hallucinations in complex geological scenarios. To address these challenges, we propose an instruction-driven agentic framework comprising three...

    arxiv.org/abs/2606.18661 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.