cs.CV · 2026-06-19 · No. 28

Computer Vision and Pattern Recognition, 2026-06-19.

22 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

22 entries
  1. 01

    UNIEGO: Proxies as Mediators for Unified Egocentric Video Representation Learning

    Wenhao Chi, Arkaprava Sinha, Dominick Reilly, Hieu Le, Srijan Das

    cs.CV · cs.LG

    Egocentric video understanding is inherently limited by the narrow perspective of wearable cameras: a single viewpoint, a single modality, a single model cannot capture the full richness of human action. We argue that a truly expressive egocentric representation must subsume complementary knowledge across viewpoints, modalities, and foundation model representations, yet remain deployable from egocentric video alone. To this end, we introduce...

    arxiv.org/abs/2606.20559 · PDF

  2. 02

    SARLO-80: Worldwide Slant SAR Language Optic Dataset 80cm

    Solène Debuysère, Nicolas Trouvé, Nathan Letheule, Elise Colin, Georgia Channing

    cs.CV · cs.AI · cs.DB

    Multimodal foundation models have advanced rapidly thanks to large optical benchmarks, but comparable resources for synthetic aperture radar (SAR) remain limited. Existing SAR--optical datasets largely rely on low-resolution, intensity-only Ground Range Detected~(GRD) products and do not preserve complex-valued SAR measurements or native acquisition geometry, which restricts physically grounded multimodal learning. In particular, large-scale...

    arxiv.org/abs/2606.20523 · PDF

  3. 03

    FreeStyle: Free Control of Style-Content Dual-Reference Generation from Community LoRA Mining

    Jinghong Lan, Wei Cheng, Yunuo Chen, Ziqi Ye, Peng Xing, Yixiao Fang, Rui Wang, Yufeng Yang, Xuanyang Zhang,...

    cs.CV · cs.AI

    Style-content dual-reference generation aims to synthesize an image that preserves the structure and semantics of a content reference while adopting the style of a separate style reference.Despite recent progress, this setting remains challenging because models must balance content fidelity, style alignment, and instruction following avoiding semantic leakage from the style reference.A key bottleneck is the lack of large-scale triplet data...

    arxiv.org/abs/2606.20506 · PDF

  4. 04

    Scalable Training of Spatially Grounded 2D Vision-Language Models for Radiology

    Yusuf Salcan, Simon Ging, Robin Schirrmeister, Philipp Arnold, Elmar Kotter, Behzad Bozorgtabar, Thomas Brox

    cs.CV · cs.CL · cs.LG

    We study how to train visually grounded vision-language models (VLMs) for radiology without manual spatial annotations. We introduce RefRad2D, a large-scale bilingual (German/English) dataset of 1.2M CT and MR image-text pairs derived from clinical practice, with task-specific VQA and spatial grounding subsets generated automatically via LLM-based curation and automated segmentation. Trained on this data, our model RadGrounder jointly...

    arxiv.org/abs/2606.20477 · PDF

  5. 05

    SPOT-E: Test-Time Entropy Shaping with Visual Spotlights for Frozen VLMs

    Bo Yin, Xiaobin Hu, Chengming Xu, Ruolin Shen, Mo Yang, Jiangning Zhang, Peng-Tao Jiang, Cheng Tan, Shuicheng YAN

    cs.CV · cs.AI

    Vision-language models (VLMs) often underperform on evidence intensive tasks because decisive visual evidence are small, localized, and easy to overlook, leading to failures in evidence readout even when high-level reasoning is intact. Prior inference-time visual interventions can improve grounding without retraining, but they are largely open-loop and lack a mechanism to verify whether highlighted evidence is actually used. We study...

    arxiv.org/abs/2606.20244 · PDF

  6. 06

    HilDA: Hierarchical Distillation with Diffusion for Advancing Self-Supervised LiDAR Pre-trainin

    Maciej Wozniak, Jesper Ericsson, Hariprasath Govindarajan, Truls Nyberg, Thomas Gustafsson, Patric Jensfelt, Olov Andersson

    cs.CV · cs.AI · cs.RO

    Leveraging Vision Foundation Models (VFMs) for camera-to-LiDAR knowledge distillation offers a promising solution to the scarcity of annotated data needed to represent the immense geometric and kinematic diversity of real-world autonomous driving (AD). However, current approaches typically treat VFMs as black-box teachers, relying exclusively on frame-wise feature similarity. Consequently, they do not fully exploit the teacher's layer-wise...

    arxiv.org/abs/2606.20189 · PDF

  7. 07

    Evaluating and Enhancing Negation Comprehension in Remote Sensing MLLMs

    Haochen Han, Jue Wang, Alex Jinpeng Wang, Fangming Liu

    cs.CV · cs.AI

    Multimodal Large Language Models (MLLMs) have demonstrated remarkable success in various Remote Sensing (RS) tasks. However, their ability to comprehend negation remains underexplored, limiting deployment in real-world applications where models must explicitly identify what is false or absent, e.g., emergency responders need to locate non-flooded routes for evacuation. To comprehensively study this limitation, we introduce RS-Neg, the first...

    arxiv.org/abs/2606.20177 · PDF

  8. 08

    EFIQA: Explainable Fundus Image Quality Assessment via Anatomical Priors

    Pengwei Wang, José Morano, Qian Wan, Hrvoje Bogunović

    cs.CV · cs.LG

    Image quality control is vital for a wide range of downstream applications. Deep learning-based image quality assessment methods typically train classifiers on dataset-specific quality labels, inheriting two limitations: (1) generalization is tied to the labeling criteria of the training set and (2) these methods cannot provide spatial feedback on where the quality is degraded, lacking explainability. In this work, we propose EFIQA, a...

    arxiv.org/abs/2606.20108 · PDF

  9. 09

    MakeupMirror: Improving Facial Attribute Preservation in Diffusion Models for Makeup Transfer

    Nefeli Andreou, Angel Martínez-González, Sabine Sternig, Matthieu Guillaumin, Epameinondas Antonakos, Michael Opitz

    cs.CV · cs.AI · cs.GR · cs.LG · cs.MM

    Makeup transfer models enable fun augmented reality (AR) experiences as well as virtual try-on (VTO) for online makeup shopping. While recent state-of-the-art diffusion based solutions such as Stable-Makeup dramatically improve the accuracy and realism of makeup transfer, they still face limitations in identity and skin color preservation, making production-level VTO for makeup shopping unrealistic. In this work, we propose MakeupMirror, a...

    arxiv.org/abs/2606.20094 · PDF

  10. 10

    The Hidden Evolution of Disguised Visual Context inside the VLM

    Wish Suharitdamrong, Tony Alex, Muhammad Awais, Sara Atito

    cs.CV · cs.AI

    Visual tokens enter Large Language Models (LLMs) as raw, foreign signals. How they are transformed into meaningful representations and interact with the language space depends entirely on the integration architecture. Whether by treating visual tokens as in-context prompts within the input sequence or injecting them directly into the LLM's intermediate layers. A controlled comparison and understanding of how these architectural choices affect...

    arxiv.org/abs/2606.20077 · PDF

  11. 11

    Variable-Length Tokenization via Learnable Global Merging for Diffusion Transformers

    Dong Hoon Lee, Seunghoon Hong

    cs.CV · cs.AI

    Latent Diffusion Models (LDMs) have become dominant in visual synthesis, but their quality-compute trade-off is largely constrained by the tokenizer's fixed compression ratio. Variable-length tokenizers (VLTs) promise adaptive compression by varying token counts, allowing diffusion models to flexibly balance quality and compute. However, conventional VLTs modulate length by truncating ordered token sequences, which makes token semantics...

    arxiv.org/abs/2606.20076 · PDF

  12. 12

    See-and-Reach: Precise Vision-Language Navigation for UAVs within the Field of View

    Fanfu Xue, En Yu, Yantian Shen, Zhikun Hu, Hongjun Wang, Yang Yang, Xindi Wang, Jiande Sun

    cs.CV · cs.AI

    UAV Vision-Language Navigation (UAV-VLN) is typically formulated as a holistic search-and-reach problem, where long-range target discovery and final target approach are optimized and evaluated jointly. This formulation makes it difficult to assess a critical capability of aerial embodied agents, namely whether a UAV can accurately ground a visible target and translate vision-language evidence into precise 3D motion once the target enters its...

    arxiv.org/abs/2606.20045 · PDF

  13. 13

    PU-UNet: Stable Multiplicative Interactions for Medical Image Segmentation

    Ziyuan Li, Osamah Sufyan, Uwe Jaekel, Babette Dellen

    cs.CV · cs.LG

    Many dense prediction networks rely on additive feature transformations and model higher-order feature interactions only implicitly. Product units provide an explicit mechanism for multiplicative feature modeling, but their logarithmic--exponential formulation can cause numerical instability, which has limited their use in deep dense prediction networks. In this work, we propose Product-Unit U-Net (PU-UNet), a residual U-Net that integrates...

    arxiv.org/abs/2606.20035 · PDF

  14. 14

    Semantic-Anchored Evidential Fusion for Domain-Robust Whole-Slide Survival Analysis

    Yucheng Xing, Ling Huang, Pei Liu, Jingying Ma, Jiaqing Xu, Kai He, Mengling Feng

    cs.CV · cs.LG

    Whole-slide images (WSIs) are widely used for computational cancer prognosis. However, most existing methods primarily focus on in-domain performance and fail to generalize across clinical centers. This limitation stems from their reliance on pixel-derived representations that are highly susceptible to domain-specific artifacts caused by staining protocols and scanner hardware. We hypothesize that high-level pathology semantics, such as tumor...

    arxiv.org/abs/2606.19966 · PDF

  15. 15

    ROSE: Benchmarking the Perception-to-Action Gap in Multimodal Models

    Yihao Wang, Zijian He, Jie Ren, Keze Wang

    cs.CV · cs.AI

    Multimodal large language models (MLLMs) are increasingly expected to act on visual information, yet the same scene may require different actions under different task contexts. How reliably can a model turn the same visual evidence into the action required by the current context? To answer this question, we introduce \textsc{ROSE} (\textbf{R}eference-conditioned \textbf{O}ddity and \textbf{S}ymbolic \textbf{E}xecution), a controlled benchmark...

    arxiv.org/abs/2606.19965 · PDF

  16. 16

    Confidence Calibration for Multimodal LLMs: An Empirical Study through Medical VQA

    Yuetian Du, Yucheng Wang, Ming Kong, Tian Liang, Qiang Long, Bingdi Chen, Qiang Zhu

    cs.CV · cs.AI

    Multimodal Large Language Models (MLLMs) show great potential in medical tasks, but their elicited confidence often misaligns with actual accuracy, potentially leading to misdiagnosis or overlooking correct advice. This study presents the first comprehensive analysis of the relationship between accuracy and confidence in medical MLLMs. It proposes a novel method that combines Multi-Strategy Fusion-Based Interrogation (MS-FBI) with auxiliary...

    arxiv.org/abs/2606.19950 · PDF

  17. 17

    Triangular Consistency as a Universal Constraint for Learning Optical Flow

    Yi Xiao, Carlos Rodriguez Coronel, Jing Zhan, Haniyeh Ehsani Oskouie, Alex Wong, Dong Lao

    cs.CV · cs.AI

    We propose triangular consistency as a first-principled constraint for optical flow, which is agnostic to network architecture, supervision type, and dataset, and applies to both image-pair and multi-frame settings. This simple but powerful constraint is to compose two flows to induce a third flow and enforce consistency among the three. The composed flows may arise from (i) image pairs, yielding cycle consistency; (ii) multiple video frames,...

    arxiv.org/abs/2606.19938 · PDF

  18. 18

    Speeding up the annotation process in semantic segmentation industrial applications

    Marta Fernandez-Moreno, Margarita Guerrero, Rosalia Rementeria, Pablo Mesejo, Raul Moreno

    cs.CV · cs.AI

    Current machine learning models commonly require large and well-annotated datasets. However, the annotation process often becomes a bottleneck, with increased complexity leading to higher chances of human errors. Within this context, our goal in this paper is to leverage unsupervised algorithms to improve data annotation efficiency for complex semantic segmentation problems in industrial materials science. Previous research has quantified...

    arxiv.org/abs/2606.19934 · PDF

  19. 19

    Spatial-Aware Reduction Framework: Towards Efficient and Faithful Visual State Space Models

    Jindi Lv, Aoyu Li, Yuhao Zhou, Zheng Zhu, Xiaofeng Wang, Qing Ye, Yueqi Duan, Wentao Feng, Jiancheng Lv

    cs.CV · cs.AI

    Mamba demonstrates strong efficiency in modeling long visual sequences. However, when token reduction is applied to structurally enhanced Mamba variants, these models exhibit a severe performance collapse. We attribute this degradation to the spatially agnostic nature of existing reduction methods, which violate the two-dimensional structural premise required by the selective scanning mechanism. In this work, we propose STORM, a spatial-aware...

    arxiv.org/abs/2606.19932 · PDF

  20. 20

    Multimodal Concept Bottleneck Models

    Tongqing Shi, Ge Yan, Tuomas Oikarinen, Tsui-Wei Weng

    cs.CV · cs.LG

    Concept Bottleneck Models (CBMs) enhance the interpretability of deep learning networks by aligning the features extracted from images with natural concepts. However, existing CBMs are constrained in their ability to generalize beyond a fixed set of predefined classes and the risk of non-concept information leakage, where predictive signals outside the intended concepts are inadvertently exploited. In this paper, we propose Multimodal Concept...

    arxiv.org/abs/2606.19882 · PDF

  21. 21

    PSCT-Net: Geometry-Aware Pediatric Skull CT Reconstruction via Differentiable Back-Projection and Attention-Guided Refinement

    Dong Yeong Kim, Jaewon Choi, Youmin Shin, Jungyu Lee, Myeongseop Kim, Jinwook Choi, Joo Whan Kim, Young-Gon Kim

    cs.CV · cs.AI

    Computed Tomography (CT) is essential for diagnosing pediatric craniofacial abnormalities, yet poses radiation risks to developing anatomies. Reconstructing 3D CT from sparse bi-planar X-rays offers a low-dose alternative but is severely ill-posed. Existing methods employ geometry-agnostic feature lifting, naively projecting 2D features into 3D without explicit spatial modeling, causing depth ambiguity and degraded osseous boundaries. We...

    arxiv.org/abs/2606.19867 · PDF

  22. 22

    CSWinUNETR: Segmentation of Thin Anatomical Structures in Medical Images

    Junho Moon, Haejun Chung, Ikbeom Jang

    cs.CV · cs.AI

    Accurate segmentation of thin, tortuous anatomical structures, such as retinal vessels, cerebral vasculature, and facial wrinkles, remains challenging due to low contrast, frequent discontinuities, and severe class imbalance. Although recent convolutional and Transformer-based models have improved performance, they often yield fragmented predictions and fail to recover fine branches. We propose CSWinUNETR, a general-purpose backbone for 2D...

    arxiv.org/abs/2606.19824 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.