cs.CV · 2026-06-17 · No. 26

Computer Vision and Pattern Recognition, 2026-06-17.

25 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

25 entries
  1. 01

    Adaptive Volumetric Mechanical Property Fields Invariant to Resolution

    Rishit Dagli, Donglai Xiang, Vismay Modi, Xuning Yang, Gavriel State, David I. W. Levin, Maria Shugrina

    cs.CV · cs.LG · cs.RO

    Accurate mechanical properties (or materials) Young's modulus ($E$), Poisson's ratio ($ν$) and density ($ρ$) are essential for reliable physics simulation of digital worlds, but most 3D assets lack this information. We propose AdaVoMP, a method for predicting accurate dense spatially-varying ($E$, $ν$, $ρ$) for input 3D objects across representations, improving the resolution, accuracy, and memory efficiency over the state-of-the-art. The...

    arxiv.org/abs/2606.18231 · PDF

  2. 02

    ReAge3D: Re-Aging 3D Faces with View Consistency

    Libing Zeng, Li Ma, Mingming He, Ning Yu, Paul Debevec, Nima Khademi Kalantari

    cs.CV · cs.AI

    We present a novel framework for realistic and controllable 3D face re-aging which produces highly detailed, identity-preserving results. Existing 3D editing methods, while effective for coarse semantic changes, are not well suited for re-aging, as even small inconsistencies across re-aged 2D views can lead to over-smoothing of subtle but perceptually important age-related details. To address this challenge, we first introduce a 2D...

    arxiv.org/abs/2606.18156 · PDF

  3. 03

    When LLMs Analyze Scars: From Images to Clinically-Meaningful Features

    Ruman Wang, Hangting Ye

    cs.CV · cs.AI · cs.LG

    Medical image classification faces a fundamental dilemma: while deep learning models achieve remarkable performance at scale, real-world clinical scenarios often suffer from severe data scarcity due to annotation costs, privacy constraints, and disease rarity. This challenge is particularly pronounced in pathological scar classification, where differentiating keloids from hypertrophic scars requires subtle expert knowledge and labeled images...

    arxiv.org/abs/2606.18063 · PDF

  4. 04

    Recover Semantics First, Generate Better: Improved Latent Modeling for 3D MRI Reconstruction and Cross-Contrast Synthesis

    Yonghao Chen, Sicheng Yang, Rui Tang, Lei Zhu

    cs.CV · cs.AI

    Multi-contrast magnetic resonance imaging (MRI) provides complementary information for clinical diagnosis. However, acquiring all MRI sequences is often time-consuming and costly. Recent generative models perform cross-contrast synthesis to address this issue by inferring absent contrasts from the available ones. Nevertheless, synthesizing 3D MRI presents significant challenges. Due to the massive volume sizes, operating directly in the pixel...

    arxiv.org/abs/2606.17989 · PDF

  5. 05

    SegDINO: Introducing Multi-Scale Structure into DINO for Efficient Medical Image Segmentation

    Sicheng Yang, Hongqiu Wang, Zhaohu Xing, Sixiang Chen, Qiuxia Yang, Yize Mao, Guang Yang, Lei Zhu

    cs.CV · cs.AI

    Self-supervised DINO models provide strong transferable visual representations, yet applying them directly to image segmentation remains challenging. Existing approaches commonly rely on heavy decoders with complex upsampling, introducing substantial parameter and computational overhead. We observe that introducing scale into DINO features is far more critical than increasing decoder capacity. In this work, we present SegDINO, an efficient...

    arxiv.org/abs/2606.17972 · PDF

  6. 06

    Robustness of Similarity-based Positional Encoding Under Rotations: Theoretical Analysis and Experimental Validation

    Andrea Santomauro, Luigi Portinale, Giorgio Leonardi

    cs.CV · cs.AI

    Positional encoding is a fundamental component of Transformer architectures, as it injects information about the spatial or sequential arrangement of inputs. Among recent alternatives to standard absolute and sinusoidal encodings, similarity-based positional encoding (simPE) has emerged as a flexible framework for representing positional structure through pairwise relations. simPE was originally designed for medical imaging applications,...

    arxiv.org/abs/2606.17961 · PDF

  7. 07

    Beyond Visual Cues: CoT-Enhanced Reasoning for Semi-supervised Medical Image Segmentation

    Yuming Chen, Yuxin Xie, Tao Zhou, Yi Zhou

    cs.CV · cs.LG

    Semi-supervised medical image segmentation has emerged as a dominant research problem in medical image analysis, mitigating annotation scarcity by leveraging consistency regularization on unlabeled data. However, existing approaches operate predominantly via visual pattern matching, relying heavily on pixel-level similarities. This visual-centric dependency often falters in clinical scenarios characterized by the visual-semantic mismatch,...

    arxiv.org/abs/2606.17958 · PDF

  8. 08

    Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model

    Jinghan Wu, Jing Li, Ivor W. Tsang, Xuetao Zhang

    cs.CV · cs.AI

    Visual information helps resolve ambiguity in coreference resolution, leading to notable performance gains. However, existing Multi-modal Coreference Resolution (MCR) methods require training with (partially) annotated data from the target dataset before they can be applied, preventing their direct usability and raising concerns about generalization. While Vision-Language Large Models (VLLMs) with billions of parameters offer promising...

    arxiv.org/abs/2606.17950 · PDF

  9. 09

    Revisiting Structural Dependency in Autoregressive Multi-Task Table Recognition via Order-Independent Cell-Level Representations

    Takaya Kawakatsu

    cs.CV · cs.LG

    Multi-task table recognition jointly addresses table structure prediction, cell localization, and cell content recognition within a unified framework. Existing approaches often rely on autoregressive decoders to generate table structures and reuse their hidden states for cell localization and content recognition. This autoregressive generation process can make cell representations order-dependent, degrading global consistency across cells....

    arxiv.org/abs/2606.17874 · PDF

  10. 10

    A Quantitative Analysis of Multimodal Biomarkers in Alzheimer's Disease

    Antonio Scardace, Daniele Ravì

    cs.CV · cs.AI

    Despite increasing adoption of multimodal approaches in Alzheimer's Disease (AD) research -- aimed at integrating molecular, structural, clinical, and genetic biomarkers to enhance disease characterization -- the relationships among these modalities remain poorly understood. A systematic analysis of their dynamic interaction is essential for improving disease modeling, identifying redundant assessments, and reducing patient burden and...

    arxiv.org/abs/2606.17867 · PDF

  11. 11

    High-Fidelity 3D Geometric Reconstruction of Pelvic Organs from MRI: A Hybrid Deep Learning and Iterative Optimization Approach

    Hui Wang, Xiaowei Li, Chenxin Zhang, Yifan Feng, Jianwei Zuo, Yumeng Tang, Xiuli Sun, Jianliu Wang, Bing Xie, Jiajia Luo

    cs.CV · cs.AI · cs.CG · cs.GR

    Patient-specific 3D reconstruction of pelvic organ geometry from MRI is important for pelvic floor modeling and downstream patient-specific analysis. However, while previous studies have focused primarily on either image segmentation or downstream use of 3D models, the reconstruction of high-fidelity, high-quality geometries remains labor-intensive and poorly standardized. The study introduced a hybrid deformable shape modeling framework that...

    arxiv.org/abs/2606.17836 · PDF

  12. 12

    Human-in-the-Loop Atlas-Based 3D Asset Segmentation for Interactive Content Workflows

    Paul Julius Kühn, Saptarshi Neil Sinha, Jakob Hansen, Robin Horst

    cs.CV · cs.AI

    Segmenting 3D assets into meaningful regions remains challenging, especially when segmentation criteria are application-dependent and require user control. We present a human-in-the-loop pipeline for generating a segmented 2D parameterized atlas from a 3D model for interactive media, game, and XR content workflows. Our method first selects a compact set of rendered views using a greedy set cover strategy over sampled surface points, and then...

    arxiv.org/abs/2606.17824 · PDF

  13. 13

    LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams

    Zhenyu Yang, Kairui Zhang, Bing Wang, Shengsheng Qian, Changsheng Xu

    cs.CV · cs.AI

    Despite the remarkable progress of Video Large Language Models (Video-LLMs), current online architectures still struggle to simultaneously process continuous video streams, decide autonomously when to respond, and preserve long-horizon contextual memory. These obstacles undermine real-time responsiveness and cause severe forgetting throughout prolonged interactions. In this work, we introduce LiveStarPro, a live streaming assistant that is...

    arxiv.org/abs/2606.17798 · PDF

  14. 14

    Structured Adversarial Camouflage via Voronoi Diagrams

    Jens Bayer, Stefan Becker, David Münch, Michael Arens, Jürgen Beyerer

    cs.CV · cs.AI

    Pixel-wise adversarial patches are computationally heavy and often visually detectable, limiting utility in security-critical systems. We present adversarial Voronoi camouflage that optimizes only seed-point locations under fixed, printable palettes using a soft assignment, producing structured, splinter camouflage-like patterns without additional regularization. Evaluated on person detection with COCO-style AP@[.5:.95], naive placement...

    arxiv.org/abs/2606.17711 · PDF

  15. 15

    Vision-language models for chest radiography do not always need the image

    Mahshad Lotfinia, Sebastian Ziegelmayer, Lisa Adams, Daniel Truhn, Andreas Maier, Soroosh Tayebi Arasteh

    cs.CV · cs.AI · cs.CL · cs.LG

    Medical vision-language models report strong chest radiograph accuracy, and this is increasingly read as evidence that they use the image. That inference is unsafe: a model exploiting finding-name priors scores like one that reads the scan, and no standard benchmark separates them. We introduce a causal audit that intervenes on the image, occluding the relevant region, occluding an irrelevant one, and swapping in another patient's same-label...

    arxiv.org/abs/2606.17710 · PDF

  16. 16

    SegTME-UNI2: A Foundation Model-Based Framework for Generalisable Multiclass Cell Segmentation and LLM-Driven Tumour Microenvironment Characterisation in Histopathology

    Wan Siti Halimatul Munirah Wan Ahmad, Faris Syahmi Samidi, Mohammad Badal Ahmmed, Vimal Angela Thiviyanathan, Selvam...

    cs.CV · cs.AI

    Characterising the tumour microenvironment (TME) from routine H&E-stained histology images requires simultaneous cell segmentation, feature extraction, and interpretable clinical reporting. We present SEGTME-UNI2, a unified framework addressing these requirements. Its core is UNI2-UPERHOVER, a dual-head segmentation model pairing the UNI2-H pathology foundation model (ViT-Giant, pretrained on >100M tiles from 100K slides) with two parallel...

    arxiv.org/abs/2606.17702 · PDF

  17. 17

    See First, Answer Later: Visual Evidence Pre-Alignment via Sufficiency-Driven RL

    Yilian Liu, Sicong Leng, Guoshun Nan, Junyi Zhu, Jiayu Huang, Minghao Sun, Xuancheng Zhu, Yisong Chen, Zexian Wei,...

    cs.CV · cs.AI

    Multimodal large language models (MLLMs) integrate strong text reasoning with visual inputs, yet their responses can be inconsistent with the underlying images, indicating ineffective utilization of visual evidence during inference. The prevailing training paradigm relies on large-scale caption-based pretraining for general alignment, followed by supervised fine-tuning and reinforcement learning to enable instruction following and complex...

    arxiv.org/abs/2606.17678 · PDF

  18. 18

    Bounding Box Label Propagation for Re-Annotation of Document Layout Analysis Datasets

    Nick Jochum, Tobias Alt-Veit, Christian Schön, Alexander Lück, René Schuster, Didier Stricker

    cs.CV · cs.AI

    Datasets in practical document processing scenarios typically grow over time, and their class annotations undergo continuous refinement. This creates significant re-annotation efforts, which are time-consuming and costly. A promising remedy is to re-annotate only a small subset of available documents manually and apply semi-supervised learning techniques that leverage both labelled and unlabelled data. Although there are numerous approaches...

    arxiv.org/abs/2606.17644 · PDF

  19. 19

    Divide, Deliberate, Decide: A Multi-Agent Framework for Fine-Grained Egocentric Action Recognition

    Alessandro Sottovia, Alessandro Torcinovich, Oswald Lanz

    cs.CV · cs.AI

    Fine-grained action recognition in egocentric video is challenging for Vision-Language Models (VLMs): actions often differ only in small visual cues, and a single model tends to be biased toward a subset of these cues. We propose Divide, Deliberate, Decide, a fully-local, zero-shot multi-agent framework in which (i) a VLM orchestrator chunks the video and proposes a top-k candidate label list per segment, (ii) an ensemble of heterogeneous VLM...

    arxiv.org/abs/2606.17627 · PDF

  20. 20

    SkillMoV: Mixture-of-View Routing with Prototype-Conditioned Gating for Unified Multi-View Proficiency Estimation

    Edoardo Bianchi, Antonio Liotta

    cs.CV · cs.AI

    Estimating human proficiency from video is a key challenge for automated skill assessment, with applications in sports coaching, music pedagogy, surgical training, and workplace learning. Existing approaches often focus on individual scenarios or rely on shared multi-view aggregation, limiting their ability to adapt to heterogeneous camera viewpoints and activity domains. We introduce SkillMoV, a unified, parameter-efficient framework for...

    arxiv.org/abs/2606.17615 · PDF

  21. 21

    Root-Selecting Fixed-Point Inversion for Rectified Flows via Trajectory Straightness

    Semin Kim, Jihwan Yoon, Seunghoon Hong

    cs.CV · cs.LG

    Finding the initial noise that generates a given data sample, known as inversion, is a key component for downstream applications such as training-free image editing. Existing fixed-point inversion methods improve inversion accuracy by formulating each inversion step as a fixed-point problem, but they lack a principled mechanism for selecting among multiple fixed-point solutions that can arise in practice. We observe that different selections...

    arxiv.org/abs/2606.17584 · PDF

  22. 22

    Geometric Consistency Protocol for Foundation Model Features in Multi-View Satellite Imagery

    Qiyan Luo, Jie Yang, Yingdong Pi, Lekang Wen, Mi Wang

    cs.CV · cs.AI

    Standardized evaluation protocols are indispensable for robust benchmarking in remote sensing, particularly as foundation features are increasingly transferred across diverse sensors and complex imaging geometries. In satellite multi-view reconstruction, conventional evaluations relying on unconstrained 2D global matching are often misleading. The Rational Function Model (RFM) and its Rational Polynomial Coefficients (RPC) dictate a curved,...

    arxiv.org/abs/2606.17564 · PDF

  23. 23

    Reinforcing Dual-Path Reasoning in Spatial Vision Language Models

    Yatai Ji, An-Chieh Cheng, Yang Fu, Yukang Chen, Han Zhang, Zhaojing Yang, Wei Huang, Ka Chun Cheung, Song Han, Vidya...

    cs.CV · cs.AI

    Spatial VLMs have made substantial progress in geometric perception, yet complex spatial reasoning requiring multi-step inference over depth, distance, and scene relations remains challenging. Moreover, different spatial queries call for fundamentally different strategies: some are best addressed through purely linguistic, step-by-step deduction, while others require explicit 3D grounding before quantitative inference. We present Dual-Path...

    arxiv.org/abs/2606.17539 · PDF

  24. 24

    OmniDrive: An LLM-Choreographed Multi-Agent World Model with Unified Latent Co-Compression for Multi-View Driving Video Generation

    Zijie Meng, Yufei Liu, Chengqian Ma, Zhiyu Li, Jiyuan Liu, Wenhua Nie, Bingcai Wei, Shuqin Chen, Weichen Xu, Jiquan...

    cs.CV · cs.AI

    Generative world models for autonomous driving face two unresolved tensions: heterogeneous control injection, where free-form language, HD-maps, trajectories, and camera poses reside in incompatible representational spaces, and post-hoc cross-view fusion, where per-camera latents fail to encode global 3-D geometry. We trace both to a single root cause: the absence of a shared symbolic interlingua aligning language, geometry, and pixels at the...

    arxiv.org/abs/2606.17536 · PDF

  25. 25

    Theoretical Grounding of Out-Of-Distribution Detection With Reinforcement Learning Optimizer

    Salimeh Sekeh, Xin Zhang

    cs.CV · cs.LG

    Out-of-distribution (OOD) detection in dynamic open-world environments requires a model to continually adapt to evolving data distributions while generalizing to covariate-shifted inputs and rejecting semantic-shifted OOD examples. Most existing OOD detection methods optimize only the current-step objective and do not explicitly account for how post-deployment environment changes affect future OOD behavior. In this paper, we establish a...

    arxiv.org/abs/2606.17477 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.