cs.CV · 2026-06-13 · No. 22

Computer Vision and Pattern Recognition, 2026-06-13.

20 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

20 entries
  1. 01

    SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning

    Seokju Cho, Ryo Hachiuma, Abhishek Badki, Hang Su, Byung-Kwan Lee, Chan Hee Song, Sifei Liu, Subhashree...

    cs.CV · cs.AI

    Spatial reasoning, the ability to determine where objects are, how they relate, and how they move in 3D, remains a fundamental challenge for vision-language models (VLMs). Tool-augmented agents attempt to address this by augmenting VLMs with specialist perception modules, yet their effectiveness is bounded by the action interface through which those tools are invoked. In this work, we study how the design of this interface shapes the agent's...

    arxiv.org/abs/2606.13673 · PDF

  2. 02

    EvTexture++: Event-Driven Texture Enhancement for Video Super-Resolution

    Dachun Kai, Jiayao Lu, Yueyi Zhang, Xiaoyan Sun

    cs.CV · cs.AI

    Event-based vision has drawn increasing attention owing to its distinctive properties, including ultra-high temporal resolution and extreme dynamic range. Recent works have introduced it to video super-resolution (VSR) to enhance flow estimation and temporal alignment. In contrast, this paper shifts the focus of event signals from motion refinement to texture enhancement in VSR. We propose EvTexture++, the first event-driven framework...

    arxiv.org/abs/2606.13580 · PDF

  3. 03

    Contrast-Informed Augmentation and Domain-Adversarial Training for Adult-to-Neonatal MR Reconstruction Generalization

    Stephen Moore, Lara Leijser, Richard Frayne, Roberto Souza

    cs.CV · cs.AI

    Purpose: To investigate whether contrast-informed data augmentation and domain-adversarial training improve the adult-to-neonatal generalization of the E2E-VarNet. Methods: Three training regimes were investigated: (1) adult-only training with unaugmented adult data, (2) mixed training with paired unaugmented and neonatal-informed augmented adult data, and (3) mixed training with a domain-adversarial objective. Models were trained on...

    arxiv.org/abs/2606.13562 · PDF

  4. 04

    MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models

    Hanyang Yu, Haitao Lin, Jingbo Zhang, Wenyao Zhang, Chenghao Gu, Heng Li, Ping Tan

    cs.CV · cs.LG · cs.RO

    World Action Models (WAMs) present a promising paradigm for robotic control via video prediction. However, current WAMs suffer from fundamental spatial bottlenecks: standard text inputs introduce referential ambiguity in cluttered scenes, while unstructured RGB predictions lack semantic grounding and remain biased by task-irrelevant backgrounds. To overcome these limitations, we introduce MaskWAM, an object-centric world-action model. By...

    arxiv.org/abs/2606.13515 · PDF

  5. 05

    Measurement-Calibrated Multi-Camera Fusion for Vision-Based Indoor Localization

    Mateo Toro Diz, Jonathan Hoss, Noah Klarmann

    cs.CV · cs.AI

    Indoor vision-based localization systems are affected by detection noise, occlusions, and limited camera coverage, leading to uncertainty at multiple stages of the pipeline. While multi-camera data fusion is widely used to mitigate these issues, it is typically treated as a black-box component and evaluated solely end-to-end, obscuring its mechanistic contributions. To address this gap, this work investigates whether explicitly characterizing...

    arxiv.org/abs/2606.13509 · PDF

  6. 06

    Heterogeneous LiDAR Early Fusion and Learned Re-Ranking Strategy for Robust Long-Term Place Recognition in Unstructured Environments

    Judith Vilella-Cantos, Juan José Cabrera, Mónica Ballesta, David Valiente, Luis Payá

    cs.CV · cs.AI · cs.RO

    Robust localization in unstructured environments, such as agricultural fields, is a critical challenge for autonomous systems. LiDAR sensors provide detailed 3D information about the environment and are invariant to lighting conditions. For this reason, LiDAR-based place recognition methods have gained significant attention. In this paper, we propose MinkUNeXt-VINE++, a novel approach that combines early fusion of heterogeneous LiDAR data...

    arxiv.org/abs/2606.13503 · PDF

  7. 07

    OmniDirector: General Multi-Shot Camera Cloning without Cross-Paired Data

    Jiwen Liu, Shujuan Li, Zhixue Fang, Xiaohan Li, Yan Zhou, Zijie Meng, Zhimin Zhang, Yawen Luo, Guoxin Zhang, Yu-Shen...

    cs.CV · cs.AI

    Cloning camera motion from reference videos is an important task in video generation, as videos provide intuitive and precise control. Existing methods either directly use parametric representations that fail to handle multi-shot generation or synthesize cross-paired data, which suffer from data scarcity, resulting in poor performance in complicated camera motion cloning. To address these issues, we introduce a general camera motion...

    arxiv.org/abs/2606.13432 · PDF

  8. 08

    SmartFont: Dynamic Condition Allocation for Few-Shot Font Generation

    Zian Yang, Zixin Wang

    cs.CV · cs.AI

    Few-shot font generation simultaneously requires global structural completeness and fine-grained local style fidelity. Existing methods usually either rely on global content-style modeling, which is robust but imperfectly disentangled, or emphasize component/local modeling, which captures fine details but relies heavily on local priors and reference coverage. We argue that the key challenge is not merely to learn purer conditions, but to...

    arxiv.org/abs/2606.13382 · PDF

  9. 09

    Dual-Domain Equivariant Generative Adversarial Network for Multimodal CT-PET Synthesis

    Gabriel Steele, Alzahra Altalib, Alessandro Perelli

    cs.CV · cs.AI · physics.med-ph

    We present a Dual-Domain Equivariant Generative Adversarial Network (DDE-GAN) for multimodal CT-PET image synthesis. Traditional GAN-based approaches often operate solely in the spatial domain and ignore geometric consistency, resulting in limited structural fidelity. DDE-GAN addresses these challenges by jointly learning from both spatial and frequency (Fourier) domains, capturing complementary anatomical and spectral information....

    arxiv.org/abs/2606.13341 · PDF

  10. 10

    HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers

    Guozhen Zhang, Xuerui Qiu, Yutao Cui, Tianhui Song, Changlin Li, Junzhe Li, Tao Huang, Xiao Zhang, Yang Li, Jianbing...

    cs.CV · cs.AI

    Holistic visual tokenizers are fundamental to unified multimodal models (UMMs) as they map diverse visual inputs into a unified representation space. In this paper, we present HYDRA-X, the first UMM that unifies image and video tokenization within a single Vision Transformer (ViT). Our design is driven by two core challenges: efficiently injecting spatiotemporal reconstruction capability into a native ViT, and embedding image- and video-level...

    arxiv.org/abs/2606.13289 · PDF

  11. 11

    Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality

    Wei Li, Zhen Huang, Xinmei Tian

    cs.CV · cs.AI · cs.CL

    Contrastively trained vision-language models like CLIP, have made remarkable progress in learning joint image-text representations, but still face challenges in compositional understanding. They often exhibit a "bag-of-words" behavior--struggling to capture the object relations, attribute-object bindings, and word order dependencies. This limitation arises not only from the reliance on global, single-vector representations for optimization,...

    arxiv.org/abs/2606.13288 · PDF

  12. 12

    Transformer-Guided Graph Attention for Direct Cardiac Mesh Reconstruction: A Structural Digital Twin Framework

    Abhishek H S, Akash Ganamukhi, Abhimanyu Suresh, Aditya G Hiremath, Prasad B Honnavalli, Adithya Balasubramanyam

    cs.CV · cs.AI

    Building patient-specific cardiac models sits at the heart of precision cardiology, yet getting those models into clinical use keeps running into the same wall: mesh generation is slow, messy, and frustrating. The standard workflow -- segmenting the image, running Marching Cubes, and then manually cleaning up the result -- is time-consuming, inconsistent across operators, and demands specialist knowledge most clinical teams do not have. We...

    arxiv.org/abs/2606.13188 · PDF

  13. 13

    Iterative Visual Thinking: Teaching Vision-Language Models Spatial Self-Correction through Visual Feedback

    Animesh Tripathy, Aswanth Krishnan

    cs.CV · cs.AI

    Vision-language models (VLMs) achieve strong singleshot spatial grounding, yet lack any mechanism to observe and correct their own predictions. We find that naively prompting a VLM to iterate over rendered visualizations of its predictions causes catastrophic failure: Acc@0.5 on referring expression comprehension collapses from 79.6% to 48.7% (a 31 percentage point drop), revealing a fundamental gap between grounding capability and...

    arxiv.org/abs/2606.13156 · PDF

  14. 14

    An Extensible and Lightweight Unified Architecture for Demosaicing Pixel-bin Image Sensors

    Saurabh Kumar, Nutan Sairam Yenneti

    cs.CV · cs.LG · eess.IV

    Pixel-bin image sensors are becoming the default choice for smartphone cameras due to their resolution vs light-gathering trade-off. However, their larger inter-color separation compared to the Bayer color filter array (CFA) makes them challenging to demosaic. Furthermore, existing deep learning-based demosaicing methods are CFA-specific, requiring multiple individual models that take up precious onboard resources and demand larger...

    arxiv.org/abs/2606.13136 · PDF

  15. 15

    Cascade Classification of Dermoscopic Images of Skin Neoplasms with Controllable Sensitivity and External Clinical Validation

    Elena S. Kozachok, Sergey S. Seregin, Aleksandr V. Kozachok, Ilya P. Latyshev, Oleg I. Samovarov

    cs.CV · cs.AI

    Purpose. To compare deep learning architectures and classification schemes for dermoscopic images of skin neoplasms and assess their generalization on transfer from open international datasets to independent clinical datasets of Russian practice. Methods. Four architectures (ViT-B/16, Swin-S, ConvNeXt-S, EfficientNetV2-S) were compared in three schemes: binary (malignant/benign), single-stage four-class (benign, MEL, SCC, BCC), and a...

    arxiv.org/abs/2606.13135 · PDF

  16. 16

    TetherCache: Stabilizing Autoregressive Long-Form Video Generation with Gated Recall and Trusted Alignment

    Yu Meng, Xiangyang Luo, Letian Li, Wenyuan Jiang, Chen Gao, Xinlei Chen, Yong Li, Xiao-Ping Zhang

    cs.CV · cs.AI

    Autoregressive video diffusion models provide a natural formulation for streaming and variable-length video generation by conditioning newly generated frames on previously generated content. However, extending these models to minute-level generation remains challenging: the limited KV-cache budget prevents the model from retaining the full history, while repeatedly conditioning on self-generated frames induces a context distribution shift...

    arxiv.org/abs/2606.13035 · PDF

  17. 17

    Quality-Preserving Imperceptible Adversarial Attack on Skeleton-based Human Action Recognition

    Ziyi Chang, Kanglei Zhou, Xiaohui Liang, Hubert P. H. Shum

    cs.CV · cs.LG

    Adversarial attacks on skeletal human action recognition have received significant attention. However, existing methods typically introduce noise-like perturbations that degrade motion quality post-attack, and thereby are inherently perceptible with recent advancements in S-HAR systems. We discover that this degradation stems from the gap between empirical and true risks during the optimization process of previous adversarial attacks. To...

    arxiv.org/abs/2606.13022 · PDF

  18. 18

    A Machine Learning Framework for Real-Time Personalized Ergonomic Pose Analysis

    Manex Atxa, Bruno Simoes, Julen Balzategui

    cs.CV · cs.AI

    This paper introduces a new methodology for real-time prediction of ergonomic and non-ergonomic human poses using volumetric video data in three dimensions. Although the methodology was designed for ergonomic assessments, it can be adapted to other applications requiring real-time analysis of human posture. One aspect that makes this system stand out is its ability to analyze 3D point clouds during the assessment, enabling computation from...

    arxiv.org/abs/2606.12988 · PDF

  19. 19

    Diffusion Transformer World-Action Model for AV Scene Prediction

    Ruslan Sharifullin, Benjamin Jiang, Kai Xi Chew

    cs.CV · cs.AI · cs.LG · cs.RO

    Action-conditioned world models let an autonomous vehicle predict future camera scenes from its own planned controls, enabling planning and simulation without real-world rollouts, but at compact, trainable scale the futures are ambiguous and the field's standard distortion metrics actively mislead: they reward a blurry regression mean over a realistic prediction. We confront this with a compact latent world model that, given the present...

    arxiv.org/abs/2606.12987 · PDF

  20. 20

    Efficient, Robust, and Anti-Collusion Fingerprinting of Image Diffusion Models

    Jianwei Fei, Yunshu Dai, Zhihua Xia, Xiaochun Cao, Jiantao Zhou, Alessandro Piva, Benedetta Tondi

    cs.CV · cs.AI · cs.CR · cs.LG

    Model fingerprinting, embedding user-specific identifiers (fingerprints) into generated outputs, has recently emerged as a popular solution to protect the intellectual property rights (IPR) of generative text-to-image (T2I) models and prevent unauthorized redistribution. In this work, we reveal a previously unexplored systematic vulnerability in existing generative model fingerprinting methods: they lack robustness against collusion attacks,...

    arxiv.org/abs/2606.12977 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.