cs.CV · 2026-05-25 · No. 9

Computer Vision and Pattern Recognition, 2026-05-25.

29 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

29 entries
  1. 01

    ETCHR: Editing To Clarify and Harness Reasoning

    Beichen Zhang, Yuhong Liu, Jinsong Li, Yuhang Zang, Jiaqi Wang, Dahua Lin

    cs.CV · cs.AI · cs.CL

    Multimodal Large Language Models have advanced visual reasoning, yet a purely textual chain of thought remains a bottleneck for questions that require fine-grained focus or view transformations. The ''think with images'' paradigm narrows this gap, but existing approaches are either constrained by fixed predefined toolkits or produce noisy intermediate images from unified multimodal methods. We pursue a third option: using a dedicated image...

    arxiv.org/abs/2605.23897 · PDF

  2. 02

    Good Token Hunting: A Hitchhiker's Guide to Token Selection for Visual Geometry Transformers

    Shuhong Zheng, Michael Oechsle, Erik Sandström, Marie-Julie Rakotosaona, Federico Tombari, Igor Gilitschenski

    cs.CV · cs.AI · cs.GR · cs.LG · cs.RO

    Visual geometry transformers have become powerful architectures for multi-view 3D reconstruction, enabling joint prediction of multiple 3D attributes in a feed-forward manner. However, their computational cost grows quadratically with the input sequence length due to the global attention layers inside these models. This limits both their scalability and efficiency. In this work, we address this challenge with a simple yet general strategy:...

    arxiv.org/abs/2605.23892 · PDF

  3. 03

    PGT: Procedurally Generated Tasks for improving visual grounding in MLLMs

    Rim Assouel, Amir Bar, Michal Drozdzal, Adriana Romero-Soriano

    cs.CV · cs.AI

    Despite remarkable progress in Multimodal Large Language Models (MLLMs), these models still struggle with fine-grained understanding tasks. In this work, we propose Procedurally Generated Tasks (PGT), a simple data-driven framework that serves a dual purpose: inducing fine-grained visual understanding and acting as a low-cost diagnostic tool to identify the source of perception failures. By overlaying unambiguous geometric primitives on...

    arxiv.org/abs/2605.23883 · PDF

  4. 04

    Not Too Generative, Not Too Discriminative: The Human Alignment Sweet Spot

    Jorge Chang Ortega, Bastien Le Lan, Thomas Serre, Victor Boutin

    cs.CV · cs.AI

    A central question in computational vision is whether human-like visual representations are better explained by discriminative or generative learning. Existing comparisons, however, often confound the learning objective with architecture, scale, and training data, leaving open whether the objective itself drives alignment. We address this confound using Joint Energy-Based Models (JEMs), which interpolate continuously between discriminative...

    arxiv.org/abs/2605.23819 · PDF

  5. 05

    PhotoFlow: Agentic 3D Virtual Photography Missions

    Jiarui Guo, Haojia Wei, Yiming Zhang, Yifei Liu, Yuning Gong, Hongjie Zhang, Xue Yang, Zhihang Zhong

    cs.CV · cs.AI · cs.MA

    Virtual photography asks an agent to enter a prepared 3D scene with no preselected camera pose or reference image, infer a suitable shot from scene information and a language intent, choose executable camera parameters, and render the final photograph. Recent progress in vision-language models makes this kind of spatial agent increasingly plausible, but the task stresses two capabilities that remain hard to evaluate together: complex 3D...

    arxiv.org/abs/2605.23771 · PDF

  6. 06

    CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception

    Liupeng Li, Haoqian Kang, Zhenyu Lu, Jinpeng Wang, Bin Chen, Ke Chen, Yaowei Wang

    cs.CV · cs.AI · cs.LG · cs.MM

    High-resolution (HR) image perception presents a key bottleneck for multimodal large language models (MLLMs). While visual search offers a promising solution, existing methods struggle with the trade-off between coverage and efficiency. Visual expert-assisted search is efficient but prone to blind spots when proposals fail, whereas scan-based search guarantees coverage at the cost of computational redundancy and semantic fragmentation. To...

    arxiv.org/abs/2605.23655 · PDF

  7. 07

    DualMem: Bypassing the Objectness Bottleneck for Calibrated Unknown-Stream Filtering in Open-World Object Detection

    Yingjun Xiao, Xi Chen, Gang Fang, Siyuan Chen

    cs.CV · cs.AI

    Open-world object detection (OWOD) requires detectors to localize known classes while identifying unknown objects for future incremental learning. We find that the unknown prediction streams of strong OWOD detectors are heavily polluted: on M-OWODB, across PROB, OW-DETR, and HypOW, future-task positive unknowns make up less than 10% of unknown predictions, whereas background false positives account for 46-71%. We show that this is not a...

    arxiv.org/abs/2605.23634 · PDF

  8. 08

    EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation

    Jente Vandersanden, Matheus Gadelha, Chun-Hao P. Huang, Hyeonho Jeong, Yulia Gryaditskaya

    cs.CV · cs.AI

    Multi-shot video generation requires maintaining a consistent appearance of recurring entities across shots while remaining faithful to shot-specific text prompts. Recent autoregressive methods reuse previously generated frames as memory. However, full-frame storage entangles persistent entity information with transient scene context, leading to irrelevant information leakage and high computational cost. We propose an entity-centric memory in...

    arxiv.org/abs/2605.23610 · PDF

  9. 09

    PathNavigate: A Training-Free Pathology Agent with Surprise-Guided Scan and Shared Slide Memory for Whole-Slide Image VQA

    Chunze Yang, Qidong Liu, Wenjie Zhao, Yue Tang, Jiusong Ge, Di Zhang, Jiashuai Liu, Lei Wu, Junbo Lu, Ni Zhang, Xian...

    cs.CV · cs.AI

    Whole-slide image visual question answering (WSI-VQA) frames pathology as an extreme-context search problem: to answer a free-form clinical query, a system must first navigate a gigapixel slide under a strict inspection budget to locate sparse, high-resolution evidence. Existing approaches largely fall into two paradigms: i) supervised pathology multimodal large language models (MLLMs) and agents can absorb localization and reasoning into...

    arxiv.org/abs/2605.23559 · PDF

  10. 10

    Multimodal Distribution Matching for Vision-Language Dataset Distillation

    Jongoh Jeong, Hoyong Kwon, Minseok Kim, Kuk-Jin Yoon

    cs.CV · cs.AI

    Dataset distillation compresses large training sets into compact synthetic datasets while preserving downstream performance. As modern systems increasingly operate on paired vision-language inputs, multimodal distillation must preserve representation quality and cross-modal alignment under tight compute and memory budgets, yet prior methods often require heavy computes and overlook their correlations. To address this, we present Multimodal...

    arxiv.org/abs/2605.23482 · PDF

  11. 11

    PhenoYieldNet: Learning Crop-Aware Phenological Responses for Multi-Crop Yield Prediction

    Yu Luo, Xiaogang Zhu, Shan Zeng, Wei Xiang, Thomas Francis Bishop, Zhiyong Wang, Kun Hu

    cs.CV · cs.AI

    Accurate crop yield prediction is crucial for sustainable agriculture and global food security. While existing methods are predominantly developed for single-crop prediction, they often struggle to generalize across diverse crop types, without addressing the unique crop phenological responses that are dynamically modulated by complex weather patterns. In this paper, we propose PhenoYieldNet, a multi-crop yield prediction framework that learns...

    arxiv.org/abs/2605.23478 · PDF

  12. 12

    One-Forcing: Towards Stable One-Step Autoregressive Video Generation

    Jiaqi Feng, Justin Cui, Yuanhao Ban, Cho-Jui Hsieh

    cs.CV · cs.AI

    Recent advances have substantially improved real-time interactive video generation in the autoregressive regime. However, most existing few-step autoregressive video generation methods, often distilled from a corresponding many-step teacher, default to a 4-step sampling configuration, which still incurs considerable latency during deployment and suffers from severe quality degradation when the number of sampling steps is further reduced,...

    arxiv.org/abs/2605.23458 · PDF

  13. 13

    Online Hand Gesture Recognition Using 3D Convolutional Neural Networks

    Yinghao Qin, Tijana Timotijevic

    cs.CV · cs.AI

    In human computer interaction, real-time detection and classification of dynamic hand gestures is challenging as: 1) the system must run in a real-time video stream and there is no noticeable lag in response after performing a gesture; 2) there is a large difference in how people perform gestures, making recognition more difficult. In this paper, an online hand gesture recognition system is proposed, which is able to localize gestures in...

    arxiv.org/abs/2605.23409 · PDF

  14. 14

    CHASD: Language Increment-Calibrated Contrastive Decoding against Hallucination in LVLMs

    Xiaoyi Huang, Kejia Zhang, Zhiming Luo

    cs.CV · cs.AI

    Large Vision-Language Models have shown strong multimodal reasoning capabilities, yet they remain susceptible to object hallucinations when language priors dominate insufficient or misaligned visual evidence. Training-free contrastive decoding methods mitigate this issue by comparing predictions from original and perturbed visual inputs, but existing approaches either apply global perturbations that may alter useful visual evidence or invoke...

    arxiv.org/abs/2605.23344 · PDF

  15. 15

    EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation

    Songlin Yang, Haobin Zhong, Ruilin Zhang, Xiaotong Zhao, Shuai Li, Kai Zheng, Xuyi Yang, Zhe Wang, Zhenchen Tang,...

    cs.CV · cs.AI

    The rapid evolution of generative video foundation models has propelled the field toward professional-grade cinematic synthesis. To achieve such demanding quality, the community transitions towards Reinforcement Learning (RL) and agentic workflows. However, reliable evaluation has emerged as a critical bottleneck. Existing benchmarks predominantly evaluate ''whether it is right'' (basic prompt-following) while fundamentally neglecting...

    arxiv.org/abs/2605.23271 · PDF

  16. 16

    ChainFlow-VLA: Causal Flow Planning with Vision-Language Models

    Xiyang Wang, Xinlin Wang, Tingguang Zhou, Gong Chen, Xingtai Gui, Zhi Xu, Xiaolei Wu, Feiyang Tan, Hangning Zhou, Mu Yang

    cs.CV · cs.AI · cs.RO

    Current end-to-end autonomous driving systems are fundamentally limited by a mismatch between temporal causal reasoning and global trajectory consistency. Autoregressive (AR) models capture interaction-aware temporal dependencies via causal factorization, but their step-wise decoding leads to error accumulation and suboptimal global structure. In contrast, diffusion models optimize trajectories globally but lack explicit causal constraints,...

    arxiv.org/abs/2605.23270 · PDF

  17. 17

    Coloring the Noise: Adversarial Sobolev Alignment for Faithful Image Super Resolution

    Hongbo Wang, Huaibo Huang, Pin Wang, Jinhua Hao, Chao Zhou, Ran He

    cs.CV · cs.AI

    Generative priors in Image Super-Resolution (SR) often compromise faithful restoration, we attribute this limitation to a fundamental spectral misalignment between isotropic objectives and the intrinsic natural image manifold. While Direct Preference Optimization offers a path to alignment, its reliance on spectrally flat Gaussian noise fails to distinguish authentic high-frequency details from hallucinations. To bridge this geometric gap, we...

    arxiv.org/abs/2605.23264 · PDF

  18. 18

    SimInsert: Seamless Video Object Insertion via Regional Sparse Attention Fusion

    Xinyu Chen, Yuyi Qian, Jiang Lin, Shenyi Wang, Gao Wang, Zhiqiu Zhang, Jizhi Zhang, Mingjie Wang, Qiang Tang, Qian...

    cs.CV · cs.AI

    Video object insertion requires ensuring spatio-temporal coherence and interactive realism, extending far beyond simple content placement. However, current approaches are often hindered by a reliance on explicit motion engineering or resource-intensive retraining, restricting their flexibility and generalization. To bridge this gap, we present \textit{SimInsert}, a training-free paradigm that efficiently decouples the task into intuitive...

    arxiv.org/abs/2605.23245 · PDF

  19. 19

    Lipschitz Optimization for Formal Verification of Homographies

    Jean-Guillaume Durand, Panagiotis Kouvaros, Maxime Gariel, Alessio Lomuscio

    cs.CV · cs.AI · cs.LG · cs.RO

    The adoption of vision neural networks in regulated industries requires formal robustness guarantees, especially in safety-critical domains such as healthcare, autonomous vehicles, and aerospace. However, current approaches are confined to incomplete statistical verification or robustness to $\ell_p$-norm and affine transforms, which cover only a narrow subset of perturbations to the image formation process. In particular, robustness to...

    arxiv.org/abs/2605.23203 · PDF

  20. 20

    Exploiting Longitudinal Context in Clinician-Verified Interactive Lesion Tracking

    Yannick Kirchhoff, Maximilian Rokuss, Daniel Philipp Mertens, David Füller, Benjamin Hamm, Andreas Schreyer, Oliver...

    cs.CV · cs.AI · cs.LG

    Tracking tumor lesions across serial CT scans is essential for oncological response assessment. Existing automated methods face a fundamental trade-off: end-to-end trackers achieve high automation but offer no opportunity to correct silent tracking failures, while decoupled registration-segmentation pipelines permit user verification yet discard the lesion's prior appearance, limiting accuracy in ambiguous cases. In this work, we propose a...

    arxiv.org/abs/2605.23118 · PDF

  21. 21

    CoReVAD: A Contextual Reasoning Framework for Training-Free Video Anomaly Detection

    Hyeongmuk Lim, Youngbum Hur

    cs.CV · cs.AI

    Existing Video Anomaly Detection (VAD) methods typically rely on task-specific training, leading to strong domain dependency and high training costs. Moreover, most existing methods output only scalar anomaly scores, providing limited insight into why specific events are considered abnormal. Recent advances in Vision-Language Models (VLMs) have enabled both anomaly detection and human-interpretable reasoning. However, many VLM-based...

    arxiv.org/abs/2605.23116 · PDF

  22. 22

    Dithering Defense: Adversarial Robustness of Vision Foundation Models via Multi-Level Floyd-Steinberg Dithering

    Yury Belousov, Brian Pulfer, Vitaliy Kinakh, Slava Voloshynovskiy

    cs.CV · cs.AI · cs.LG

    Vision foundation models are widely used as frozen backbones across many downstream tasks, making them a single point of failure under adversarial attack. We study multi-level Floyd-Steinberg error-diffusion dithering as a lightweight, model-agnostic input transformation that disrupts adversarial perturbations while preserving semantic content. Unlike prior work, which was limited to binary dithering, grayscale CIFAR-10, and a single small...

    arxiv.org/abs/2605.23065 · PDF

  23. 23

    The TIME Machine: On The Power of Motion for Efficient Perception

    Mantas Skackauskas, Xinyue Hao, Laura Sevilla-Lara

    cs.CV · cs.AI · cs.LG

    Video representation learning has seen tremendous progress in recent years. This has been driven by many factors, including the scale of training and the success of visual models trained contrastively with language. While these factors have pushed the boundaries of what video models can do, they also introduce their own set of limitations: first, scaling video models can reach prohibitive costs and second, learning from language restricts the...

    arxiv.org/abs/2605.23045 · PDF

  24. 24

    Suicide Risk Assessment from AI-powered Video Surveillance: An Interpretable Framework for Prevention in Metro Stations

    Safwen Naimi, Wassim Bouachir, Guillaume-Alexandre Bilodeau, Brian Mishara

    cs.CV · cs.AI

    Understanding and monitoring human behavior in metro stations play an important role in supporting suicide prevention efforts, where early identification of high-risk situations can enable timely intervention. This requires assessing suicide risk from a surveillance video by jointly reasoning about the behavior of each passenger, his/her spatial context, and temporal dynamics. However, this assessment using videos captured by surveillance...

    arxiv.org/abs/2605.22904 · PDF

  25. 25

    Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision?

    Zixuan Lan, Luzhe Sun, Matthew R. Walter, Jiawei Zhou

    cs.CV · cs.AI · cs.CL

    Benchmark accuracy is often implicitly assumed to reflect grounded visual understanding in vision-language models (VLMs), yet it remains unclear to what extent such scores truly reflect reliance on visual evidence. Motivated by a surprising observation that removing a substantial fraction of image tokens only degrades model performance very slightly on a widely used hallucination benchmark, we systematically investigate this mismatch in a set...

    arxiv.org/abs/2605.22903 · PDF

  26. 26

    AnyMo: Geometry-Aware Setup-Agnostic Modeling of Human Motion in the Wild

    Baiyu Chen, Zechen Li, Wilson Wongso, Lihuan Li, Xiachong Lin, Hao Xue, Benjamin Tag, Flora Salim

    cs.CV · cs.AI · cs.CL · cs.HC

    As wearable and mobile devices become increasingly embedded in daily life, they offer a practical way to continuously sense human motion in the wild. But inertial signals are highly dependent on the sensing setup, including body location, mounting position, sensor orientation, device hardware, and sampling protocol. This setup dependence makes it difficult to learn motion representations that transfer across devices and datasets, and limits...

    arxiv.org/abs/2605.22715 · PDF

  27. 27

    Swift Sampling: Selecting Temporal Surprises via Taylor Series

    Dahye Kim, Bhuvan Sachdeva, Karan Uppal, Naman Gupta, Vineeth N. Balasubramanian, Deepti Ghadiyaram

    cs.CV · cs.AI

    While most frames in long-form video are redundant, the critical information resides in temporal surprises: moments where the actual visual features deviate from their predicted evolution. Inspired by the human brain's predictive coding, we introduce Swift Sampling, an elegant, training-free frame selection algorithm that automatically identifies high-information moments in a video. Specifically, we model a video as a differentiable...

    arxiv.org/abs/2605.22678 · PDF

  28. 28

    SceneAligner: 3D-Grounded Floorplan Localization in the Wild

    Junhyeong Cho, Ruojin Cai, Hadar Averbuch-Elor

    cs.CV · cs.AI · cs.LG

    Many public buildings provide floorplans with a "you are here" indicator to help visitors orient themselves. Floorplan localization seeks to computationally replicate this capability by determining where visual observations were captured within a floorplan. However, existing methods typically assume controlled small-scale environments and precise vectorized floorplans, limiting their ability to operate in large-scale buildings and rasterized...

    arxiv.org/abs/2605.22581 · PDF

  29. 29

    VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis

    Jinho Park, Youbin Kim, Hogun Park, Eunbyung Park

    cs.CV · cs.AI

    Spatio-temporal reasoning is a core capability for Multimodal Large Language Models (MLLMs) operating in the real world. As such, evaluating it precisely has become an essential challenge. However, existing spatio-temporal reasoning benchmark datasets primarily rely on static image sets or passively curated video data, which limits the evaluation of fine-grained reasoning capabilities. In this paper, we introduce VGenST-Bench, a video...

    arxiv.org/abs/2605.22570 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.