cs.CV · 2026-05-23 · No. 8

Computer Vision and Pattern Recognition, 2026-05-23.

32 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

32 entries
  1. 01

    AnyMo: Geometry-Aware Setup-Agnostic Modeling of Human Motion in the Wild

    Baiyu Chen, Zechen Li, Wilson Wongso, Lihuan Li, Xiachong Lin, Hao Xue, Benjamin Tag, Flora Salim

    cs.CV · cs.AI · cs.CL · cs.HC

    As wearable and mobile devices become increasingly embedded in daily life, they offer a practical way to continuously sense human motion in the wild. But inertial signals are highly dependent on the sensing setup, including body location, mounting position, sensor orientation, device hardware, and sampling protocol. This setup dependence makes it difficult to learn motion representations that transfer across devices and datasets, and limits...

    arxiv.org/abs/2605.22715 · PDF

  2. 02

    Swift Sampling: Selecting Temporal Surprises via Taylor Series

    Dahye Kim, Bhuvan Sachdeva, Karan Uppal, Naman Gupta, Vineeth N. Balasubramanian, Deepti Ghadiyaram

    cs.CV · cs.AI

    While most frames in long-form video are redundant, the critical information resides in temporal surprises: moments where the actual visual features deviate from their predicted evolution. Inspired by the human brain's predictive coding, we introduce Swift Sampling, an elegant, training-free frame selection algorithm that automatically identifies high-information moments in a video. Specifically, we model a video as a differentiable...

    arxiv.org/abs/2605.22678 · PDF

  3. 03

    SceneAligner: 3D-Grounded Floorplan Localization in the Wild

    Junhyeong Cho, Ruojin Cai, Hadar Averbuch-Elor

    cs.CV · cs.AI · cs.LG

    Many public buildings provide floorplans with a "you are here" indicator to help visitors orient themselves. Floorplan localization seeks to computationally replicate this capability by determining where visual observations were captured within a floorplan. However, existing methods typically assume controlled small-scale environments and precise vectorized floorplans, limiting their ability to operate in large-scale buildings and rasterized...

    arxiv.org/abs/2605.22581 · PDF

  4. 04

    VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis

    Jinho Park, Youbin Kim, Hogun Park, Eunbyung Park

    cs.CV · cs.AI

    Spatio-temporal reasoning is a core capability for Multimodal Large Language Models (MLLMs) operating in the real world. As such, evaluating it precisely has become an essential challenge. However, existing spatio-temporal reasoning benchmark datasets primarily rely on static image sets or passively curated video data, which limits the evaluation of fine-grained reasoning capabilities. In this paper, we introduce VGenST-Bench, a video...

    arxiv.org/abs/2605.22570 · PDF

  5. 05

    Case-Aware Medical Image Classification with Multimodal Knowledge Graphs and Reliability-Guided Refinement

    Yiming Xu, Yixuan Liu, Yuhang Zhang, Ling Zheng, Yihan Wang, Qi Song

    cs.CV · cs.AI

    Deep learning has brought significant progress to medical image classification, yet most existing methods still rely on isolated visual evidence and cannot effectively leverage similar cases or external knowledge. In clinical practice, diagnosis is typically supported by historical similar cases and their associated symptoms. To simulate this diagnostic process, we propose a framework that performs case-aware reasoning using multimodal...

    arxiv.org/abs/2605.22547 · PDF

  6. 06

    Making the Discrete Continuous: Synthetic RAW Augmentations for Fine-Grained Evaluation of Person Detection Performance in Low Light

    Valeria Pais, Malena Mendilaharzu, Daniele Faccio, Luis Oala, Christoph Clausen, Bruno Sanguinetti

    cs.CV · cs.AI · cs.LG · physics.optics

    Real-world deployment of AI vision models is both fueled and limited by the data available for training and testing. Real datasets are sparse and uneven: long-tailed or unbalanced distributions hinder generalization, and the low number of samples in low density regions makes it hard to run evaluations. Synthetic data can fill these gaps, providing us with a way to sample the input space more continuously and improve data coverage for...

    arxiv.org/abs/2605.22455 · PDF

  7. 07

    Pre-VLA: Preemptive Runtime Verification for Reliable Vision-Language-Action and World-Model Rollouts

    Zhen Sun, Yongjian Guo, Haoran Sun, Luqiao Wang, Wei Lu, Jiachi Ji, Shengzhe Ji, Junwu Xiong, Zhijun Meng

    cs.CV · cs.AI · cs.RO

    While large vision-language-action (VLA) models and generative world models (WM) have advanced long-horizon embodied intelligence, their practical deployment remains challenged by uncertainty in learning-based action generation. Low-quality actions may cause physical failures during execution or lead to misleading world-model rollouts with redundant rendering costs. To address this issue, we propose Pre-VLA, a unified runtime verification...

    arxiv.org/abs/2605.22446 · PDF

  8. 08

    FastTab: A Fast Table Recognizer with a Tiny Recursive Module and 1D Transformers

    Laziz Hamdi, Amine Tamasna, Pascal Boisson, Thierry Paquet

    cs.CV · cs.AI

    Table structure recognition (TSR) requires both table-level coherence (row/column counts, headers, spanning cells) and precise separator localization. We introduce FastTab, a grid-centric TSR model that avoids autoregressive HTML decoding by combining (i) a lightweight Tiny Recursive Module (TRM) for global reasoning and (ii) axial 1D Transformer encoders that capture long-range dependencies along rows and columns. The model predicts...

    arxiv.org/abs/2605.22422 · PDF

  9. 09

    Diffusion-guided Generalizable Enhancer for Urban Scene Reconstruction

    Henry Che, Jingkang Wang, Yun Chen, Ze Yang, Sivabalan Manivasagam, Raquel Urtasun

    cs.CV · cs.AI · cs.RO

    Urban scene reconstruction from real-world observations has emerged as a powerful tool for self-driving development and testing. While current neural rendering approaches achieve high-fidelity rendering along the recorded trajectories, their quality degrades significantly under large viewpoint shifts, limiting the applicability for closed-loop simulation. Recent works have shown promising results in using diffusion models to enhance quality...

    arxiv.org/abs/2605.22420 · PDF

  10. 10

    Towards Clinically Interpretable Ophthalmic VQA via Spatially-Grounded Lesion Evidence

    Xingyue Wang, Bo Liu, Meng Wang, Zhixuan Zhang, Chengcheng Zhu, Huazhu Fu, Jiang Liu

    cs.CV · cs.AI

    Visual Question Answering (VQA) holds great promise for clinical support, particularly in ophthalmology, where retinal fundus photography is essential for diagnosis. However, ophthalmic VQA benchmarks primarily emphasize answer accuracy, neglecting the explicit visual evidence necessary for clinical interpretability. In this work, we introduce FundusGround, a new benchmark for clinically interpretable ophthalmic VQA with spatially-grounded...

    arxiv.org/abs/2605.22414 · PDF

  11. 11

    VEELA: A Clinically-Constrained Benchmark for Liver Vessel Segmentation in Computed Tomography Angiography

    Ziya Ata Yazıcı, N. Sinem Gezer, İlkay Öksüz, İlker Özgür Koska, Tuğçe Toprak, Pervin Bulucu, Ufuk Beşenk, A. Emre...

    cs.CV · cs.AI

    Accurate segmentation of hepatic and portal vessels in contrast-enhanced computed tomography angiography (CTA) remains challenging due to complex vascular topology, peripheral visibility limitations, and acquisition-induced ambiguities. While existing public datasets offer valuable benchmarks, few include clinically realistic annotation constraints. We introduce VEELA (Vessel Extraction and Extrication for Liver Analysis), a rigorously...

    arxiv.org/abs/2605.22357 · PDF

  12. 12

    Bernini: Latent Semantic Planning for Video Diffusion

    Bernini Team, Chenchen Liu, Junyi Chen, Lei Li, Lu Chi, Mingzhen Sun, Zhuoying Li, Yi Fu, Ruoyu Guo, Yiheng Wu, Ge...

    cs.CV · cs.AI · cs.MM

    Multimodal large language models (MLLMs) and diffusion models have each reached remarkable maturity: MLLMs excel at reasoning over heterogeneous multimodal inputs with strong semantic grounding, while diffusion models synthesize images and videos with photorealistic fidelity. We argue that these two families can be unified through a simple division of labor: MLLMs perform semantic planning, while diffusion models render pixels from high-level...

    arxiv.org/abs/2605.22344 · PDF

  13. 13

    4D-GSW: Kinematic-Aware Spatio-Temporal Consistent Watermarking for 4D Gaussian Splatting

    Sifan Zhou, Hang Zhang, Yuhang Wang, Ming Li

    cs.CV · cs.AI

    While 4D Gaussian Splatting (4DGS) has revolutionized high-fidelity dynamic reconstruction, safeguarding the intellectual property of these assets remains an open challenge. Conventional steganographic techniques often neglect the underlying kinematic manifolds, triggering non-physical artifacts such as severe temporal flickering and "FVD collapse". To address this, we propose \textbf{4D-GSW}, a kinematic-aware watermarking framework designed...

    arxiv.org/abs/2605.22342 · PDF

  14. 14

    MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering

    Junbin Xiao, Jiajun Chen, Tianxiang Sun, Xun Yang, Angela Yao

    cs.CV · cs.AI · cs.MM

    Long streaming video QA remains challenging due to growing visual tokens and limited reasoning length of large language models (LLMs). KV-caching stores the Key-Value (KV) of the historical tokens via LLM prefill and enables more efficient streaming QA. However, existing methods cache every one or two frames, causing redundant memory usage and losing fine-grained spatial details within frame or temporal contexts across frames. This paper...

    arxiv.org/abs/2605.22269 · PDF

  15. 15

    OSS: Open Suturing Skills Vision-Based Assessment Challenge 2024-2025

    Hanna Hoffmann, Setareh Bady, Claas de Boer, Max Kirchner, Jan Egger, Rainer Röhrig, Frank Hölzle, Lennart Johannes...

    cs.CV · cs.AI · cs.LG

    Achieving high levels of surgical skill through effective training is essential for optimal patient outcomes. Automated, data-driven skill assessment holds significant potential to improve surgical training. While machine learning-based methods are increasingly popular for assessing skills in minimally invasive surgery, their application to open surgery remains limited. We present the results of a dedicated MICCAI challenge designed to...

    arxiv.org/abs/2605.22200 · PDF

  16. 16

    TextTeacher: What Can Language Teach About Images?

    Tobias Christian Nauen, Stanislav Frolov, Brian Bernhard Moser, Federico Raue, Ahmed Anwar, Andreas Dengel

    cs.CV · cs.AI · cs.LG

    The platonic representation hypothesis suggests that sufficiently large models converge to a shared representation geometry, even across modalities. Motivated by this, we ask: Can the semantic knowledge of a language model efficiently improve a vision model? As an answer, we introduce TextTeacher, a simple auxiliary objective that injects text embeddings as additional information into image classification training. TextTeacher uses readily...

    arxiv.org/abs/2605.22098 · PDF

  17. 17

    LVDrive: Latent Visual Representation Enhanced Vision-Language-Action Autonomous Driving Model

    Xiaodong Mei, Diankun Zhang, Hongwei Xie, Guang Chen, Hangjun Ye, Dan Xu

    cs.CV · cs.AI

    Vision-Language-Action (VLA) models have emerged as a promising framework for end-to-end autonomous driving. However, existing VLAs typically rely on sparse action supervision, which underutilizes their powerful scene understanding and reasoning capabilities. Recent attempts to incorporate dense visual supervision via world modeling often overemphasize pixel-level image reconstruction, neglecting semantically meaningful scene representation...

    arxiv.org/abs/2605.22089 · PDF

  18. 18

    JMed48k: A Multi-Profession Japanese Medical Licensing Benchmark for Vision-Language Model Evaluation

    Yue Xun, Junyu Liu, Qian Niu, Xinyi Wang, Zheng Yuan, Zirui Li, Zequn Zhang, Bowen Zhao, Shujun Wang, Irene Li, Kan...

    cs.CV · cs.AI

    We introduce JMed48k, a multi-profession Japanese healthcare licensing benchmark for evaluating vision-language models. Built from official PDF materials released by the Japanese Ministry of Health, Labour and Welfare, JMed48k contains 48,862 exam questions and 20,142 images from 11 national licensing examinations between 2005 and 2025, with visual content annotated under an 8-type taxonomy. From this corpus, we derive JMed48k-Eval, a recent...

    arxiv.org/abs/2605.22080 · PDF

  19. 19

    Echo4DIR: 4D Implicit Heart Reconstruction from 2D Echocardiography Videos

    Yanan Liu, Qinya Li, Hao Zhang, Kangjian He, Xuan Yang, Hao Li, Dan Xu, Lei Li

    cs.CV · cs.AI

    Reconstructing 4D (3D+t) cardiac geometry from sparse 2D echocardiography is highly desirable yet fundamentally challenged by geometric ambiguity and temporal discontinuity. To tackle these issues, we propose Echo4DIR, a novel test-time 4D implicit reconstruction framework. Specifically, we learn robust 3D shape priors from statistical shape models (SSMs) via a cardiac conditional SDF, constructing an Epipolar Mask Encoder module with...

    arxiv.org/abs/2605.22066 · PDF

  20. 20

    GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation

    Jiahao Yang, Zihan Wang, Xiangyang Li, Xing Zhu, Yujun Shen, Yinghao Xu, Shuqiang Jiang

    cs.CV · cs.AI

    Despite significant progress in Vision-Language Navigation (VLN), existing approaches still rely on dense RGB videos that produce excessive patch tokens and lack explicit spatial structure, resulting in substantial computational overhead and limited spatial reasoning. To address these issues, we introduce the Geometry-Aware BEV (GA-BEV) - a compact, 3D-grounded feature representation that integrates both explicit and implicit geometric cues...

    arxiv.org/abs/2605.22036 · PDF

  21. 21

    AgroVG: A Large-Scale Multi-Source Benchmark for Agricultural Visual Grounding

    Haocheng Li, Juepeng Zheng, Zenghao Yang, Kaiqi Du, Guilong Xiao, Gengmeng Pu, Haohuan Fu, Jianxi Huang

    cs.CV · cs.AI

    Visual grounding, the task of localizing objects described by natural-language expressions, is a foundational capability for agricultural AI systems, enabling applications such as selective weeding, disease monitoring, and targeted harvesting. Reliable evaluation of agricultural visual grounding remains challenging because agricultural targets are often small, repetitive, occluded, or irregularly shaped, and instructions may refer to one,...

    arxiv.org/abs/2605.22034 · PDF

  22. 22

    FRED: A Multi-Modal Autonomous Driving Dataset for Flooded Road Environments

    Connor Malone, Sebastien Demmel, Sebastien Glaser

    cs.CV · cs.AI · cs.RO

    The Flooded Road Environments Dataset (FRED) is, to our knowledge, the first multi-modal autonomous driving dataset specifically targeting the collection of data from scenarios involving water hazards on the road. The dataset contains images from a 2.3 MP FLIR Blackfly USB3 camera, 64-beam 360$^\circ$ point clouds from an Ouster OS1-64 LiDAR, and data from an iXblue ATLANS-C IMU corrected by a Geoflex RTK GNSS, from five separate locations...

    arxiv.org/abs/2605.22018 · PDF

  23. 23

    Virtual 3D H&E Staining from Phase-contrast Back-illumination Interference Tomography

    Anthony Song, Boyan Zhou, Mayank Golhar, Marisa Morakis, Alex Baras, Nicholas Durr

    cs.CV · cs.AI

    Three-dimensional (3D) histopathology of unprocessed tissues has the potential to transform disease management by enabling volumetric characterization of tissue microarchitecture and in-vivo assessment. Back-illumination Interference Tomography (BIT) is a new phase microscopy technology that provides rapid, non-destructive volumetric imaging of unprocessed tissues. However, translating BIT volumes into clinically interpretable H&E images...

    arxiv.org/abs/2605.22000 · PDF

  24. 24

    Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning

    Dazhao Du, Jian Liu, Jialong Qin, Tao Han, Bohai Gu, Fangqi Zhu, Yujia Zhang, Eric Liu, Xi Chen, Song Guo

    cs.CV · cs.AI

    Video large language models (Video LLMs) achieve strong benchmark accuracy, yet often answer video questions through shortcuts such as single-frame cues and language priors rather than by tracking spatiotemporal dynamics. This issue is exacerbated in RL post-training, where correctness-only rewards can further reinforce shortcut policies that obtain high reward without tracking video dynamics. We address this by asking a controlled...

    arxiv.org/abs/2605.21988 · PDF

  25. 25

    Interpreting and Enhancing Emotional Circuits in Large Vision-Language Models via Cross-Modal Information Flow

    Chengsheng Zhang, Chenghao Sun, Zhining Xie, Xinmei Tian

    cs.CV · cs.AI

    Large Vision-Language Models (LVLMs) represent a significant leap towards empathetic agents, demonstrating remarkable capabilities in emotion understanding. However, the internal mechanisms governing how LVLMs translate abstract visual stimuli into coherent emotional narratives remain largely unexplored, primarily due to the scarcity of visual counterfactuals and the diffuse nature of emotional expression. In this paper, we bridge this gap by...

    arxiv.org/abs/2605.21980 · PDF

  26. 26

    Video as Natural Augmentation: Towards Unified AI-Generated Image and Video Detection

    Zhengcen Li, Chenyang Jiang, Liangxu Su, Tong Shao, Shiyang Zhou, Ming Tao, Jingyong Su

    cs.CV · cs.AI

    AI-generated content (AIGC) is rapidly improving, creating an urgent need for detectors that generalize across data sources, deployment pipelines, and visual modalities. A strongly generalizable detector should remain robust under distributional variations. However, we identify a consistent failure mode: SOTA AI-generated image detectors often collapse when applied to frames extracted from videos. Through systematic analysis, we show that...

    arxiv.org/abs/2605.21977 · PDF

  27. 27

    MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues

    Dazhao Du, Liao Duan, Jian Liu, Tao Han, Yujia Zhang, Eric Liu, Xi Chen, Song Guo

    cs.CV · cs.AI

    Video temporal grounding (VTG), which localizes the start and end times of a queried event in an untrimmed video, is a key test of whether multimodal large language models (MLLMs) understand not only what happens but also when it happens. Although modern MLLMs describe video content fluently, their timestamp predictions remain unreliable, while existing remedies either require costly post-training on temporal annotations or rely on coarse...

    arxiv.org/abs/2605.21954 · PDF

  28. 28

    SDGBiasBench: Benchmarking and Mitigating Vision--Language Models' Biases in Sustainable Development Goals

    Zihang Lin, Huaiyuan Qin, Muli Yang, Hongyuan Zhu

    cs.CV · cs.AI

    Assessing progress toward the Sustainable Development Goals (SDGs) requires multi-step reasoning over visual cues, contextual knowledge, and development indicators, where incomplete evidence use and imperfect evidence integration can introduce hidden prediction biases. Real-world SDG monitoring further spans both qualitative judgments and quantitative estimation. However, existing benchmarks typically evaluate these aspects in isolation,...

    arxiv.org/abs/2605.21919 · PDF

  29. 29

    MAVEN: A Multi-stage Agentic Annotation Pipeline for Video Reasoning Tasks

    Han Zhang, Wanting Jiang, Tomasz Kornuta, Tian Zheng, Vidya Murali

    cs.CV · cs.AI

    Training Vision Language Models (VLMs) for video event reasoning requires high-quality structured annotations capturing not only what happened, but when, where, why, and with what consequence, at a scale manual labelling cannot support. We present MAVEN (Multi-stage Agentic Video Event aNnotation), a multi-stage agentic pipeline that turns raw videos into multi-task training data with Chain-of-Thought (CoT) reasoning traces, organized around...

    arxiv.org/abs/2605.21917 · PDF

  30. 30

    Two-Stage Multimodal Framework for Emotion Mimicry Intensity Prediction

    Dinithi Dissanayake, Shaveen Silva, Ovindu Atukorala, Prasanth Sasikumar, Suranga Nanayakkara

    cs.CV · cs.AI · cs.HC

    We present our submission to the Hume-ABAW10 Emotional Mimicry Intensity (EMI) Challenge, which aims to predict six continuous emotion intensity dimensions: Admiration, Amusement, Determination, Empathic Pain, Excitement, and Joy, from in-the-wild multimodal video clips. We propose a staged multimodal framework that combines textual, acoustic, and visual representations, with an optional motion branch. Our approach first trains...

    arxiv.org/abs/2605.21869 · PDF

  31. 31

    Learning Emergent Modular Representations in Multi-modality Medical Vision Foundation Models

    Yuting He, Chenyu You, Shuo Li

    cs.CV · cs.AI

    Multi-modality medical vision (MV) foundation models (FM) are fundamentally challenged by pronounced Non-IID feature statistics across heterogeneous imaging modalities. Monolithic self-supervised optimization on such data induces conflicting gradients, driving representations to collapse toward modality-dominant shortcuts. This work reframes this failure as an imbalance between specialization and coordination in emergent modularity, and...

    arxiv.org/abs/2605.21861 · PDF

  32. 32

    CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models

    Zhi Liu

    cs.CV · cs.AI

    Vision-Language-Action (VLA) models have rapidly converged on a small set of architectural patterns: discrete-token autoregression (e.g. OpenVLA) and continuous-action flow-matching (e.g. pi-0.5). Yet preference alignment via Direct Preference Optimisation (DPO) -- the de-facto post-training step in language models -- has been studied almost exclusively on autoregressive VLAs. We present CrossVLA, an empirical study of cross-paradigm VLA...

    arxiv.org/abs/2605.21854 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.