cs.CV · 2026-06-09 · No. 18

Computer Vision and Pattern Recognition, 2026-06-09.

18 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

18 entries
  1. 01

    OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics

    Mingxian Lin, Shengju Qian, Yuqi Liu, Yi-Hua Huang, Yiyu Wang, Wei Huang, Yitang Li, Fan Zhang, Zeyu Hu, Lingting...

    cs.CV · cs.AI

    Vision-language model (VLM) agents are increasingly deployed in interactive game environments. Yet game benchmarks for VLM agents typically report a single first-attempt score per (agent, game) pair, focus on single-agent Solo play, and lack unified protocols for evaluating heterogeneous agent classes (commercial VLMs, open-weight VLMs, and specialized game policies) on the same footing. We address these gaps with OmniGameArena, a real-time...

    arxiv.org/abs/2606.09826 · PDF

  2. 02

    PTL-Diffusion: Manifold-Aware Diffusion with Periodic Terminal Laws

    Danqi Zhuang, Jisui Huang, Xiaoyue Xi, Andrew Kiggins, Xiaojie Wang, Ke Chen, Yue Wu

    cs.CV · cs.AI · math.PR

    Standard diffusion models typically use a single time-homogeneous Gaussian terminal distribution as the reference law for generation. While this choice is analytically convenient and empirically powerful, it provides little explicit structure for data concentrated near low-dimensional manifolds, where different regions of the data distribution may correspond to distinct local geometric or semantic factors. As a result, the reverse model must...

    arxiv.org/abs/2606.09816 · PDF

  3. 03

    Echo-Memory: A Controlled Study of Memory in Action World Models

    Wayne King, Zeyue Xue, Yuxuan Bian, Jie Huang, Haoran Li, Yaowei Li, Yaofeng Su, Yuming Li, Haoyu Wang, Shiyi Zhang,...

    cs.CV · cs.GR · cs.LG

    We present \textbf{Echo-Memory}, a controlled study of memory mechanisms in action-conditioned world models. These models generate multi-segment videos from a first frame, text prompt, and camera-action sequence, but their central failure is often memory rather than local image synthesis: after the camera leaves and returns, the scene or salient object may silently change. Existing memory designs are hard to compare because gains are...

    arxiv.org/abs/2606.09803 · PDF

  4. 04

    Hybrid Robustness Verification for Spatio-Temporal Neural Networks

    Sherwin Varghese, Matthew Wicker, Alessio Lomuscio

    cs.CV · cs.AI · cs.LG

    With AI increasingly deployed in safety-critical systems, providing formal robustness guarantees for the underlying models is essential. Existing verification methods either rely on overly conservative approximations or incur prohibitive computational costs. For example, the use of lp-norm perturbations in video settings encodes the belief that the adversary can inject noise in every video frame. In practice, adversarial perturbations exhibit...

    arxiv.org/abs/2606.09746 · PDF

  5. 05

    Visual Prompting Meets Feature Reconstruction-Based Anomaly Detection with Dual-Teacher Supervision

    Mateo Diaz-Bone, Daniel Caraballo, Florian Scheidegger, Thomas Frick, Mattia Rigotti, Andrea Bartezzaghi, Roy Assaf,...

    cs.CV · cs.AI

    Recent Anomaly Detection methods achieve perfect detection and segmentation scores on well-established datasets, such as MVTec. However, many of these methods face challenges when foundational assumptions - such as consistent object scale, viewpoint, background, illumination, and centered placement - are violated. Those variations that occur render anomaly detection methods unusable in many real-world scenarios. To address these limitations,...

    arxiv.org/abs/2606.09670 · PDF

  6. 06

    Do Video Foundation Models Understand Intuitive Physics? A Layerwise Probing Analysis

    Samuele Punzo, Niccolò Caselli, Ippokratis Pantelidis, Francesco Massafra, Salvatore Lo Sardo, Mohammadreza Salehi

    cs.CV · cs.AI · cs.LG

    We study whether pretrained video foundation models encode intuitive-physics information in their frozen representations, and how this information varies across model families, layers, and probe types. Using frozen-feature probing on IntPhys2 and Minimal Video Pairs (MVP), we compare predictive joint-embedding models (V-JEPA), masked reconstruction models (VideoMAE), and a diffusion-based video generator (LTX-Video). V-JEPA achieves the...

    arxiv.org/abs/2606.09646 · PDF

  7. 07

    ATN3D: Density-Aware LiDAR-Radar Early 3D Object Detection Under Extreme Sparsity

    Debojyoti Biswas, Xianbiao Hu

    cs.CV · cs.AI

    3D object detection is the backbone of perception for automated vehicles (AV) and broader intelligent transportation systems applications. Long-range detection is challenging because sensing evidence is sparse; yet this ``long-range'' scenario is routine in traffic. Although >30m is often labeled long-range in computer vision, on roadways it affords only approx. 1-2s for perception and decision-making. Under such extreme sparsity, two core...

    arxiv.org/abs/2606.09634 · PDF

  8. 08

    Streaming Interventions: Can Video Large Language Models Correct Mistakes as They Occur?

    Apratim Bhattacharyya, Shweta Mahajan, Sanjay Haresh, Rajeev Yasarla, Reza Pourreza, Litian Liu, Risheek Garrepalli,...

    cs.CV · cs.LG

    Learning everyday skills, like cooking a dish, relies increasingly on instructional media such as online videos. This opens the door to the use of video (and multimodal) large language models (LLMs) as task guidance assistants. A crucial capability for the real-world success of a prospective task guidance assistant is it's ability to intervene proactively as soon as a mistake is apparent in order to guide the user. To evaluate this crucial...

    arxiv.org/abs/2606.09547 · PDF

  9. 09

    Real-time body pose non-verbal communication with a consistency-based reliability measure

    Alina Marcu, Dragos Costea, Cristina Lazar, Marius Leordeanu

    cs.CV · cs.AI · cs.RO

    Body movement communicates intent at distances and in conditions where neither the face, nor speech can be captured. We study the recognition of communicative intent from 2D body pose alone. We argue that body motion is a reliable signal especially in scenarios that require real time low-cost on-device person-to-robot communication in long distance environments, such as rescue missions. However, existing resources do not isolate this signal....

    arxiv.org/abs/2606.09390 · PDF

  10. 10

    PhysScene: A Scene Graph Dataset for Scientific Visual Reasoning in Physics Experiments

    Minghao Zou, Qingtian Zeng, Shangkun Liu, Yanda Meng, Guanghui Yue, Baoquan Zhao, Abdulmotaleb El Saddik, Wei Zhou

    cs.CV · cs.AI

    Scene Graphs (SGs) provide structured representations of visual scenes by modeling objects and their pairwise relationships. Despite recent progress, existing datasets primarily focus on generic natural contexts, leaving domain-specific and function-oriented scenes largely underexplored. This limitation restricts the evaluation of relational reasoning in scientific experimental scenes, thereby hindering the development of intelligent...

    arxiv.org/abs/2606.09368 · PDF

  11. 11

    Zero-Shot Semantic Re-Identification for Autonomous Driving: A VLM Baseline Study

    Eduardo Borges, Manuel Abreu, Luís Garrote, Urbano J. Nunes

    cs.CV · cs.LG

    Re-Identification (ReID) in autonomous driving is typically formulated as a visual matching problem, where observations of vehicles, pedestrians, and cyclists are associated across time, frames, or camera views using learned appearance embeddings, often complemented by motion, geometric, or multimodal cues. However, purely visual representations may be sensitive to viewpoint, occlusion, illumination, and sensor-domain variations, limiting...

    arxiv.org/abs/2606.09362 · PDF

  12. 12

    Beyond Humans: Multispecies Animal Face Recognition Using Transfer Learning

    Maria De Marsico, Anil K. Jain, Annalaura Miglino

    cs.CV · cs.AI

    Individual animal recognition can be useful in the search for lost or stolen pets, the tracking of individuals of endangered species, and the recognition of animals in crowded farms. Present recognition techniques mostly use physical devices, e.g., microchips, often impractical and difficult to apply. These could be replaced by remote recognition via the animal's face; if accurate enough, it provides several advantages: it is non-invasive,...

    arxiv.org/abs/2606.09353 · PDF

  13. 13

    Proposal Refinement for Few-Shot Object Detection

    Yuan Zeng, Bin Song, Jie Guo, Yuwen Chen

    cs.CV · cs.AI

    Few-shot object detection has gained widely attention in recent years. Some excellent algorithms have been proposed to handle this task. However, most of these algorithms rely on the performance of few-shot classification. Unlike previous attempts, our work focuses on the problem of unbalanced distribution of region proposals between the novel classes and the base classes. In order to alleviate this unbalanced distribution, we propose the...

    arxiv.org/abs/2606.09245 · PDF

  14. 14

    EgoTactile: Learning Grasp Pressure for Everyday Objects from Egocentric Video

    Yuan Zeng, Yujia Shi, Tiao Tan, Xingting Li, Yaqi Qin, Zongqing Lu, Wenming Yang, Jing-Hao Xue, Qingmin Liao

    cs.CV · cs.AI

    Estimating full-hand grasp pressure from egocentric video is critical for immersive VR and robotic manipulation, yet dense tactile sensing often relies on intrusive hardware. Existing vision-based methods predominantly rely on planar surfaces or fingertip contacts, failing to generalize to complex 3D object interactions. Therefore, we introduce EgoTactile, a benchmark pairing egocentric video with full-hand pressure supervision for diverse...

    arxiv.org/abs/2606.09243 · PDF

  15. 15

    Counterfactual Reasoning for Fine-Grained Evidence Disentanglement in VideoQA

    Zhou Du, Hamid Krim, Xiao Wu, Zhaoquan Yuan, Liangwei Li, Keisuke Fujii

    cs.CV · cs.LG

    Recent advances in video multimodal models have significantly improved VideoQA performance. However, these systems often rely on spurious statistical correlations rather than answer-relevant causal evidence, resulting in unfaithful and brittle reasoning, especially in complex real-world scenarios. Existing methods either rely on cross-modality correlations, costly curated training resources, or insufficient causal assumptions and constraints,...

    arxiv.org/abs/2606.09181 · PDF

  16. 16

    Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models

    Danya Li, Xiang Su, Yan Feng, Rico Krueger

    cs.CV · cs.AI

    Egocentric vision offers a first-person view of human perception and decision making, yet its potential for traffic-safety prediction remains underexplored. In this work, we study the decoding of pedestrian crossing intentions from short egocentric video clips. We approach this by formulating the task as a closed-ended visual question answering (VQA) problem and leveraging vision language models (VLMs) to predict the pedestrians' intent. We...

    arxiv.org/abs/2606.09142 · PDF

  17. 17

    An Enhanced Geometric-Spectral Feature Learning Framework for Airborne Multispectral Point Cloud Classification

    Xian Li, Yanfeng Gu, Aleksandra Pižurica

    cs.CV · cs.AI

    Multispectral point cloud (MPC) is composed of 3D spatial-spectral information, which holds tremendous potential for accurate land-cover classification. However, the representation power of classification models is limited by inherent high-dimensional and heterogeneous spatial-spectral information, unbalanced sample distribution, and inter-class spectral similarity of airborne MPCs. We build two MPC datasets and propose an enhanced...

    arxiv.org/abs/2606.09123 · PDF

  18. 18

    Driving Video Retrieval for Complex Queries with Structured Grounding

    Manyi Yao, Sparsh Garg, Christian Shelton, Amit Roy-Chowdhury, Abhishek Aich

    cs.CV · cs.IR · cs.LG

    Video retrieval at scale is central to data curation and safety validation in autonomous driving, where users want to find not only scenes but also dynamic events such as cut-ins and hard braking. Existing vision-language and keyword-based retrieval methods often miss these events because the relevant motion may not be explicitly described in text or captured by lexical overlap. Rule-based retrieval can encode such events more directly, but...

    arxiv.org/abs/2606.09109 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.