cs.CV · 2026-06-11 · No. 20

Computer Vision and Pattern Recognition, 2026-06-11.

29 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

29 entries
  1. 01

    Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models

    Cheng-Yu Yang, Shao-Yuan Lo, Yu-Lun Liu

    cs.CV · cs.AI

    Vision-language models (VLMs) project images into hundreds to thousands of visual tokens, making decoder inference expensive in both attention computation and KV-cache memory. Existing visual-token reduction methods largely follow a rank-and-remove paradigm: they score visual tokens, keep a compact subset, and permanently discard the rest. We show that this irreversible action is fragile because visual-token importance changes across decoder...

    arxiv.org/abs/2606.12412 · PDF

  2. 02

    Illumination-Robust Camera-Based Heart-Rate Estimation for Physiological Sensing in Robots

    Zhi Wei Xu, Torbjörn E. M. Nordling

    cs.CV · cs.AI

    Physiological awareness is important for service, social, and assistive robots that interact with humans in everyday environments. Remote photoplethysmography (rPPG) enables non-contact heart-rate (HR) estimation from an RGB camera, making it a promising sensing modality for robot-mounted vision systems. However, illumination variation remains a major barrier to robust deployment. This paper presents an end-to-end spatial-temporal transformer...

    arxiv.org/abs/2606.12378 · PDF

  3. 03

    Atlas H&E-TME: Scalable AI-Based Tissue Profiling at Expert Pathologist-Level Accuracy

    Kai Standvoss, Miriam Hägele, Rosemarie Krupar, Julika Ribbat-Idel, Jennifer Altschüler, Gerrit Erdmann, Hans...

    cs.CV · cs.AI · cs.LG

    Hematoxylin and eosin (H&E) staining is the cornerstone of histopathology, yet scalable, quantitative analysis of H&E whole-slide images (WSIs) remains a central challenge in computational pathology. We present Atlas H&E-TME, an AI-based system built on the Atlas family of pathology foundation models that predicts tissue quality, tissue region, and cell type labels across multiple cancer types, yielding over 4,500 quantitative readouts per...

    arxiv.org/abs/2606.12346 · PDF

  4. 04

    Natural-Language Temporal Grounding in Hour-Long Videos is a Search Problem: A Benchmark and Empirical Decomposition

    Sukmin Seo, Geewook Kim

    cs.CV · cs.AI

    Temporal grounding--returning the interval $[t_s, t_e]$ for a natural-language query over a video--is the language interface to long-form video, yet has been studied on short videos; the dynamics of hour-scale natural-language grounding remain underexplored. We take the position that at hour-scale, the binding constraint is search, not recognition: Video-LLMs are bottlenecked not by localizing a nearby event, but--given a natural-language...

    arxiv.org/abs/2606.12300 · PDF

  5. 05

    Finding Sparse Subnetworks in One Training Cycle via Progressive Magnitude-Based Pruning

    Romana Qureshi, Hafida Benhidour, Said Kerrache, Nahlah Aljeraisy

    cs.CV · cs.LG

    Neural network pruning reduces model size by removing less important parameters while aiming to preserve predictive performance. Although the Lottery Ticket Hypothesis (LTH) shows that sparse subnetworks can match dense networks when trained from suitable initializations, its iterative pruning procedure requires multiple complete training cycles. This work evaluates progressive magnitude-based pruning as a single-cycle alternative. The method...

    arxiv.org/abs/2606.12278 · PDF

  6. 06

    Adapting Prithvi-EO for Fallow Detection for Food-Water Nexus: ViT-Adapter Necks and Parameter-Efficient Backbone tuning of Geospatial Foundation Model

    Sk Muhammad Asif, Orhun Aydin

    cs.CV · cs.AI

    Understanding spatial distribution of fallow land is important for optimizing the food-water (FW) nexus, given fallowing's role in crop rotation and water conservation. Fallow is a low accuracy class in USDA Cropland Data Layer (CDL). Geospatial foundation model (GFM), Prithvi-EO has shown strong transferability across computer vision tasks. However, its Vision Transformer (ViT) backbone produces features at a single spatial scale that are...

    arxiv.org/abs/2606.12218 · PDF

  7. 07

    Making Foresight Actionable: Repurposing Representation Alignment in World Action Models

    Lu Qiu, Yizhuo Li, Yi Chen, Yuying Ge, Yixiao Ge, Xihui Liu

    cs.CV · cs.AI · cs.RO

    World Action Models (WAMs) offer a promising route for robot manipulation by using video generation models to model future scene evolution before producing control actions. However, our empirical observations reveal a phenomenon: generating plausible visual futures does not always guarantee the extraction of accurate actions. To diagnose this failure, we conduct action-head attention analysis and causal interventions. We find that the action...

    arxiv.org/abs/2606.12217 · PDF

  8. 08

    MLT-Dedup: Efficient Large-Scale Online Video Deduplication via Multi-Level Representations and Spatial-Temporal Matching

    David Yuchen Wang, Haoying Li, Hailun Xu, Wei Chee Yew, Zirui Zhu, Sanjay Saha, Hao Hei, Kanchan Sarkar, Kun Xu

    cs.CV · cs.IR · cs.LG

    The explosive growth of user-generated video content on online platforms is accompanied by the emergence of numerous near-duplicate videos--videos that are identical or highly similar but differ by partial edits. These duplicates degrade user experience and increase storage and bandwidth costs, making large-scale video deduplication a critical task. Existing video deduplication frameworks face a fundamental challenge in retrieving sufficient...

    arxiv.org/abs/2606.12215 · PDF

  9. 09

    Beyond Dark Knowledge: Mixup-Based Distillation for Reliable Predictions

    José Medina, Paul Honeine, Abdelaziz Bensrhair, Amnir Hadachi

    cs.CV · cs.LG

    Knowledge Distillation (KD) and mixup have proven effective at inducing smoothness in class boundaries; KD captures inherent class relationships in probability distributions, and mixup enforces them through convex combinations of inputs. Their interaction, however, remains poorly understood, particularly when mixup is applied only during student training. In this setting, the teacher is queried on inputs drawn from a vicinal distribution it...

    arxiv.org/abs/2606.12171 · PDF

  10. 10

    OpenMedReason: Scientific Reasoning Supervision for Medical Vision-Language Models

    Negin Baghbanzadeh, Pritam Sarkar, Michael Colacci, Abeer Badawi, Adibvafa Fallahpour, Arash Afkanpour, Leonid...

    cs.CV · cs.AI · cs.CL · cs.LG

    High-stakes clinical use of large vision-language models (LVLMs) requires reasoning that is grounded in visual evidence and clinical knowledge, not just correct final answers. We introduce OpenMedReason, a large-scale, open multimodal medical reasoning corpus comprising approximately 450K image-question-answer instances whose reasoning traces are primarily derived from curated biomedical, human-authored scientific articles. OpenMedReason...

    arxiv.org/abs/2606.12169 · PDF

  11. 11

    MSUE: Multi-Modal Soccer Understanding Expert

    Litao Li, Yibo Yu, Yufeng Hu, Zhuo Yang, Jiali Wen, Yixin Chen, Yixi Zhou

    cs.CV · cs.AI

    This paper presents our solution to the 2026 SoccerNet VQA Challenge. We first develop a cost-effective data synthesis pipeline driven by a Vision-Language Model (VLM), which systematically restructures raw domain data into diverse VQA samples, including concise answers and long-form responses. Second, we propose MSUE, a multi-expert question answering architecture that employs a Large Language Model (LLM) to dynamically dispatch questions to...

    arxiv.org/abs/2606.12106 · PDF

  12. 12

    Non-frontal face recognition using GANs and memristor-based classifiers

    Semih Vazgecen, Cristian Sestito, Spyros Stathopoulos, Themis Prodromakis

    cs.CV · cs.AI · eess.IV

    Face recognition systems have advanced significantly through deep learning techniques, delivering high performance and robustness in complex scenarios. However, these approaches incur substantial computational overhead, limiting their in situ applicability in resource-constrained platforms such as drones, where they can address challenges including non-frontal facial imagery. Memristor-based neuromorphic systems have emerged as a compelling...

    arxiv.org/abs/2606.12074 · PDF

  13. 13

    Metadata-Aware Multi-Prompt Reasoning for Zero-Shot Accident Understanding

    Tarandeep Singh, Soumyanetra Pal, Soham Biswas, Nishanth Chandran

    cs.CV · cs.AI · stat.ML

    In this paper, we address the problem of zero-shot understanding of accidents from surveillance videos by identifying when an impact event occurs, what type of impact it is, and where in the frame it occurs using natural language. We propose a three-stage pipeline that decomposes the accident understanding into when, what, and where. The first stage extracts a short temporal window around the impact using vision-language similarity. In the...

    arxiv.org/abs/2606.12047 · PDF

  14. 14

    Corpus Augmentation for Sign Language Translation via LLM-Guided Video Stitching

    Zsolt Robotka, Ádám Rák, Jalal Al-Afandi, András Horváth, György Cserey

    cs.CV · cs.LG

    Sign language translation (SLT) converts sign language video into spoken language text and holds significant promise for improving accessibility and enabling communication between signing and non-signing communities. While large weakly-aligned datasets have enabled pre-training at scale and gloss-free methods have reduced reliance on expert annotation, high-quality parallel sign video-text pairs for fine-tuning remain scarce, limiting...

    arxiv.org/abs/2606.11925 · PDF

  15. 15

    Task-Aligned Stability Analysis of Vision-Language Models for Autonomous Driving Hazard Detection

    Everett Richards

    cs.CV · cs.AI · cs.RO

    Vision-language models (VLMs) are increasingly used for scene understanding in autonomous driving, but robustness analysis often relies on task-agnostic embedding stability alone. We study whether corruption-induced embedding drift predicts changes in a task-aligned hazard score derived from CLIP image-text similarities. Using controlled corruptions on BDD100K road scenes, we compare embedding drift against margin drift, defined as the change...

    arxiv.org/abs/2606.11889 · PDF

  16. 16

    Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning

    Zhirui Chen, Ziwei Chen, Ling Shao

    cs.CV · cs.AI

    Multi-modal large language models (MLLMs) depend on in-context learning (ICL) for rapid task adaptation, but their scalability is severely limited by finite context windows and the growing cost of key-value (KV) caches in long multi-modal sequences. Existing memory compression approaches typically rely on rigid token removal or sample-dependent importance estimation, which introduces bias, disrupts semantic structure, particularly for visual...

    arxiv.org/abs/2606.11853 · PDF

  17. 17

    LASA: A Weak Supervision Method for Open-Vocabulary Scene Sketch Semantic Segmentation

    Liwen Yi, Xianlin Zhang, Yue Zhang, Yue Ming, Xueming Li

    cs.CV · cs.AI

    Open-vocabulary scene sketch semantic segmentation aims to assign dense semantic labels to sparse line drawings based on flexible category vocabularies specified at inference time, without relying on pixel-level annotations during training. Unlike natural images, sketches lack texture and color cues, making semantic understanding heavily dependent on stroke layout and spatial configuration, a challenge that renders single-layer...

    arxiv.org/abs/2606.11837 · PDF

  18. 18

    TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization

    Zixiong Hao, Zhencun Jiang

    cs.CV · cs.AI

    Text-conditioned 3D generation has progressed rapidly for images and isolated objects, but producing a hand-object mesh remains challenging: the output must preserve language semantics, cross-view consistency, object geometry, articulated hand shape, and physically plausible contact. We present TextHOI-3D, a staged framework that uses generated multi-view observations as an explicit interface between text-conditioned visual generation and...

    arxiv.org/abs/2606.11805 · PDF

  19. 19

    MultiToP: Learning to Patch Visual Tokens to Mitigate Hallucinations in Video Large Multimodal Models

    Yuansheng Gao, Wenbin Xing, Jiahao Yuan, Kaiwen Zhou, Han Bao, Zonghui Wang, Wenzhi Chen

    cs.CV · cs.AI · cs.CL

    Video Large Multimodal Models have achieved remarkable progress in video understanding, yet they remain prone to hallucinations, where generated responses are not faithfully supported by the input video. In this paper, we propose MultiToP, a multimodal-context-aware visual token patching framework that mitigates hallucinations by refining unreliable visual tokens before language generation. MultiToP introduces a lightweight Visual Token...

    arxiv.org/abs/2606.11792 · PDF

  20. 20

    AnchorEdit: Maintaining Temporal Consistency in Multi-turn Image Editing via Causal Memory

    Hang Xu, Xiaoxiao Ma, Guohui Zhang, Yu Hu, Siming Fu, Jie Huang, Lin Song, Haoyang Huang, Nan Duan, Feng Zhao

    cs.CV · cs.AI

    Multi-turn image editing is essential for iterative design, yet current models often struggle with identity drift and error accumulation over successive steps. While existing research leverages video priors for consistency, their reliance on bidirectional attention is fundamentally misaligned with the causal, sequential nature of interactive editing. In this paper, we propose AnchorEdit, the first autoregressive (AR) diffusion-based framework...

    arxiv.org/abs/2606.11751 · PDF

  21. 21

    From Prompts to Tokens: Internalizing Causal Supervision in Vision-Language Model for Multi-Image Causal Reasoning

    Haoping Yu, Yuanxi Li, Jing Ma

    cs.CV · cs.AI

    Visual causal reasoning is essential for understanding and intervening in the physical world, requiring identification of causal variables from visual inputs and reasoning over intervention effects. Despite recent progress, large vision--language models (VLMs) remain brittle at such tasks, especially for interventional and counterfactual queries over multi-image inputs. Most existing explorations inject causal knowledge via textual prompts,...

    arxiv.org/abs/2606.11745 · PDF

  22. 22

    Multi-View In-Cabin Monitoring System for Public Transport Vehicles

    Evgeny Gorelik, Kenny Dean Karrow, Fikret Sivrikaya, Sahin Albayrak, Christian Baumann

    cs.CV · cs.AI

    We introduce a multi-view in-cabin monitoring dataset for public transportation with synchronized RGB and depth images from four inward-facing cameras and a rotating LiDAR covering the vehicle interior of a digitalized and partly automated German city bus. The dataset contains 9.136 synchronized samples with annotations and is accompanied by a calibration and pseudo-labeling pipeline that generates 3D human pose estimates and oriented 3D...

    arxiv.org/abs/2606.11739 · PDF

  23. 23

    Ouroboros-Spatial: Closing the Data-Model Loop for Spatial Reasoning

    Enhan Zhao, Wei Wu, Yuanrui Zhang, Xueliang Zhao, Di He

    cs.CV · cs.AI

    Spatial reasoning remains a persistent challenge for multimodal large language models (MLLMs). Existing approaches largely rely on large-scale, statically curated datasets, where all training samples are treated uniformly regardless of the model's evolving capabilities. This static paradigm is inherently data-inefficient: training capacity is often spent on samples that are either trivial or overly difficult for the model at its current...

    arxiv.org/abs/2606.11719 · PDF

  24. 24

    MedCTA: A Benchmark for Clinical Tool Agents

    Tajamul Ashraf, Hyewon Jeong, Fida Mohammad Thoker, Bernard Ghanem

    cs.CV · cs.AI · cs.CL

    To make clinically grounded decisions, medical AI agents are expected to go beyond simple recognition and be capable of tool retrieval, evidence acquisition, and integration. Existing benchmarks largely evaluate isolated perception or single-turn question answering, and therefore provide limited visibility into failures of planning, tool recruitment, and rollout reliability. We introduce MedCTA, a benchmark for evaluating medical tool agents...

    arxiv.org/abs/2606.11702 · PDF

  25. 25

    DroneShield-AI: A Multi-Modal Sensor Fusion Framework for Real-Time Autonomous Drone Threat Detection, Behavioral Intent Classification, and Swarm Intelligence in Contested Airspace

    Marius Bayizere

    cs.CV · cs.LG · cs.RO

    Unmanned Aerial Vehicle (UAV) threats have emerged as a defining security challenge of the 21st century. This paper presents DroneShield-AI, a unified open framework integrating six processing layers: RF signal classification, acoustic motor-signature detection, YOLOv8-based visual detection, evidence-weighted sensor fusion, a Behavioral Intent Classification Engine (BICE), and a Graph Neural Network Swarm Intelligence Module (GNN-SIM). BICE...

    arxiv.org/abs/2606.11687 · PDF

  26. 26

    Reason, Then Re-reason: Cross-view Revisiting Improves Spatial Reasoning

    Chaofan Ma, Zhenjie Mao, Yuhuan Yang, Fanqin Zeng, Yue Shi, Yingjie Zhou, Xiaofeng Cao, Jiangchao Yao

    cs.CV · cs.AI

    Spatial reasoning from egocentric videos is inherently challenging because the observable evidence is constrained by the camera trajectory. Existing methods rely on single-turn inference, forcing models to resolve geometric ambiguity through semantic priors rather than verifiable evidence. We argue that spatial reasoning should be revisitable: conclusions formed under limited evidence should remain open to revision when complementary...

    arxiv.org/abs/2606.11683 · PDF

  27. 27

    Parameter-Efficient Adapter Tuning for Tabular-Image Multimodal Learning

    Jiaqi Luo

    cs.CV · cs.LG

    Tabular-image multimodal learning aims to improve predictive modeling by jointly using structured tabular attributes and visual data. Although pretrained encoders provide strong modality-specific representations, full fine-tuning can be computationally expensive, while keeping encoders frozen may limit task-specific adaptation. We propose the Tabular-Image Adapter (TI-Adapter), a modality-specific adapter-based fine-tuning framework for...

    arxiv.org/abs/2606.11682 · PDF

  28. 28

    ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation

    Zijie Meng, Jiwen Liu, Yufei Liu, Chengzhuo Tong, Xiaoqiang Liu, Yuanxing Zhang, Yulong Xu, Pengfei Wan

    cs.CV · cs.AI

    Subject-preserving video generation is not solved by frontal-face similarity alone: a generated person must remain recognizable across motion, large viewpoint changes, expression shifts, occlusion, scale variation, and conflicts among text, first-frame, and identity references. We argue that the central bottleneck is the point-reference paradigm, which collapses identity into a single static observation entangled with pose, accessories,...

    arxiv.org/abs/2606.11670 · PDF

  29. 29

    Learning Instance-Adaptive Low-Rank Orthogonal Subspaces for Clothes-Changing Person Re-Identification

    Dong-Woo Kim, Tae-Kyun Kim

    cs.CV · cs.LG

    Clothes-changing person re-identification (CC-ReID) aims to recognize individuals despite drastic appearance changes caused by clothing variation. While existing methods rely on adversarial learning to disentangle clothing features, we propose Ortho-ReID, which explicitly models a low-rank clothing subspace from VLM text descriptions and extracts clothing-invariant representations via direct geometric constraints. A critical component is our...

    arxiv.org/abs/2606.11661 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.