cs.CV · 2026-07-14 · No. 53

Computer Vision and Pattern Recognition, 2026-07-14.

21 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

21 entries
  1. 01

    Evidence-Backed Video Question Answering

    Shijie Wang, Honglu Zhou, Ziyang Wang, Ran Xu, Caiming Xiong, Silvio Savarese, Chen Sun, Juan Carlos Niebles

    cs.CV · cs.AI

    Current Video Large Language Models (Video LLMs) excel in question answering (QA) but largely operate as black boxes, providing textual answers without verifiable visual grounding. Existing explainability efforts rely on textual rationales or sparse bounding boxes, which struggle to capture complex video dynamics such as occlusions and non-rigid deformations. We propose Evidence-Backed Video Question Answering (E-VQA), a novel task requiring...

    arxiv.org/abs/2607.11862 · PDF

  2. 02

    LoRA-Based Cascaded Multimodal Fusion for Action Recognition in Medical Training Environments

    Divya Mereddy, Jeevan Beedareddy

    cs.CV · cs.AI

    This paper presents a cascaded Low-Rank Adaptation (LoRA)-based multimodal fusion framework for action and activity recognition in healthcare-oriented training environments. The proposed architecture combines parameter-efficient modality-specific adaptation with sequential fusion, enabling modalities to be integrated in stages without retraining previously learned components. Rather than assuming a fixed fusion structure, the framework first...

    arxiv.org/abs/2607.11839 · PDF

  3. 03

    MM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling Agents

    Kaixin Ma, Di Feng, Alexander Metz, Jiarui Lu, Eshan Verma, Afshin Dehghan

    cs.CV · cs.AI

    We introduce MM-ToolSandBox, a benchmark and evaluation framework for visually grounded tool-calling agents. The framework provides a stateful execution environment spanning 500+ tools across 16 application domains, supporting multi-image, multi-turn tasks where agents must ground progressively arriving visual inputs into executable tool calls while handling realistic conversational phenomena (goal revisions, error corrections, state...

    arxiv.org/abs/2607.11818 · PDF

  4. 04

    StoryTeller: Training-Free Narrative Grounding for Long-Form Audio Description

    Seung Hyun Hahm, Minh T. Dinh, SouYoung Jin

    cs.CV · cs.AI

    Long-form audio description (AD) requires more than describing visible actions: it must preserve characters, events, relationships, and story context across scenes so that blind and low-vision (BLV) audiences can follow a film. Modern video-language models (VLMs) are effective on short clips, but they often treat each moment independently, producing descriptions that miss who characters are, why events matter, and how the current scene...

    arxiv.org/abs/2607.11798 · PDF

  5. 05

    Technical Report on the CVPR 2026@AdvML Workshop Challenge

    Tianyuan Zhang, Zonglei Jing, Jiangfan Liu, Ligong Zhang, Ke Ma, Chengzhi Sun, Xiaohai Xu, Zhirui Zhang, Qianqian...

    cs.CV · cs.AI

    Vision-language agents (VLAs) are increasingly used to interpret complex driving scenes and support safety-critical reasoning. This report presents the CVPR 2026@AdvML Workshop Challenge on adversarial multimodal attacks against autonomous-driving VLAs. Built on DriveLM-style multi-view visual question answering, the challenge represents each scene with six synchronized camera images and a structured collection of driving-related...

    arxiv.org/abs/2607.11560 · PDF

  6. 06

    Training-Free Off-Screen Player Imputation for Broadcast-Based Spatial Football Analytics

    Seongjin Choi

    cs.CV · cs.LG

    Spatial football metrics such as pitch control assume access to the positions of all 22 players, yet the most widely available source of positional data -- the broadcast main camera -- shows only 10-16 of them at any moment. We quantify the resulting distortion with an open, reproducible benchmark: a simulated broadcast viewport applied to open full-pitch tracking data (Metrica Sports; three matches, one held out from method development)....

    arxiv.org/abs/2607.11548 · PDF

  7. 07

    Adaptive Routing for Efficient Diffusion Transformer-Based PNI Prediction

    Youngung Han, Dohyun Kweon, Kyeonghun Kim, Hyunsu Go, Jina Jeong, Suah Park, Induk Um, Junga Kim, Anna Jung, Yului...

    cs.CV · cs.LG

    Perineural invasion (PNI) is a critical prognostic factor in cholangiocarcinoma. However, its preoperative prediction from magnetic resonance imaging (MRI) remains challenging due to subtle imaging features that extend beyond tumor boundaries into surrounding regions. Conventional convolutional neural networks are limited in capturing long-range spatial dependencies. Transformer-based architectures improve global modeling of volumetric MRI by...

    arxiv.org/abs/2607.11533 · PDF

  8. 08

    Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos

    Gong Sitong, Tianyu Yan, Caixin Kang, Bo Zheng, Xiang Ruan, Huchuan Lu, Kaipeng Zhang, Yoichi Sato, Yifei Huang

    cs.CV · cs.AI

    When should an intelligent assistant speak up without being asked? Continuous egocentric video offers rich, evolving context that enables a new form of assistance: one that is proactive rather than merely reactive. Yet existing approaches either wait passively for user queries or treat every detected event as requiring a response, without considering the user's history, current activity, or whether assistance would actually be welcome. We...

    arxiv.org/abs/2607.11523 · PDF

  9. 09

    Towards Efficient Convolutional Neural Network for Embedded Hardware via Multi-Dimensional Pruning

    Hao Kong, Di Liu, Xiangzhong Luo, Shuo Huai, Ravi Subramaniam, Christian Makaya, Qian Lin, Weichen Liu

    cs.CV · cs.LG

    In this paper, we propose TECO, a multi-dimensional pruning framework to collaboratively prune the three dimensions (depth, width, and resolution) of convolutional neural networks (CNNs) for better execution efficiency on embedded hardware. In TECO, we first introduce a two-stage importance evaluation framework, which efficiently and comprehensively evaluates each pruning unit according to both the local importance inside each dimension and...

    arxiv.org/abs/2607.11473 · PDF

  10. 10

    Uncertainty Quantification for EO Regression Tasks: Building Height, Tree Canopy Height and Above-ground Biomass Estimation

    Ritu Yadav, Andrea Nascetti, Yifang Ban

    cs.CV · cs.AI

    Earth Observation regression tasks such as building height, canopy height, and above-ground biomass estimation underpin critical applications in urban planning, forest monitoring, and climate policy, where both accuracy and reliability are critical. Yet most deep learning models yield only deterministic predictions, providing no indication of per-pixel reliability. These regression tasks are inherently challenging due to heterogeneous land...

    arxiv.org/abs/2607.11412 · PDF

  11. 11

    Longitudinal Multi-View Breast Cancer Risk Prediction

    Solveig Thrun, Zijun Sun, Suaiba A. Salahuddin, Kristoffer Wickstrøm, Elisabeth Wetzer, Stine Hansen, Robert...

    cs.CV · cs.AI

    Accurate breast cancer risk prediction from screening mammography is critical for enabling personalized screening intervals and early detection. Recent deep learning methods have shown the value of longitudinal data and explicit temporal alignment. However, existing approaches either perform explicit alignment using a single mammographic view or model multiple views without explicit longitudinal alignment, limiting their ability to exploit...

    arxiv.org/abs/2607.11343 · PDF

  12. 12

    A Unified Framework for Comprehensive Cardiac CT Segmentation and Phenotyping: Human-in-the-Loop Data Annotation, Vision Foundation Model Development, Multicenter Evaluation and Clinical Validation

    Pooya Mohammadi Kazaj, Leo Fridolin Weber, Wen Xie, Seyed Amir Ahmad Safavi-Naini, Anselm Stark, Giovanni Baj, Ali...

    cs.CV · cs.AI

    Comprehensive quantification of cardiac structures from computed tomography (CT) remains limited not by data availability but by the scalability of measurements, which makes routine use impractical. Here we present a unified framework for comprehensive cardiac CT segmentation and phenotyping that combines a human-in-the-loop annotation pipeline, a cardiac CT augmentation technique, and a self-supervised foundation model pre-trained on 60,000...

    arxiv.org/abs/2607.11287 · PDF

  13. 13

    LaGuadia: Language-Guided Adaptive Distillation from Pathology Foundation Models

    Gangsu Kim, Won-Ki Jeong

    cs.CV · cs.LG

    Pathology Foundation Models (PFMs) offer powerful Whole Slide Image (WSI) representations but suffer from massive computational costs. While Knowledge Distillation (KD) can create efficient student models, existing multi-teacher methods often use suboptimal uniform weighting that ignores tissue heterogeneity. We propose LaGuadia (Language-Guided Adaptive DistillAtion), a framework that develops a compact pathology image encoder by dynamically...

    arxiv.org/abs/2607.11257 · PDF

  14. 14

    HandFlow: Fully Generative 4D Hand Recovery with Flow Matching

    Mingxi Xu, Bowen Duan, Yi Gu, Zhengyang Shen, Renjing Xu, Yutao Yue

    cs.CV · cs.AI

    Accurate monocular 4D hand reconstruction remains challenging. Per-frame discriminative regressors lack temporal context and often produce jittery predictions. Temporal models improve consistency by aggregating information across frames, but they are typically deterministic regressors, making them vulnerable to ambiguous observations caused by occlusion and motion blur. Generative modeling offers a natural alternative by learning a prior over...

    arxiv.org/abs/2607.11221 · PDF

  15. 15

    Controlling Motion Transfer in Diffusion Transformers via Attention Heads

    Sunyoung Jung, Jiwoo Park, Yoonseok Choi, Kyobin Choo, Ming-Hsuan Yang, Seong Jae Hwang

    cs.CV · cs.AI

    Diffusion Transformers (DiTs) have advanced video generation with high-quality, temporally coherent results. However, extending them to motion transfer, which requires following reference motion while aligning with a target prompt, remains challenging due to limited understanding of motion and structure representations within DiTs. We analyze video DiTs at the attention-head level and identify distinct heads specialized for motion and spatial...

    arxiv.org/abs/2607.11081 · PDF

  16. 16

    Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video

    Mohammad Al-Ratrout, Shayla Sharmin, Aditya Raikwar, Roghayeh Leila Barmaki

    cs.CV · cs.AI

    Can a Video Large Language Model (Video-LLM) follow one person through a long video, keeping track of who they are well enough to report, in order, how their outfit changes across a full TV episode? Benchmarks increasingly score this kind of task, and the strongest open-source 7--8B models now reach 37--38% on InfiniBench's global appearance task, which asks exactly that. But does that score come from tracking the named character, or from...

    arxiv.org/abs/2607.11078 · PDF

  17. 17

    Reference-Based Face Super-Resolution Using the Spatial Transformer

    Varun Ramesh Jois, Antonella DiLillo, James Storer

    cs.CV · cs.LG

    Face super-resolution is the task of increasing the resolution of an image containing a face thereby adding finer detail. It is a ubiquitous task in many computer vision applications and quite often the user isn't even aware that it is being performed. However, doing it with high fidelity is challenging as it is an ill-posed problem. In this paper we present a reference-based solution for face super-resolution that uses higher resolution...

    arxiv.org/abs/2607.11025 · PDF

  18. 18

    SynCLIP: Synonym-Coherent Language-Image Pretraining for Robust Open-Vocabulary Dense Perception

    Mingjie Xie, Guangjun He, Dongli Xu, Youtian Lin, Hongjue Li, Pengming Feng, Jian Guan, Yue Deng

    cs.CV · cs.AI

    Open-vocabulary dense perception (OVDP) aims to localize objects unseen during training by leveraging textual knowledge. Despite the remarkable progress of recent CLIP-based approaches, we identify a critical limitation: synonym-induced grounding inconsistency, where semantically equivalent expressions yield disparate spatial attention patterns. This inconsistency undermines the robustness and performance of existing methods in real-world...

    arxiv.org/abs/2607.11008 · PDF

  19. 19

    LoSA-Net: A Localized and Scale-Adaptive Network for Boundary-Sensitive Prediction of Perineural Invasion in 3D MRI

    Youngung Han, Hyunsu Go, Kyeonghun Kim, Induk Um, Junga Kim, Jaewon Jung, Woo Kyoung Jeong, Won Jae Lee, Pa Hong,...

    cs.CV · cs.AI

    Perineural invasion (PNI) is a clinically relevant indicator of tumor aggressiveness and can influence surgical decision-making, motivating interest in reliable preoperative assessment. The subtle MRI features of PNI, however, often resemble nearby anatomy, complicating noninvasive prediction. These fine perineural cues are easily attenuated by routine downsampling or overly global feature aggregation, reducing the effectiveness of...

    arxiv.org/abs/2607.10992 · PDF

  20. 20

    MMA-Former: Multi-Window Mixture-of-Head Attention Transformer for Adaptive PNI Prediction in 3D MRI

    Youngung Han, Induk Um, Kyeonghun Kim, Junga Kim, Hyunsu Go, Jaewon Jung, Woo Kyoung Jeong, Won Jae Lee, Pa Hong,...

    cs.CV · cs.AI

    Perineural invasion (PNI) is a critical prognostic factor in cholangiocarcinoma. Non-invasive prediction from 3D MRI is challenging, demanding models that efficiently capture both fine-grained details and global context. We propose the Multi-window Mixture-of-Head Attention Transformer (MMA-Former), a novel end-to-end 3D architecture featuring a Coarse-Fine Transformer (CFT) structure for parallel multi-scale feature extraction. We advance...

    arxiv.org/abs/2607.10988 · PDF

  21. 21

    EquiFusion: Kinematics-Agnostic Human Motion Prediction via Equivariant Latent Diffusion

    Cecilia Curreli, Florian Hofherr, Dominik Muhle, Abhishek Saroha, Riccardo Marin, Daniel Cremers

    cs.CV · cs.AI · cs.HC · cs.LG

    Existing Stochastic 3D Human Motion Prediction models are fundamentally constrained by hard-coding the skeleton kinematics, severely limiting generalization, preventing cross-dataset training, and requiring complex data retargeting. We introduce EquiFusion, the first kinematics-agnostic model to solve this bottleneck, implementing a latent diffusion model with a permutation equivariant architecture. EquiFusion treats the kinematics'...

    arxiv.org/abs/2607.10984 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.