cs.CV · 2026-07-20 · No. 59

Computer Vision and Pattern Recognition, 2026-07-20.

27 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

27 entries
  1. 01

    An Exam for Active Observers

    Jiarui Zhang, Muzi Tao, Shangshang Wang, Ollie Liu, Xuezhe Ma, Willie Neiswanger

    cs.CV · cs.AI · cs.CL · cs.LG

    Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapshot. Decades of psychophysics and cognitive science have argued that this active observation is essential for a wide range of tasks. Whether today's multimodal large language models (MLLMs) exercise active observation is an empirical question that current vision-language benchmarks do not answer. We introduce ActiveVision, a...

    arxiv.org/abs/2607.16165 · PDF

  2. 02

    HCIG: A Hierarchical Cross-Modal Incongruity Graph Network for Multimodal Sarcasm and Cyberbullying Detection

    Bhavana Verma, Priyanka Meel, Dinesh Kumar Vishwakarma

    cs.CV · cs.AI · cs.CL

    Multimodal sarcasm and cyberbullying detection remain challenging because the intended meaning often emerges from incongruity between textual and visual information rather than from either modality alone. Existing multimodal approaches primarily rely on feature fusion or cross-modal attention, which may not effectively capture hierarchical semantic inconsistencies across different levels of representation. To address this limitation, this...

    arxiv.org/abs/2607.16076 · PDF

  3. 03

    Spatial Normalization for Cross-Domain Retinal Layer Segmentation in Optical Coherence Tomography

    Iker Moran-Cavero, Monica Hernandez, Elvira Mayordomo, Naiara Artiaga, Beatriz Pardiñas, Beatriz Cordon, Elena Garcia-Martin

    cs.CV · cs.AI

    Retinal layer segmentation in Optical Coherence Tomography (OCT) is a fundamental step for extracting quantitative biomarkers of retinal structure. Indeed, there is a growing interest in the analysis of OCTs in the context of neurodegenerative diseases. However, segmentation remains challenging due to speckle noise, shadowing artifacts, low contrast between adjacent layers, anatomical variability across subjects, and domain shifts arising...

    arxiv.org/abs/2607.16065 · PDF

  4. 04

    DPNeXt: A Lightweight Multi-Scale Feature Fusion Framework for Efficient ViT-Based Multi-Task Dense Prediction

    Jehun Kang, Jungha Wang, Youngjun Hwang, David Hyunchul Shim

    cs.CV · cs.AI · cs.RO

    Multi-Task Learning (MTL) in robotics perception systems supports comprehensive 3D spatial scene understanding by integrating semantic segmentation and depth estimation. While Vision Foundation Models (VFMs) are increasingly adopted as robust feature encoders, existing decoding strategies present a critical bottleneck. To address this, we propose DPNeXt, a streamlined multi-scale feature fusion decoder and efficient alternative to the...

    arxiv.org/abs/2607.16012 · PDF

  5. 05

    CanonicalPhys: Pose-Robust Remote Photoplethysmography via Canonical-Space Priors

    Hui Wei, Seyedata Jodeiri Seyedian, Xiaobai Li, Guoying Zhao

    cs.CV · cs.LG

    Deep remote photoplethysmography (rPPG) attains sub-bpm heart-rate error on frontal, stationary faces yet degrades sharply under head pose: on MMPD, the state-of-the-art FactorizePhys backbone's MAE grows $1.60\times$ from frontal ($|\text{yaw}|{<}15^\circ$) to large-yaw ($|\text{yaw}|{\geq}45^\circ$) frames. We argue that pose is a \emph{coordinate-structural} nuisance rather than a data-augmentation problem: in image coordinates the same...

    arxiv.org/abs/2607.15995 · PDF

  6. 06

    More with Less: a Large Scale Remote Sensing VLM with a Simple Recipe

    Stefan Maria Ailuro, Mario Markov, Mohammad Mahdi, Luc Van Gool, Danda Pani Paudel

    cs.CV · cs.LG

    Remote sensing vision-language models are increasingly expected to support open-ended reasoning over Earth Observation data and a variety of tasks. Most recent progress in this area has been driven by remote-sensing-specific architectural designs, often introducing new encoders, alignment modules, or task-specific fusion mechanisms. In this work, we challenge the necessity of such architectural specialization. We show that a generally capable...

    arxiv.org/abs/2607.15942 · PDF

  7. 07

    Orbis 2: A Hierarchical World Model for Driving

    Sudhanshu Mittal, Arian Mousakhan, Silvio Galesso, Karim Farid, Jonannes Dienert, Rajat Sahay, Thomas Brox

    cs.CV · cs.AI · cs.LG · cs.RO

    Current world models operate at a single level of abstraction, with most prioritizing perceptual fidelity while lacking the spatial reasoning and semantic understanding required for real-world downstream tasks. We present a hierarchical driving world model that factorizes future prediction across two levels operating at distinct temporal and abstraction scales: a high-level predictor that forecasts coarse scene structure over extended...

    arxiv.org/abs/2607.15898 · PDF

  8. 08

    EgoExoMoCap: Distributed Ego-Exo Human Motion Capture

    Jiaxi Jiang, Bharat Lal Bhatnagar, Nan Yang, Lingni Ma, Sebastian Starke, Robin Kips, Nadine Bertsch, Christian...

    cs.CV · cs.AI · cs.GR · cs.HC · cs.RO

    Human motion capture from head-mounted devices (HMDs) offers a scalable way to acquire real-world human motion and interaction data, which is crucial for applications in embodied AI and VR/AR. Existing approaches focus on either egocentric body tracking, estimating the motion of the subject wearing the device, or exocentric tracking, capturing the movements of people in the wearer's surroundings. So far, these two paradigms have largely been...

    arxiv.org/abs/2607.15868 · PDF

  9. 09

    Von Mises-Fisher Mixture Model with Dynamic Shrinkage for Realistic Test-Time Transduction

    Jiazhen Huang, Zhiming Liu, Changhu Wang, Wei Ju, Ziyue Qiao, Xiao Luo

    cs.CV · cs.LG

    A range of methods aim to enhance the performance of vision-language models (VLMs) at test time. Among them, transduction has emerged as a promising paradigm due to its strong compatibility and efficiency. However, realistic evaluations often involve highly imbalanced class distributions, which cause performance degradation or even collapse. In this work, we systematically revisit transduction from the perspective of penalized likelihood...

    arxiv.org/abs/2607.15851 · PDF

  10. 10

    Test-Time Noise Guided Adaptation for Realistic Autoregressive Video Generation

    Dimitrios Karageorgiou, Symeon Papadopoulos, Ioannis Kompatsiaris, Efstratios Gavves

    cs.CV · cs.AI

    Autoregressive video diffusion models have enabled the generation of arbitrarily long videos by removing conditioning on future frames, thus greatly improving computational efficiency. Yet, they suffer from error accumulation over time, as the denoised sequence gradually drifts away from the conditioning distribution seen during training. Recent advances attempt to reduce this error by anchoring each generated frame to the learned manifold of...

    arxiv.org/abs/2607.15849 · PDF

  11. 11

    On the Geometry of Learned Representations in Event-Based Multi-Modal Egomotion Estimation

    Stefano Silvestrini, Michele Ceresoli

    cs.CV · cs.AI

    Classical approaches to event-based egomotion estimation, including those adopted by the top-performing teams of the ELOPE challenge, rely on geometric optimization frameworks such as contrast maximization, homography estimation, or dense optical flow combined with analytic motion inversion. This work investigates the geometric structure that emerges inside a multi-modal network for egomotion estimation. Event tensors, inertial measurements,...

    arxiv.org/abs/2607.15794 · PDF

  12. 12

    Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding

    Wei Feng, Xin Wang, Yu-Wei Zhan, Yuwei Zhou, Wenwu Zhu

    cs.CV · cs.AI

    Video Large Language Models (Video LLMs) have made significant advancements in various video understanding tasks. However, long-video scenarios remain challenging due to the tension between limited visual token budgets and the need to capture multiple key events. Existing approaches typically process long videos in two stages, i.e., i) select keyframes and ii) perform detailed perception, which exhibit limitations: they lack a modular...

    arxiv.org/abs/2607.15778 · PDF

  13. 13

    GeoChrono: Benchmarking and Rethinking Long-Term Temporal Understanding in Remote Sensing

    Yujie Li, Jiancheng Pan, Zhiwei Wei, Jiuniu Wang, Mugen Peng, Wenjia Xu

    cs.CV · cs.AI

    Remote sensing offers an unparalleled vantage point for observing the Earth's long-term surface evolution, yet it demands that a model not only perceive land cover at isolated moments, but also track changes, memorize evolution histories, and reason across time and space. However, existing studies lack a systematic evaluation that dissects these distinct competencies. To fill this gap, we introduce ChronoBench, a multidimensional benchmark...

    arxiv.org/abs/2607.15768 · PDF

  14. 14

    Debiasing Text-to-Image Evaluation via Implicit Cultural Alignment Reward Modeling

    Bo-An Chang, Yu-Chih Chen

    cs.CV · cs.AI · cs.LG · cs.MM

    As Text-to-Image (T2I) systems rapidly advance, evaluating the cultural authenticity of synthesized content has become increasingly important for fair and trustworthy generative AI. Existing T2I evaluation metrics and multimodal judges often rely on visual-semantic representations that underrepresent implicit cultural norms, leading to biased preference judgments and the omission of fine-grained cultural cues. In addition, visual question...

    arxiv.org/abs/2607.15740 · PDF

  15. 15

    Efficient Difficulty-Aware Dynamic Routing for Diffusion-Based Real-World Image Super-Resolution

    Xue Wu, Kang Zhao, Kafeng Wang, Jianfei Chen, Jingwei Xin, Nannan Wang, Xinbo Gao

    cs.CV · cs.AI

    Diffusion-based methods have achieved impressive performance in real-world image super-resolution (Real-ISR) by leveraging large pre-trained stable diffusion (SD) models as powerful generative priors. However, these methods still face two key limitations. First, existing SD-based one-step and multi-step Real-ISR approaches adopt a unified processing paradigm for all input samples, ignoring the varying restoration difficulty across images....

    arxiv.org/abs/2607.15711 · PDF

  16. 16

    Hierarchical Specialised Ensembles for Classification of Zebrafish Phenotypes Using the Selected Image Recognition Methods

    Piotr S. Maciąg, Monika Maciąg, Magdalena Majdan

    cs.CV · cs.LG

    We propose and evaluate three hierarchical ensemble setups for zebrafish phenotype classification from embryo images. In all setups, stage 1 uses a single four-class classifier to assign images to one of the exclusive phenotypes: Normal, Chorion, Dead, or Other. Images classified as Other are then processed in stage 2, where the ensemble design differs across setups: a single multi-label classifier, two specialized multi-label classifiers, or...

    arxiv.org/abs/2607.15698 · PDF

  17. 17

    Ask Twice, Look Twice: Prompt Echoing Resolves the Question-First Paradox in Vision-Language Models

    Rakshanda Hassan Abhinandan, John Galeotti, Deva Ramanan, Gautam Rajendrakumar Gare

    cs.CV · cs.AI · cs.LG · eess.IV

    Where should the question go in a vision-language model (VLM) prompt: before the image or after it? Intuition says before: knowing what is asked should tell the model where to look. Yet across visual question answering benchmarks, question-first prompting consistently underperforms the image-first ordering recommended for frontier VLMs, a phenomenon we term the question-first paradox. We trace the paradox to a conflict between two stages of...

    arxiv.org/abs/2607.15565 · PDF

  18. 18

    E3DGS: Unified Geometric-Photometric Equivariance for 3D Gaussian Splatting via Color-as-Geometry Embedding

    Chankyo Kim, Maani Ghaffari

    cs.CV · cs.LG

    3D Gaussian Splatting (3DGS) captures scenes by coupling explicit geometry (position, covariance) with view-dependent photometry (Spherical Harmonics). However, building $\mathrm{SE}(3)$-equivariant architectures on these primitives presents a fundamental representation bottleneck. Color has been treated as a signal rather than a geometric entity, making it nontrivial to unify symmetry across geometry and appearance as the camera frame...

    arxiv.org/abs/2607.15536 · PDF

  19. 19

    SLAPBench: Benchmarking Multimodal Large Language Models for Four-Finger SLAP Fingerprint Verification

    Bibesh Pyakurel, M. G. Sarwar Murshed

    cs.CV · cs.AI

    Four-finger SLAP fingerprints are flat live-scan impressions of the index, middle, ring, and little fingers of one hand, used for identity verification in border control and law enforcement. No benchmark has evaluated whether multimodal large language models (MLLMs) can verify identity from SLAP images. We introduce SLAPBench, the first benchmark for MLLM-based four-finger SLAP fingerprint verification, built from NIST SD302b with 7,832 pairs...

    arxiv.org/abs/2607.15517 · PDF

  20. 20

    Intentional Electromagnetic Interference Attacks on Facial Recognition

    Tyler Fitzsimmons, Adam Czajka

    cs.CV · cs.CR · cs.LG

    Attacks on general computer vision algorithms are often relegated to the digital domain, with the optimization performed purely in the digital world and then translated to physical mediums for implementation. In the field of biometrics, including facial recognition, physical presentation attacks targeting biometric sensors are dominant and present significant opportunity and risk. This paper highlights a critical vulnerability in the...

    arxiv.org/abs/2607.15512 · PDF

  21. 21

    LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o, and Claude Sonnet 4

    Mobina Kashaniyan, Amirhossein Ghassemi, Nasser Mozayani

    cs.CV · cs.AI · cs.LG

    We present a fully automated closed-loop AutoML framework that uses GPT-5, GPT-4o, and Claude Sonnet 4 as autonomous neural architecture designers for cross-lingual handwritten optical character recognition. Each large language model independently generates, trains, evaluates, and iteratively refines neural network architectures using performance feedback from previous trials. The framework is evaluated on Arabic, Persian, and English...

    arxiv.org/abs/2607.15509 · PDF

  22. 22

    Unsupervised Keypoints for Real-Time Fall Detection: Comparative Analysis Under Real-world Conditions with Predictive Bandwidth Reduction

    Tasmiah Haque, Jacob Kosinski, Sumit Mohan, Srinjoy Das, Mohammad Abdullah Al-Mamun

    cs.CV · cs.LG

    Falls among older adults are a major safety challenge, but continuous monitoring is difficult to sustain. Video captures fall-related posture and motion, yet deployment is limited by privacy, computation, and bandwidth. Supervised pose estimation is anatomically interpretable but vulnerable to occlusion and partial body visibility. We propose a privacy-preserving framework that replaces RGB transmission with compact motion representations...

    arxiv.org/abs/2607.15400 · PDF

  23. 23

    Partial Information Decomposition as a Multi-Contrast 3D MRI Selection Strategy for Resource-Constrained Deep Neural Network Training in Brain Tumor Segmentation

    Agamdeep Chopra, Mehmet Kurt

    cs.CV · cs.AI

    Multi-contrast 3D MRI segmentation can be computationally demanding when all available sequences are used. We evaluate a pre-training Partial Information Decomposition framework that ranks input pairs according to their redundant, unique, and synergistic information about regional tumor burden and selects the highest-ranked pair for downstream training. Applied to T1n, T1c, T2w, and T2-FLAIR MRI, the framework selected T1c+T2-FLAIR. We then...

    arxiv.org/abs/2607.15396 · PDF

  24. 24

    MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators

    Yushi Huang, Xiangxin Zhou, Jun Zhang, Liefeng Bo, Tianyu Pang

    cs.CV · cs.LG

    MeanFlow generators achieve fast few-step sampling by predicting average velocities over time intervals, making them attractive for efficient generation. Reinforcement learning (RL) has become a powerful way to align diffusion and flow models with human preferences and task-specific objectives. In particular, DiffusionNFT offers an efficient forward-process RL framework that does not require reverse-process trajectories or likelihood...

    arxiv.org/abs/2607.15273 · PDF

  25. 25

    Online Neural Space Time Memory for Dynamic Novel View Synthesis

    Baback Elmieh, Lynn Tsai, Zeman Li, Srinivas Kaza, Tiancheng Sun, Gabor Csapo, Ali Behrouz, Yuan Deng, Stephen...

    cs.CV · cs.GR · cs.LG

    Online novel view synthesis from multi-view streaming videos faces a fundamental trade-off: maintaining a persistent, long-horizon memory to reconstruct temporarily occluded regions while operating under strict real-time constraints. While Test-Time Training (TTT) offers a powerful memory mechanism, standard models mandate gradient-based memory updates at every frame to adapt to the changing motion in dynamic scenes. The computational cost of...

    arxiv.org/abs/2607.15271 · PDF

  26. 26

    SceneBind: Binding What and Where Across Vision, Audio and Language

    Mingfei Chen, Zijun Cui, Ruoke Zhang, Hyeonggon Ryu, Eli Shlizerman

    cs.CV · cs.AI · cs.MM · cs.SD

    We present SceneBind, an omni-modal representation of realistic scenes with joint semantic and 3D spatial understanding across vision, audio and language. Existing omni-modal encoders excel at instance-level semantics (i.e., what is present), but often lack explicit spatial structure (i.e., where it is). SceneBind addresses this gap by representing each scene as a semantic-spatial entity, combining a global semantic embedding with...

    arxiv.org/abs/2607.15265 · PDF

  27. 27

    Symbal: Detecting Systematic Misalignments in Model-Generated Captions

    Maya Varma, Jean-Benoit Delbrouck, Sophie Ostmeier, Akshay Chaudhari, Curtis Langlotz

    cs.CV · cs.AI

    Multimodal large language models (MLLMs) often introduce errors when generating image captions, resulting in misaligned image-text pairs. Our work focuses on a class of captioning errors that we refer to as systematic misalignments, where a recurring error in MLLM-generated captions is closely associated with the presence of a specific visual feature in the paired image. Given a vision-language dataset with MLLM-generated captions, our aim in...

    arxiv.org/abs/2607.15216 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.