cs.CV · 2026-07-29 · No. 68

Computer Vision and Pattern Recognition, 2026-07-29.

24 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

24 entries
  1. 01

    VetClaw: An Edge-Cloud Multimodal Agentic System for Veterinary Disease Screening

    Syed Mhamudul Hasan, Anas AlSobeh, Hussein Zangoti, Abdur R. Shahid

    cs.CV · cs.LG

    We present VetClaw, an edge-cloud multimodal agentic system for early veterinary disease screening. VetClaw uses a camera module as an edge sensing device and sends captured images, together with optional symptom descriptions, to a server-hosted vision-language model for zero-shot disease classification. The system separates agent interaction from workflow orchestration: OpenClaw provides scheduling, tool access, user interaction, and...

    arxiv.org/abs/2607.26042 · PDF

  2. 02

    Pictura: Perspective-View Self-Play at Scale for Driving

    Yuan Yin, Elias Ramzi, Marc Lafon, Valentin Charraut, Victor Bares, Yihong Xu, Éloi Zablocki, Alexandre Boulch,...

    cs.CV · cs.AI · cs.RO

    Self-play in simulation produces robust driving policies at scale. Demonstrations of such behavior have been made using privileged vectorized observations such as exact poses and velocities, even for occluded agents. This assumes that perception is solved and introduces a representation gap with the partial observation of a deployed agent driving from the perspective view of egocentric cameras. A common fix, distilling the privileged policy...

    arxiv.org/abs/2607.26005 · PDF

  3. 03

    Parallel Decoding Distillation for Fast Image and Video Generation

    Neta Shaul, Chao Liu, Arash Vahdat, Julius Berner

    cs.CV · cs.LG

    Generation in video diffusion or flow models is computationally expensive due to the slow and iterative sampling process. Current state-of-the-art (SOTA) acceleration methods heavily rely on variational score distillation (VSD) and adversarial losses to distill diffusion models into few-step generators. Albeit achieving high-quality video generation, these training losses are notoriously hard to optimize and suffer from mode collapse, leading...

    arxiv.org/abs/2607.26004 · PDF

  4. 04

    Schrödinger's Cat: Probabilistic Representation and Prediction of Potential Scene Kinematics

    Timy Phan, Jannik Wiese, Björn Ommer

    cs.CV · cs.LG

    Predicting how a scene may evolve from partial observations requires reasoning about multiple possible futures rather than committing to a single trajectory. Existing approaches either generate appearance-dominated video predictions or sample a small number of trajectories without explicitly modeling the distribution of possible motion. We introduce Goal-Aware Representations of Future kInEmatic Latent Distributions (GARFIELD), a...

    arxiv.org/abs/2607.25984 · PDF

  5. 05

    Quasi-SVD: Learning a Lie-constrained matrix factorisation for real-time imaging

    Christopher Hahne

    cs.CV · cs.LG · math.NA

    Singular Value Decomposition (SVD) underlies matrix factorisation tasks across computational imaging, with medical applications increasingly demanding real-time processing. Yet SVD algorithms are inherently sequential, constraining real-time GPU throughput and limit online deployment in clinical pipelines. This study introduces Quasi-SVD, a differentiable, fully parallelized matrix factorization framework for GPUs. Rather than enforcing...

    arxiv.org/abs/2607.25967 · PDF

  6. 06

    Knowledge-Guided Multimodal Reasoning over Interacting Streams for Video-Level Ambivalence and Hesitancy Recognition

    Podakanti Satyajith Chary, Barath Parthiban, Pranesh Velmurugan, Adeeba Khan, Nagarajan Ganapathy

    cs.CV · cs.AI

    Ambivalence and hesitancy (A/H) are conflicting affective states that precede the delay or abandonment of health behaviour change. Recognition of A/H at the video level is difficult, since the signal arises from disagreement across and within facial, vocal, linguistic, and bodily modalities, and manifests differently across individuals. The proposed PRISM-AH (Predictive Reasoning over Interacting Streams for Multimodal Ambivalence/Hesitancy...

    arxiv.org/abs/2607.25961 · PDF

  7. 07

    MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities

    Mingqiao Ye, Zhaochong An, Zhitong Gao, Xian Liu, François Fleuret, Chuan Li, Amir Zadeh, Serge Belongie, Afshin...

    cs.CV · cs.AI · cs.LG

    Any-to-any models predict any modality from any combination of others within a single network, a formulation used in multimodal vision and vision-language models, and increasingly in scientific domains such as ecology and astronomy. Existing any-to-any models are typically trained from scratch using encoder-decoder or diffusion architectures, impacting their performance and preventing them from using strong pre-trained decoder-only models as...

    arxiv.org/abs/2607.25948 · PDF

  8. 08

    Face De-Identification: A Domain-Centric Survey from Capture to Processing

    Hui Wei, Hao Yu, Guoying Zhao

    cs.CV · cs.AI

    Face de-identification (De-ID) aims to remove or conceal personally identifiable facial features in images or videos to prevent identity recognition while preserving utility for downstream tasks. With the rising emphasis on data privacy and responsible AI, face De-ID has emerged as an active research area spanning computer vision and privacy-preserving communities. Early approaches, and many contemporary ones, operate in the digital domain by...

    arxiv.org/abs/2607.25926 · PDF

  9. 09

    Evaluating VLMs for Autonomous Agent-Driven Geometry Clipping Detection in Video Game QA

    Carlos Celemin, Benedict Wilkins, Adrián Barahona-Ríos, Saman Zadtootaghaj, Nabajeet Barman

    cs.CV · cs.AI

    In this work, we study the use of Vision-Language Models (VLMs) for anomaly detection in an agent-driven game Quality Assurance (QA) pipeline focusing on geometry clipping. In this evaluation, a custom exploration agent navigates a game level to collect visual observations, while the automatic annotation pipeline provides frame-level clipping labels. This setup allows us to evaluate recent VLMs on a controlled anomaly detection task without...

    arxiv.org/abs/2607.25921 · PDF

  10. 10

    Image Quality Dependent Degradation for AI Systems

    Yannick Kees, Elena Hoemann, Frank Köster, Sven Hallerbach

    cs.CV · cs.AI

    Perception is one of the primary applications where neural networks outperform conventional algorithms. One example is AI systems for automated driving, which can detect pedestrians based on image data and avoid them accordingly. A substantial challenge with these AI systems is that their output depends heavily on the quality of the input images. For example, if an image is of inferior quality due to heavy contamination, such as noise or...

    arxiv.org/abs/2607.25736 · PDF

  11. 11

    OmniPhys: Knowledge-Graph-Driven Benchmarking and Collective Optimization for Physical Commonsense in Text-to-Image Generation

    Yajing Xu, Yarong Lan, Jiaoyan Chen, Yichi Zhang, Jeff Z. Pan, Mingchen Tu, Zhizhen Liu, Wen Zhang, Huajun Chen

    cs.CV · cs.AI

    While text-to-image models exhibit remarkable visual fidelity, they frequently violate fundamental physical commonsense. Existing benchmarks often rely on coarse-grained descriptions, failing to diagnose the mastery of specific physical principles. Moreover, the high stochasticity of generative processes causes current prompt optimization methods to suffer from gradient hallucinations, where optimizers are misled by transient visual artifacts...

    arxiv.org/abs/2607.25641 · PDF

  12. 12

    The LAIA Dataset: Labelled Attention for Intelligent Automobiles

    A. Contreras, D. Porres, R. Abad, P. Cano, G. Villalonga, A. M. López, A. Hernández-Sabaté

    cs.CV · cs.AI · cs.SE

    The development of autonomous vehicles (AVs) usually relies heavily on data-driven artificial intelligence (AI) models that require large volumes of sensor data with ground-truth annotations. While modular architectures are widely used, end-to-end driving paradigms offer a promising alternative by directly mapping sensor inputs to control actions. However, their adoption is limited by challenges in interpretability and explainability. To...

    arxiv.org/abs/2607.25570 · PDF

  13. 13

    Less is More: Modality-Decoupling for General AIGC Audio-Video Detection

    Jielun Peng, Yabin Wang, Yaqi Li, Jincheng Liu, Xiaopeng Hong, Athanasios V. Vasilakos

    cs.CV · cs.AI

    Generative AI has rapidly expanded audio-visual forgery beyond human-centric deepfakes into general scenes. Existing AIGC detection methods assume audio-visual content correspondence, identifying forgeries by spotting cross-modal inconsistencies. However, we empirically find that this assumption does not consistently hold in general scenarios. We argue that, for general audio-visual AIGC detection, decision-level fusion is a more robust...

    arxiv.org/abs/2607.25543 · PDF

  14. 14

    Visual prompt engineering for video models

    Robert Geirhos, Yuxuan Li, Thaddäus Wiedemer, Neha Kalibhat, Zi Wang, Mani Malek, Oyvind Tafjord, Kevin Swersky,...

    cs.CV · cs.AI

    In the age of foundation models, a model is only as good as its prompt. For this reason, prompt engineering has become an essential technique for improving language model performance. Since video models are currently becoming foundation models for visual tasks (e.g., visual reasoning), we here ask whether they similarly benefit from visual prompt engineering: automatically modifying the task image to improve model performance. For example,...

    arxiv.org/abs/2607.25537 · PDF

  15. 15

    Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation

    Weiming Zhuang, Jiabo Huang, Jingtao Li, Zhizhong Li, Chen Chen, Sina Sajadmanesh, Lingjuan Lyu

    cs.CV · cs.AI

    Unifying visual understanding and generation in one model holds immense promise, but remains challenging and expensive due to heavy compute and data demands and conflicts between the visual features needed for these two capabilities. To address these challenges, we present Argus-Unified, a compact, effective and unified multimodal model built with low demand on computation and data. Instead of aligning modalities from scratch, Argus-Unified...

    arxiv.org/abs/2607.25527 · PDF

  16. 16

    ReLATE: Reliability-Guided Evidence Fusion for Robust UAV--Satellite cross-view Geo-Localization

    Haochen Jiang, Jialei Pan, Yuzhe Sun, Zhe Dong, Lecheng Ren, Yanfeng Gu, Tianzhu Liu

    cs.CV · cs.AI

    Unmanned aerial vehicle (UAV)-satellite cross-view geo-localization matches UAV images against satellite imagery and has achieved impressive accuracy on clean (non-degraded) image benchmarks. In real-world flights, however, UAV observations are frequently affected by adverse weather, illumination changes, platform motion, sensor noise, and compression, while the robustness of existing methods under such degradations remains largely...

    arxiv.org/abs/2607.25524 · PDF

  17. 17

    I2VShield: An Efficient Proactive Defense Framework against DiT-based Image-to-Video Models

    Yimao Guo, Zuomin Qu, Wei Lu

    cs.CV · cs.AI

    The rapid advancement of video generation models has led to the increasing misuse of image-to-video (I2V) models. Although substantial progress has been made in detecting AI-generated videos, proactive defenses against I2V models remain underexplored. In particular, current proactive defenses against I2V models predominantly rely on gradient-based adversarial attacks, which require defenders to possess GPUs with substantial memory resources...

    arxiv.org/abs/2607.25522 · PDF

  18. 18

    Beyond Counts: A Distributional Robustness Margin For Pathology Foundation Models

    Clément Grisi, Jeroen van der Laak, Geert Litjens

    cs.CV · cs.AI

    Pathology foundation models are approaching clinical deployment, yet remain vulnerable to systematic non-biological variation across centres. Differences in tissue preparation, staining and scanning are strongly encoded in their representations, enabling shortcut learning and weakening generalisation across cohorts and institutions. The Robustness Index (RI) quantifies whether local representation geometry is dominated by biology or by...

    arxiv.org/abs/2607.25497 · PDF

  19. 19

    Seen, Said, or Forgotten? A Causal Audit of Visual KV Memory Across Dialog Turns

    Hong Chen, Kang Chen, Yuxuan Fan, Bo Wang, Yubo Gao, Yuanlin Chu, Xuming Hu

    cs.CV · cs.AI

    Stateful multimodal assistants encode an image once but may answer questions about it many turns later. Attention-guided visual-KV eviction assumes that evidence irrelevant now will remain dispensable, although future questions are unknown. We ask when a visual fact is actually safe to forget and introduce the Causal Visual Memory Audit (CVMA), a paired single-prefill framework that tests what later answers lose when a visual region, the...

    arxiv.org/abs/2607.25467 · PDF

  20. 20

    Balanced Soft mixture-of-expert model for Glaucoma Detection

    Sai Venkatesh Chilukoti, Krishna Rauniyar, Min Shi, Xiali Hei

    cs.CV · cs.AI

    Glaucoma is a group of eye diseases that damage the optic nerve, often caused by elevated intraocular pressure. It is a leading cause of irreversible vision loss and is typically developed slowly and painlessly, making it difficult to notice until significant damage has occurred. Therefore, early detection is crucial to prevent or slow the progression of vision loss. In recent years, deep learning based uni-modal models have improved the...

    arxiv.org/abs/2607.25324 · PDF

  21. 21

    CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition

    Lai Wei, Chengqi Li, Jiapeng Li, Ruina Hu, Yue Wang, Weiran Huang

    cs.CV · cs.AI · cs.CL · cs.LG

    Real-world tasks often require models to learn from task-specific context rather than relying only on pre-trained knowledge. While recent work has highlighted this capability as context learning, existing evaluations mainly focus on textual contexts. In many practical settings, however, the context to be learned from is multimodal: scientific findings are conveyed through figures and tables, financial indicators are scattered across converted...

    arxiv.org/abs/2607.25294 · PDF

  22. 22

    FunnelAL: Retrieve-then-Rank Active Learning for Single-Class Discovery

    Reihaneh Rostami, Brian Goodwin

    cs.CV · cs.IR · cs.LG

    We present FunnelAL, a retrieve-then-rank active learning system for single-class discovery, which adapts the multi-stage funnel architecture of industrial recommender systems to data annotation. Large-scale supervised learning faces two challenges: efficiently finding relevant samples in a massive corpus, and distinguishing true positives from visually confusable negatives when embeddings do not cleanly separate classes. Conventional active...

    arxiv.org/abs/2607.25276 · PDF

  23. 23

    ScaleResfusion: Residual Rectified Flow based on Residual Vector Field

    Zhenning Shi, Chen Xu, Junhao Zhang, Kefei Zhang, Linjie Liu, Zhedong Zheng, Tao Li

    cs.CV · cs.AI

    Real-world Image Restoration (Real-IR) aims to recover high-quality (HQ) images from complex and unknown degradations. Although recent diffusion-based methods have substantially improved perceptual quality, their current designs leave two key challenges unresolved. Methods that start from Gaussian noise are slow and often less faithful to the degraded input. Residual-based methods usually train from scratch, which makes it hard to exploit...

    arxiv.org/abs/2607.25275 · PDF

  24. 24

    FORGE: Frame Orthogonality in Relevance Geometry for Long-Form Video Understanding

    Ghazal Kaviani, Ghassan AlRegib

    cs.CV · cs.LG

    Multimodal large language models (MLLMs) have enabled long-form video understanding at a scale that was not previously possible. However, the density of relevant content decreases sharply as video sequence length increases, and exposing the model to more irrelevant content measurably reduces its accuracy. In this paper, we address the problem of maximizing query-relevant information in a frame subset selected at inference time, without...

    arxiv.org/abs/2607.25266 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.