cs.CV · 2026-08-06 · No. 76

Computer Vision and Pattern Recognition, 2026-08-06.

23 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

23 entries
  1. 01

    Predicting Brain Morphometry with MT-GNN: Mesh Evolution in Continuous Time with Graph-Based Metric Tensor Embeddings

    Hao Ding, Daniel Semchin, Paul M. Thompson, Boris Gutman

    cs.CV · cs.LG

    Predicting how a subcortical structure's shape will evolve from a few prior scans could support prognosis and clinical-trial enrichment. Existing longitudinal mesh predictors either extrapolate shape trajectories via high-dimensional embeddings or regress vertex deformations directly. We instead predict the surface's intrinsic geometry in continuous time: a single per-structure graph network predicts the future per-vertex first fundamental...

    arxiv.org/abs/2608.05132 · PDF

  2. 02

    OPD-V: Visual On-Policy Self-Distillation with Modality Balance

    Aniri, Jinhe Bi, Peng Liao, Zengjie Jin, Volker Tresp, Fei Shen, Yunpu Ma, Tat-Seng Chua

    cs.CV · cs.AI

    On-Policy Self-Distillation (OPSD) has become a standard post-training approach for improving visual reasoning in multimodal large language models (MLLMs). Existing methods draw privileged information from diverse input sources to guide self-distillation. Yet these designs overlook Modality Imbalance, a challenge inherent to MLLM reasoning. When textual information dominates generation, the model cannot fully integrate its multimodal input....

    arxiv.org/abs/2608.05131 · PDF

  3. 03

    Robust and Efficient Motion Reasoning for Privacy-Aware Classroom Incident Recognition

    Paritosh Parmar, Landy Lan, Hong Yang, Chen Yi, Chiat Pin Tay

    cs.CV · cs.AI · cs.ET · cs.HC · cs.LG

    Can computer vision help make classrooms safer? In this pilot study, we investigate privacy-aware and computationally efficient classroom incident recognition from CCTV-style observations. This setting remains underexplored, with limited benchmarks and few methods designed for the privacy, efficiency, and generalization demands of real-world deployment. We introduce a novel hybrid benchmark combining generative CCTV-style videos with...

    arxiv.org/abs/2608.05115 · PDF

  4. 04

    VQ-VAD: Vector-quantized Motion Representation Learning for Human-centric Video Anomaly Detection

    Narges Rashvand, Ghazal Alinezhad Noghre, Shanle Yao, Gabriel Maldonado, Hamed Tabkhi

    cs.CV · cs.AI

    Video Anomaly Detection (VAD) is inherently challenging due to the scarcity of anomalies and the large visual variability in surveillance footage, including changes in lighting, viewpoint, and human appearance. To mitigate visual noise and address privacy concerns, recent work has shifted to pose-based VAD, which focuses on motion dynamics rather than raw video data. However, existing pose-based approaches model human behavior in continuous...

    arxiv.org/abs/2608.05069 · PDF

  5. 05

    Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes

    Junlin Han, Shengbang Tong, David Fan, Minghao Chen, Philip Torr, Filippos Kokkinos, Mike Lewis

    cs.CV · cs.LG · cs.MM

    Vision offers a critical axis for advancing foundation models, driving a shift towards natively unified multimodal pretraining. Despite this momentum, the design space and the fundamental mechanisms of how modalities interact during unified training remain underexplored. We provide empirical clarity through a systematic exploration of multimodal pretraining. Our controlled experiments on both synthetic and large-scale real-world datasets...

    arxiv.org/abs/2608.05000 · PDF

  6. 06

    Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation

    Zehua Chen, Junyou Wang, Yuxuan Jiang, Zhenying Fang, Yusheng Dai, Jianfei Chen, Ziwei Liu, Jun Zhu

    cs.CV · cs.LG · cs.SD

    Video-to-audio (V2A) generation extends image-to-audio generation (I2A) by introducing consecutive frames that provide essential temporal cues for audio synthesis. However, existing conditional diffusion-based V2A methods typically enhance visual conditioning with additional audio-visual supervision, acoustic structure prediction, or reasoning from large multimodal models, requiring extra networks or strong inductive biases. Inspired by...

    arxiv.org/abs/2608.04902 · PDF

  7. 07

    Towards a satellite image manipulation and deepfake localization benchmark dataset

    Jacob Arndt, Debvrat Varshney, Philipe Dias, Nivedita Nukavarapu

    cs.CV · cs.AI

    Verifying the authenticity of satellite imagery has become increasingly critical given advances in generative artificial intelligence. Highly realistic synthetic imagery produced for malicious purposes (deepfakes) can have major consequences in the remote sensing domain, where this data is a fundamental source of information for science applications, planning, logistics, and monitoring. The remote sensing community lacks high-quality,...

    arxiv.org/abs/2608.04840 · PDF

  8. 08

    FUSEP: A Multi-Center Benchmark for Diverse Tasks in Early Pregnancy Fetal Ultrasound Screening

    Bin Pu, Jiewen Yang, Liwen Wang, Ying Tan, Guannan He, Xingbo Dong, Qika Lin, Jiarong Guo, Lixian Yang, Zuozhu Liu,...

    cs.CV · cs.AI

    A large number of infants with congenital anomalies are born each year globally, especially in areas with underdeveloped medical resources. Currently, fetal ultrasound screening is the most common modality for early pregnancy anatomy detection. This modality can detect anomalies earlier and provide opportune treatment advice. However, the lack of an ultrasound dataset on early fetal gestation has slowed down the development of automated...

    arxiv.org/abs/2608.04766 · PDF

  9. 09

    Teaching MLLMs to Say No: Generalized Referring Expression Comprehension via Refusal Calibrated GRPO

    Xuzheng Yang, Jun Ling, Tao Huang, Caiyan Qin, Peng Wang

    cs.CV · cs.AI

    We tackle the challenging yet underexplored task of Generalized Referring Expression Comprehension (GREC), which requires a model to localize the object described by a textual expression when it exists (positive sample) and to refuse output when it does not (negative sample). Although Multimodal Large Language Models (MLLMs) excel at localizing existing objects, they often fail to reject nonexistent ones due to the absence of negative samples...

    arxiv.org/abs/2608.04698 · PDF

  10. 10

    CSGen: A Multi-Domain Curvilinear Structure Generation Model via Hierarchical Multimodal Diffusion

    Zhe Shan, Ziming Yang, Lei Zhou, Wenwen Zhang, Cong Lin, Xia Xie

    cs.CV · cs.AI

    Curvilinear structure analysis is an important and fundamental task in multimedia. However, the controllable generation of images with precise curvilinear structure objects remains an open challenge. To address this, we propose CSGen, a hierarchical multimodal diffusion model that synthesizes high-fidelity images precisely aligned with multiple control conditions. The CSGen is built upon three key innovations: 1) We construct a multi-domain...

    arxiv.org/abs/2608.04655 · PDF

  11. 11

    DisMix: Order-Aware Mixup for Medical Imaging via Disentangling Ordinal and Non-Ordinal Features

    Dileepa Pitawela, Gustavo Carneiro, Hsiang-Ting Chen

    cs.CV · cs.AI

    Image mixup is a widely adopted data augmentation strategy, yet it is ill-suited for ordinal classification tasks such as medical disease grading, where labels encode a progression of severity. By indiscriminately blending disease-severity cues (ordinal) with appearance-level variation (non-ordinal), standard mixup produces samples that distort the very ordinal structure that underpins clinical severity grading. We introduce DisMix, an...

    arxiv.org/abs/2608.04652 · PDF

  12. 12

    The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering

    Yuqian Fu, Tianwen Qian, Yanjun Li, Yu Li, Kunyu Peng, Xu Zheng, Yongqin Xian, Alessio Tonioni, Yanwei Fu, Xiaoling...

    cs.CV · cs.AI

    EgoCross is a cross-domain egocentric video question answering benchmark designed to evaluate whether multimodal large language models can generalize beyond common daily-life scenarios. The first EgoCross Challenge was hosted at the Third EgoVis Workshop at CVPR 2026 and evaluated models on first-person videos from four target domains: surgery, industrial assembly, extreme sports, and animal perspectives. Each test example consists of an...

    arxiv.org/abs/2608.04589 · PDF

  13. 13

    PhysMind: From Video to Executable Worlds for Training-Free Physical Reasoning

    Chen Yang, Shenxiang Zeng, Haoyang Zhao, Zhouyuan Xu, Youquan He, Haoyu Li, Mingyi Deng, Jiansheng Fan, Chen Wang

    cs.CV · cs.AI

    Reliable physical reasoning from video requires understanding how objects move, interact, and respond to interventions. Existing vision-language models (VLMs) often struggle to interpret these dynamics and reason reliably about future and counterfactual outcomes. We introduce PhysMind, a training-free agentic framework that constructs one reusable, question-agnostic executable world per video. PhysMind recovers a temporally consistent dynamic...

    arxiv.org/abs/2608.04575 · PDF

  14. 14

    CARVE: Cross-Slice Anisotropic Reallocation of Visual Evidence for Efficient 3D Medical Volume Understanding

    Zhenyu Yi, Qiang Hu, Zhenhao Li, Jiaxuan Zhao, Yusong Sun, Lichi Zhang

    cs.CV · cs.AI · cs.CL

    Slice-based MLLMs leverage mature 2D encoders by representing 3D volumes as sequences of 2D slices. However, this slice-wise formulation produces thousands of visual tokens that burden the LLM backbone, many of which capture overlapping visual evidence across adjacent slices. To understand how effectively a growing visual token budget improves performance, we perform scaling analyses on two 3D medical VQA benchmarks and find diminishing...

    arxiv.org/abs/2608.04515 · PDF

  15. 15

    GeoReward: Mitigating Contextual Variable Overestimation in Vision-Language Models for Cross-Market Preference Prediction

    Shuo Liu, Huixiang Cai, Weiru Zhang, Xiaoyi Zeng

    cs.CV · cs.AI

    Vision-language models excel in many multimodal tasks but remain prone to a subtle yet impactful failure mode: they tend to overestimate dominant visual-textual cues while underestimating sparse but decision-critical contextual variables. This issue, which we term Contextual Variable Overestimation (CVE), becomes particularly evident in real-world applications such as predicting advertisement image preferences across diverse geographic...

    arxiv.org/abs/2608.04504 · PDF

  16. 16

    DIVE: Dynamic Iterative Visual Evidence Construction for Efficient Vision-Language Models

    Chen Zhong, Xiao An, Zijie Wang, Jiepan Li, Guangyi Yang, Wei He

    cs.CV · cs.LG

    Visual inputs in vision-language models (VLMs) are often encoded into substantially longer token sequences than text, making visual tokens a major bottleneck for efficient inference. Abundant recent methods address this bottleneck by scoring token importance and pruning low-scoring tokens in a single pass. However, one-shot scoring is insufficient because a token's prompt-relevant usefulness depends on the evidence already retained. Motivated...

    arxiv.org/abs/2608.04496 · PDF

  17. 17

    EndoVLM: An Endoscopy Vision-Language Pre-training Model via Anatomy-Guided Sparsity and Progressive Alignment

    Zhenyu Yi, Jianwei Xu, Yue Hu, Zhongwei Qiu, Sijing Li, Liang Huang, Bin Lv, Ling Zhang, Yingda Xia

    cs.CV · cs.AI · cs.CL

    The development of foundation models (FMs) is crucial for advancing endoscopic image analysis. However, existing endoscopy FMs mainly rely on self-supervised learning from uni-modal images or videos, overlooking the rich semantic knowledge contained in clinical reports. Furthermore, effectively leveraging these records is hindered by a fundamental modality gap: structured anatomical descriptions are not naturally mapped to specific frames...

    arxiv.org/abs/2608.04472 · PDF

  18. 18

    Beyond Global Routing Aggregation: Phase-Aware Expert Merging for MoE Vision-Language Models

    Hongyu Zhang, Cheng Yan, Xiang Xia, Wuyang Zhang

    cs.CV · cs.LG

    Mixture-of-experts vision-language models (MoE-VLMs) increase model capacity with sparse expert activation, yet deployment requires storing the full expert pool. Training-free expert merging reduces this burden, and many routing-based methods aggregate routing statistics across all tokens to determine merge compatibility. However, MoE-VLM inference is phase-structured: image-context tokens carry visual content, question tokens specify the...

    arxiv.org/abs/2608.04454 · PDF

  19. 19

    TwinIR: Coordinated Invisible Dual-Point Attacks on Online HD Map Construction

    Haibo Hu, Jianghuai Deng, Chen Tang, Yang Lou, Qian Xu, Jianping Wang

    cs.CV · cs.AI

    Online HD map construction is critical to prediction and planning in autonomous driving. We find that existing physical attacks against online map construction are limited by a cross-boundary compensation effect: after the target boundary is perturbed, another visible boundary may retain sufficient geometric cues for the model to recover the original road geometry. Based on this observation, we propose TwinIR, a new mechanism-guided physical...

    arxiv.org/abs/2608.04453 · PDF

  20. 20

    Q-CueGraph: Query-Conditioned Visual Evidence Graphs for Multimodal Reasoning

    Pengcheng Pan, Xinfang Zhang

    cs.CV · cs.AI · cs.CL

    High-resolution pixels and crop or zoom tools give multimodal large language models the ability to inspect an image, but they do not provide a reliable task-conditioned policy for deciding where to inspect. Q-CueGraph makes this decision explicit. It maps a question and an image representation to budgeted, coordinate-level observations for a frozen reader. Text-rich images use a reusable OCR/layout graph; natural-image search instantiates...

    arxiv.org/abs/2608.04452 · PDF

  21. 21

    When does training on downscaled images yield the same gradients?

    Seunghyun Ji

    cs.CV · cs.AI

    Diffusion transformers deliver strong image generation, but their training cost grows superlinearly with resolution. Recent work justifies training or sampling at reduced resolution on a spectral premise: at high noise, a downscaled latent preserves almost the full surviving signal. Whether a downscaled step also preserves the native training gradient signal, however, has remained unresolved. We reduce how that signal changes under...

    arxiv.org/abs/2608.04448 · PDF

  22. 22

    Image Classification Using CNN-QNN Hybrid Model with Optimized Correlated Features

    Minseo Seong, Youngwook Kim

    cs.CV · cs.AI

    We propose a method to optimize the correlation among convolutional neural network (CNN) features that are used as inputs to quantum neural network (QNN) to enhance image classification accuracy. Unlike prior approaches that employ orthogonal decomposition as preprocessing, we intentionally introduce correlated features that are more physically compatible with QNN. This design leverages the QNN's inherent ability to exploit quantum...

    arxiv.org/abs/2608.04379 · PDF

  23. 23

    iStructTab: Structured Feature Sequencing for Multimodal Learning of Image and Tabular Data

    Al Zadid Sultan Bin Habib, Md Younus Ahamed, Prashnna Gyawali, Gianfranco Doretto, Donald A. Adjeroh

    cs.CV · cs.AI · cs.LG · stat.ML

    Multimodal learning of images and tabular data is often impaired by ineffective representations, resulting in redundancy, dispersion, and generalization problems. To tackle this challenge, we introduce Graph-Enhanced Descriptor Sequencing (GEDS), a structured feature sequencing algorithm grounded in principles from the Column Permutation Problem (CPP). GEDS refines statistical descriptors of the features through similarity graph-based...

    arxiv.org/abs/2608.04348 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.