cs.CV · 2026-07-02 · No. 41

Computer Vision and Pattern Recognition, 2026-07-02.

35 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

35 entries
  1. 01

    World from Motion: Generative Dynamic Gaussian Reconstruction from Monocular Video

    Liyuan Zhu, Shengyu Huang, Amrita Mazumdar, Tianye Li, Zan Gojcic, Gordon Wetzstein, Iro Armeni, Shalini De Mello,...

    cs.CV · cs.AI · cs.GR

    We present World from Motion, a method for generating freely renderable dynamic 3D Gaussian representations from monocular videos. Our approach conditions a video model on dense, pixel-aligned renderings that encode appearance, geometry, and 3D scene motion along both input and target camera trajectories to correct rendering artifacts and fill in missing regions from an initial reconstruction. To train this model, we construct a dataset of...

    arxiv.org/abs/2607.01202 · PDF

  2. 02

    Autonomous Scientific Discovery via Iterative Meta-Reflection

    Bingchen Zhao, Sara Beery, Oisin Mac Aodha

    cs.CV · cs.AI

    Autonomous scientific discovery systems offer the potential to accelerate research by automating the process of hypothesis generation and validation. However, current systems operate within constrained search spaces or require predefined research questions, limiting their capacity for true open-ended inquiry. Furthermore, while they generate hypotheses iteratively, they largely lack the ability to explicitly synthesize their own accumulated...

    arxiv.org/abs/2607.01131 · PDF

  3. 03

    LongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language Models

    Arpita Nema, Hanwei Zhu, Xi Zhang, Weisi Lin

    cs.CV · cs.AI

    The evaluation of long-term video quality understanding remains an open challenge for large vision-language models (LVLMs). Existing video quality benchmarks predominantly focus on short clips and isolated distortions, overlooking the temporal continuity, cumulative degradation, and reasoning complexity inherent in long-duration content. To address these limitations, we present LongVQUBench, a comprehensive benchmark for long-term video...

    arxiv.org/abs/2607.01086 · PDF

  4. 04

    EchoRisk: A Multicentre Echocardiography Dataset and Benchmark for Cardio-Oncology

    Grigorios Kalliatakis, Georgia Karanasiou, Georgios Manikis, Manolis Tsiknakis, Dimitrios Fotiadis, Dorothea...

    cs.CV · cs.AI

    Therapy-induced cardiotoxicity is the leading non-oncological cause of treatment interruption in breast cancer patients, yet early, automated risk stratification from routine cardiac imaging remains an unsolved problem. We present EchoRisk, the first curated, multicentre, longitudinal echocardiography dataset with explicit cardiotoxicity labels, released as the primary technical reference for the EchoRisk-MICCAI 2026 challenge. The dataset...

    arxiv.org/abs/2607.01039 · PDF

  5. 05

    Foundation Models vs. Radiomics for Lung Computed Tomography: A Benchmark of Feature Extractors, Classification Heads, and Segmentation Choices

    Nils Neukirch, Martin Maurer, Nils Strodthoff

    cs.CV · cs.LG

    Radiomics is the established approach for CT-based lung cancer phenotyping, yet comparisons with foundation models rarely isolate contributions of feature extractor, classification head, and segmentation choice, or test cross-cohort robustness. We benchmark five feature extractors (Curia, Curia-2, DINOv3, Radiomics2D, Radiomics3D), seven classification heads (TabPFN, TabICL, XGBoost, CatBoost, Random Forest, logistic regression, Ridge), and...

    arxiv.org/abs/2607.01001 · PDF

  6. 06

    TRCGL-Net: A Long-Tailed Multi-Label Chest X-Ray Classification Framework with Generative Data Augmentation and Label Co-Occurrence Modeling

    Tong Shao, Hongshun Ling, Li Zhang, Jinjing Wu, Junke Wang, Yuan Gao, Fang Wang

    cs.CV · cs.AI

    Chest X-ray multi-label classification is a core task in intelligent medical imaging diagnosis. However, real clinical data often exhibit extreme long-tailed distributions, leading to degraded performance on rare diseases in tail classes. This issue is not only driven by data scarcity but also by two intrinsic factors:1) attenuation of tail-class lesion representations under complex anatomical backgrounds, and 2) dominance of head classes in...

    arxiv.org/abs/2607.00975 · PDF

  7. 07

    Learning Cardiac Motion Priors for Implicit Neural Representations

    Andrew Bell, George Webber, Andrew P King, Steffen E Petersen, Muhummad Sohaib Nazir, Alistair Young

    cs.CV · cs.AI

    Implicit neural representations (INRs) are well suited to cardiac motion estimation, providing continuous, compact representations of motion fields. However, fitting an INR to each image sequence is time-consuming and sensitive to the optimisation trajectory. Learned priors can help guide optimisation towards plausible motion fields and enable faster adaptation, but learning priors for cardiac motion INRs remains under-explored. In this work,...

    arxiv.org/abs/2607.00955 · PDF

  8. 08

    Post-Training Pruning for Diffusion Transformers

    Chengzhi Hu, Xuewen Liu, Jing Zhang, Mengjuan Chen, Zhikai Li, Qingyi Gu

    cs.CV · cs.AI

    Diffusion Transformers (DiTs) have demonstrated impressive performance in image generation but suffer from substantial computational overhead and resource consumption. Post-training pruning offers a promising solution; however, due to DiTs' unique architectural design and parameter distribution, traditional pruning methods are inapplicable, leading to significant performance degradation. Specifically, prior methods developed for LLMs, which...

    arxiv.org/abs/2607.00927 · PDF

  9. 09

    DeWorldSG: Depth-Aware 3D Semantic Scene Graph Generation via World-Model Priors

    Seok-Young Kim, Abdelrahman Elskhawy, Taewook Ha, Dooyoung Kim, Eunjae Shin, Benjamin Busam, Woontack Woo

    cs.CV · cs.AI

    We present DeWorldSG, a novel framework that generates spatio-temporally robust 3D Semantic Scene Graphs from RGB-D sequences. Existing methods often struggle to construct reliable 3D scene graphs due to unstable 3D object representations and missing relations caused by frame-wise inference. DeWorldSG addresses these issues by estimating instance-level geometric 3D Gaussian distributions through depth-guided filtering and representing each...

    arxiv.org/abs/2607.00889 · PDF

  10. 10

    Improving Sparse-View 3DGS Generalization via Flat Minima Optimization

    Kangmin Seo, Sangeek Hyun, MinKyu Lee, Jae-Pil Heo

    cs.CV · cs.AI

    Recent advances in neural rendering have established 3D Gaussian Splatting (3DGS) as a highly efficient representation for novel view synthesis, enabling fast training and real-time rendering with strong fidelity. However, when supervision is limited to sparse input views, 3DGS tends to overfit to the observed images and generalize poorly to unseen viewpoints. We address this challenge from the perspective of flat minima (FM) optimization,...

    arxiv.org/abs/2607.00885 · PDF

  11. 11

    MoVA: Learning Asymmetric Dual Projections for Modular Long Video-Text Alignment

    Peiyuan Zhu, Shaoan Xie, Zijian Li, Yifan Shen, Namrata Deka, Harsh Shrivastava, Guangyi Chen, Kun Zhang

    cs.CV · cs.LG

    Contrastive pre-training has propelled video-text alignment, yet models often inherit the critical limitations of their image-text predecessors like CLIP, resulting in entangled representations. These challenges are severely exacerbated by two fundamental properties in the video domain: Temporal Misalignment, where textual descriptions often correlate only to specific, constrained temporal windows, leaving other frames text-irrelevant; and...

    arxiv.org/abs/2607.00858 · PDF

  12. 12

    Mirror-Fusion Attention for Reflection-Aware Self-Supervised Representation Learning

    Ruixin Li, Jin Liu, Yuling Shi, Stefano Lodi

    cs.CV · cs.LG

    Most self-supervised learning (SSL) methods encourage invariance across augmentations, but strict flip invariance can suppress informative left--right correspondences in approximately bilateral data such as medical images and human faces. We propose Mirror-Fusion-Augmented Self-Supervised Learning (MFASSL), a Vision Transformer framework that injects a soft reflection prior into standard SSL without redesigning the backbone. MFASSL constructs...

    arxiv.org/abs/2607.00850 · PDF

  13. 13

    Pano2World: End-to-End 3D Generation via Unified Multi-View Sequences

    Zhenjia Li, Jinrang Jia, Yifeng Shi

    cs.CV · cs.AI

    A single panorama captures the full visual sphere from one camera center, yet confines users to looking around in place without enabling true scene exploration. Converting a single panorama into a persistent, renderable 3D representation for free-viewpoint navigation has attracted growing interest; existing methods either adopt iterative per-view completion that propagates inpainting results to update the underlying geometry, leading to...

    arxiv.org/abs/2607.00832 · PDF

  14. 14

    LeVLJEPA: End-to-End Vision-Language Pretraining Without Negatives

    Lukas Kuhn, Giuseppe Serra, Randall Balestriero, Florian Buettner

    cs.CV · cs.AI

    Vision-language pretraining remains dominated by contrastive objectives, whereas vision-only self-supervised learning has largely adopted non-contrastive methods. At the same time, the role of vision-language encoders has shifted: they are increasingly deployed not as zero-shot classifiers but as the frozen visual backbone of vision-language models and dense prediction systems, which consume the full grid of patch tokens rather than a single...

    arxiv.org/abs/2607.00784 · PDF

  15. 15

    Soft Mixture-of-Recursions: Going Deeper with Recursive Vision Transformers

    Sang In Lee, Jihun Park

    cs.CV · cs.LG

    Recent recursive Transformer studies have primarily reused shared parameters across computation steps to construct compact, parameter-efficient models. In this work, we leverage recursion to build effectively deeper Transformers with stronger representational capacity. However, in Vision Transformers, simply increasing recursion depth does not reliably improve performance, as existing recursive approaches do not fully utilize the intermediate...

    arxiv.org/abs/2607.00774 · PDF

  16. 16

    Active Learning for Cascaded Object Detection: Balancing Coverage and Uncertainty in Table Extraction Pipelines

    Eliott Thomas, Mickael Coustaty, Aurelie Joseph, Gaspar Deloin, Vincent Poulain d'Andecy, Jean-Marc Ogier

    cs.CV · cs.AI

    Table extraction from business documents relies on a cascaded pipeline where Table Detection (TD) first localizes tables and Table Structure Recognition (TSR) then recovers their internal layout. Building task-specific training sets for this pipeline is costly, particularly for TSR which requires fine-grained structural annotations. Active learning (AL) can reduce this annotation burden, yet most AL strategies are designed for single-model...

    arxiv.org/abs/2607.00747 · PDF

  17. 17

    GaussianFusion: Unified 3D Gaussian Representation for Multi-Modal Fusion Perception

    Xiao Zhao, Chang Liu, Mingxu Zhu, Zheyuan Zhang, Linna Song, Qingliang Luo, Chufan Guo, Kuifeng Su

    cs.CV · cs.AI

    The bird's-eye view (BEV) representation enables multi-sensor features to be fused within a unified space, serving as the primary approach for achieving comprehensive 3D perception. However, the discrete grid representation of BEV leads to significant detail loss and limits feature alignment and cross-modal information interaction in multimodal fusion perception. In this work, we break from the conventional BEV paradigm and propose a new...

    arxiv.org/abs/2607.00746 · PDF

  18. 18

    Prototype Memory-Guided Training-Free Anomaly Classification and Localization in Prenatal Ultrasound

    Huanwen Liang, Yuhao Huang, Xiliang Zhu, Yuanji Zhang, Xuedong Deng, Xinru Gao, Guowei Tao, Yuhan Zhang, Dong Ni

    cs.CV · cs.AI

    Prenatal anomaly classification and localization is of critical importance for fetal health and pregnancy management. Although ultrasound (US) is the primary modality for prenatal screening, accurate diagnosis remains challenging due to the low prevalence and high heterogeneity of anomalies. Existing deep learning methods for prenatal tasks rely on large-scale annotated datasets, which are difficult to obtain in practice. Although few-shot...

    arxiv.org/abs/2607.00744 · PDF

  19. 19

    ConRTF: Edge-Constrained Boundary Distribution Refinement for Realtime TransFormer Table Structure Recognition

    Eliott Thomas, Tri-Cong Pham, Mickael Coustaty, Aurelie Joseph, Gaspar Deloin, Vincent Poulain d'Andecy, Jean-Marc...

    cs.CV · cs.AI

    Table Structure Recognition (TSR) aims to recover the row and column layout of tables from document images, a key step in document understanding pipelines. Accurate TSR depends on precise boundary localization: small errors in row or column boundaries can propagate into incorrect cell assignments and structural inconsistencies. Yet detection-based approaches treat table elements as generic objects, ignoring a fundamental property of table...

    arxiv.org/abs/2607.00734 · PDF

  20. 20

    Partial Skeleton Visibility for Action Recognition: A Constrained Field-of-View Approach

    Yingjie Dai, Tianyang Xu, Yanglin Deng, Xiao-Jun Wu, Josef Kittler

    cs.CV · cs.AI

    Skeleton-based action recognition has achieved remarkable success by exploiting joint coordinates and their topological connections, yet prevailing methods overwhelmingly assume complete and clean skeleton inputs. In real-world deployments, such as egocentric vision, crowded surveillance, wearable devices, or edge robotics, limited field-of-view (FoV) frequently causes substantial joint visibility dropout, leading to severe performance...

    arxiv.org/abs/2607.00716 · PDF

  21. 21

    Creating Impactful Autonomous Driving Datasets: A Strategic Guide from Research Gap to Benchmark

    Richard Schwarzkopf, Jonas Merkert, Frank Bieder, Annika Bätz, Alexander Blumberg, Carlos Fernandez, Felix Hauser,...

    cs.CV · cs.AI · cs.RO

    Well-designed autonomous driving datasets have fundamentally shaped research progress, yet existing literature primarily describes what datasets contain rather than how to strategically design impactful ones. This is especially limiting for small and medium-sized labs and startups that cannot afford to misallocate scarce resources. We argue that impactful dataset creation begins with a diagnosis: whether a research question is blocked by a...

    arxiv.org/abs/2607.00710 · PDF

  22. 22

    LUMA: Benchmarking Segmentation via a Lightweight Universal Mask Adapter

    Tobias Christian Nauen, Anosh Billimoria, Federico Raue, Stanislav Frolov, Brian B. Moser, Andreas Dengel

    cs.CV · cs.AI · cs.LG · cs.PF

    Comparing transformer backbones for image segmentation is confounded: each is paired with a different decoder, recipe, and pretraining, so reported differences rarely reflect the backbone itself. We introduce the Lightweight Universal Mask Adapter (LUMA), a lightweight, backbone-agnostic mask-transformer head that treats any backbone as a black-box feature extractor, letting a set of queries read from its features through cheap...

    arxiv.org/abs/2607.00687 · PDF

  23. 23

    Identifying Latent Concepts and Structures for Generalized Category Discovery

    Boyang Dai, Chaoqi Chen, Yizhou Yu

    cs.CV · cs.AI

    Generalized Category Discovery (GCD) aims to recognize known classes while autonomously discovering novel ones in open-world settings. However, current approaches primarily focus on designing clustering objectives, often overlooking a critical bottleneck: standard vision backbones yield high-rank, entangled token representations that are ill-suited for unsupervised discovery of latent concepts and structures. In this paper, we propose...

    arxiv.org/abs/2607.00620 · PDF

  24. 24

    EgoGapBench: Benchmarking Egocentric Action Selection in Multi-Agent Scenes

    Jihyeok Jung, Jeewu Lee, Sanghyeop Kim, Chanhee Han, Seong Joon Oh

    cs.CV · cs.AI

    Existing egocentric benchmarks have primarily constructed the egocentric setting from first-person-view data, which makes it difficult to evaluate egocentric perspective itself in isolation. However, understanding first-person-view input and taking an egocentric perspective are separable abilities, especially when first-person body cues are absent or when other agents are present. To isolate egocentric perspective understanding, we introduce...

    arxiv.org/abs/2607.00547 · PDF

  25. 25

    Cross4D-JEPA: Dense Cross-modal Correspondence Distillation for 4D Point Cloud Representation Learning

    Trung Thanh Nguyen, Hai Nguyen-Truong, Tu Vo, Hoang M. Truong, Tuan-Anh Vu

    cs.CV · cs.AI

    Automatic understanding of dynamic 4D point clouds, the 3D-point sequences captured over time by depth sensors and LiDAR, is central to robotics and embodied perception. Yet annotating them densely is expensive, making self-supervised pretraining the natural route to transferable representations. Existing pretext tasks, however, are almost entirely intra-modal, and the few methods that transfer knowledge from 2D foundation models rely on a...

    arxiv.org/abs/2607.00514 · PDF

  26. 26

    MindEdit-Bench: Benchmarking Object-Level Counterfactual Spatial Reasoning in VLMs from In-the-Wild Photos

    Leyuan Yu, Xiao Tang, Minghao Liu, Xinyuan Li, Xiaokai Bai, Sheng Zhou, Qunshu Lin, Weihao Xuan, Naoto Yokoya

    cs.CV · cs.AI · cs.CL

    Benchmarks for vision-language models (VLMs) mostly test observational spatial reasoning: models describe relations already visible in the input. Existing what-if tasks typically vary the observer while keeping the scene fixed. Can VLMs instead predict the consequences of hypothetically moving or rotating an object? We introduce MindEdit-Bench, a benchmark of six spatial reasoning tasks built from three-photo smartphone triplets of newly...

    arxiv.org/abs/2607.00491 · PDF

  27. 27

    StochasT: Learning with Stochastic Turn Depth for Visual Instruction Tuning

    Yuan Qing, Chengzhi Mao, Boqing Gong

    cs.CV · cs.CL · cs.LG

    Large Vision-Language Models (LVLMs) rely extensively on Visual Instruction Tuning (VIT) to elicit their multimodal reasoning capabilities. However, we find a discrepancy: VIT often packs multiple language tasks about the same image for conversational, multi-turn training, whereas existing benchmarks evaluate LVLMs in isolated, single-turn scenarios. The models can suffer from visual attention decay and contextual overfitting during...

    arxiv.org/abs/2607.00465 · PDF

  28. 28

    VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement

    Seohyun Lee, Seoung Choi, Dohwan Ko, Jongha Kim, Hyunwoo J. Kim

    cs.CV · cs.AI

    As video corpora continue to expand in both scale and task complexity, there is increasing demand for approaches that retrieve relevant videos from large-scale corpora (inter-video reasoning) and subsequently perform fine-grained, query-conditioned tasks (intra-video reasoning) within the retrieved content, such as temporal grounding. However, existing approaches typically treat retrieval as a preprocessing step, and consequently, when the...

    arxiv.org/abs/2607.00446 · PDF

  29. 29

    Information-Regularized Attention for Visual-Centric Reasoning

    Guohao Sun, Xiaofang Wang, Yash Patel, Mengchen Liu, Zhiqiang Tao, Praveen Krishnan

    cs.CV · cs.LG

    Vision-language models (VLMs) have become a paradigm for multimodal learning, yet remain unstable due to object hallucination, weak visual grounding, and catastrophic forgetting after full-parameter instruction tuning. We claim these failures result from a lack of explicit control over visual representation learning during the standard next-token prediction objective. As a result, visual embeddings thus become passively optimized and prone to...

    arxiv.org/abs/2607.00434 · PDF

  30. 30

    EO-VGGT: Orbital Ray-Conditioned 3D Foundation Models for Satellite Multi-View Reconstruction

    Qiyan Luo, Yingdong Pi, Lekang Wen, Jie Yang, Xiaoyu Wang, Haiming Zhang, Mi Wang

    cs.CV · cs.AI

    In the era of satellite constellations, multi-view optical satellite imagery is pivotal for Earth Observation (EO) and high-quality Digital Surface Model (DSM) reconstruction. Although feed-forward 3D foundation models have transformed computer vision, their deployment in satellite remote sensing is inherently constrained by the structural discrepancy between implicit perspective assumptions and explicit orbital pushbroom geometry. This...

    arxiv.org/abs/2607.00417 · PDF

  31. 31

    MindAU: EEG-Conditioned Facial Action Unit Editing via Dual-Stream Manifold Alignment

    Zhenhang Li, Xin Zhou, Hao Deng, Lijun Yin

    cs.CV · cs.LG

    Recent brain decoding studies have made substantial progress in reconstructing externally perceived visual content from neural signals. However, using electroencephalography (EEG) recordings to guide facial expression editing remains largely unexplored and poses a distinct challenge: rather than recovering what a subject sees, it requires identifying facial-action related patterns from noisy EEG signals and grounding them in localized,...

    arxiv.org/abs/2607.00410 · PDF

  32. 32

    The Illusion of High Utility in Safety Alignment of Text-to-Image Diffusion Models

    Adeel Yousaf, Soumik Ghosh, James Beetham, Amrit Singh Bedi, Mubarak Shah

    cs.CV · cs.AI · cs.LG

    Safety alignment of text-to-image (T2I) diffusion models aims to suppress harmful generations while preserving utility on benign prompts. Recent methods often appear to deliver high safety with high utility, but this conclusion rests largely on coarse global utility metrics (e.g., FID, CLIPScore) that are insensitive to fine-grained semantic correctness, creating an illusion of high utility. We show that when utility is measured with...

    arxiv.org/abs/2607.00402 · PDF

  33. 33

    Learning to Compose: Revisiting Proxy Task Design for Zero-Shot Composed Image Retrieval

    Jingjing Zhang, Lei Zhang, Zheren Fu, Zhendong Mao

    cs.CV · cs.AI · cs.CL · cs.IR · cs.MM

    Composed Image Retrieval (CIR) retrieves a target image from a reference image and a textual modification. While supervised CIR relies on costly triplets, Zero-Shot CIR (ZS-CIR) alleviates this reliance through proxy tasks trained on image-text pairs. However, existing proxy tasks primarily enhance visual and textual representations to accommodate a predefined composition mechanism such as pseudo-word injection into a frozen text encoder or...

    arxiv.org/abs/2607.00374 · PDF

  34. 34

    MEPA: Multi-Scale Representation Alignment for Visual Autoregressive Modeling with Mixture of Experts

    Nuoyan Zhou, Zhijun Tu, Lei Yu, Kun Cheng, Jie Hu, Nannan Wang, Xinghao Chen

    cs.CV · cs.AI

    Visual AutoRegressive modeling (VAR) has pioneered a coarse-to-fine multi-scale autoregressive generative paradigm, demonstrating strong capabilities in image generation. However, VAR still suffers from inherent deficiencies in multi-scale representation learning. Specifically, lower scales primarily capture global semantics, while higher scales focus on fine-grained details. Employing a shared architecture across scales induces optimization...

    arxiv.org/abs/2607.00371 · PDF

  35. 35

    RetailSMV: Exocentric vs. Egocentric Adaptation of Foundation Video World Models in Retail

    Amirreza Rouhi, Rajat Aggarwal, Parikshit Sakurikar, Anoop M. Namboodiri, Sashi P. Reddi

    cs.CV · cs.AI

    Foundation video diffusion models are increasingly viewed as world simulators for embodied agents, yet their pretraining on internet-scale generic video leaves them poorly aligned with real-world deployment domains. We study parameter-efficient adaptation of a pretrained foundation video world model to retail scenes: when synchronized egocentric and exocentric video of the same activity are available, which viewpoint of training data produces...

    arxiv.org/abs/2607.00310 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.