cs.CV · 2026-06-23 · No. 32

Computer Vision and Pattern Recognition, 2026-06-23.

20 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

20 entries
  1. 01

    Semantic Browsing: Controllable Diversity for Image Generation

    Sara Dorfman, Maya Vishnevsky, Omer Dahary, Or Patashnik, Daniel Cohen-Or

    cs.CV · cs.AI · cs.GR · cs.LG

    Modern text-to-image models excel in visual fidelity and prompt adherence. However, this strict adherence comes at the cost of diversity: generated samples tend to collapse into a single visual interpretation. Existing methods to improve diversity produce outputs driven by incidental variations rather than meaningful design choices. This motivates a new variant of the diversity task where structure is enforced on the generated samples. We...

    arxiv.org/abs/2606.23679 · PDF

  2. 02

    AIR: Adaptive Interleaved Reasoning with Code in MLLMs

    Cong Han, Xiaohan Lan, Haibo Qiu, Yujie Zhong

    cs.CV · cs.AI

    Following the paradigm shift initiated by OpenAI o3, interleaved reasoning with code to enhance multimodal large language models (MLLMs) has become a pivotal research frontier. The existing literature focuses primarily on tool-use within vision-perception tasks. However, such approaches typically rely on predefined heuristics for visual manipulation and are inherently incapable of addressing numerical computation problems due to their...

    arxiv.org/abs/2606.23678 · PDF

  3. 03

    Hedgementation = Hedgerow Segmentation: A Remote Sensing Benchmark

    Nathan Senyard, Salem Hamdani, Astrid Zhang, Derek Wang, Evan Shelhamer, Mathias Lécuyer, Joséphine Gantois

    cs.CV · cs.LG

    We propose Hedgementation: a new benchmark to evaluate machine learning models for hedgerow mapping from remote sensing data at country scale and 10m$^2$ spatial resolution. We combine and harmonize multiple remote sensing data products and ground truth labels sourced from a hedgerow inventory in France. We measure the ability of three baseline models to generalize across spatial distance, and across climatic zones, a more explicitly...

    arxiv.org/abs/2606.23615 · PDF

  4. 04

    Data Selection Through Iterative Self-Filtering for Vision-Language Settings

    Andrei Liviu Nicolicioiu, Sarvjeet Singh Ghotra, Morgane M. Moss, Aaron Courville

    cs.CV · cs.AI · cs.LG

    The availability of large amounts of clean data is paramount to training neural networks. However, at large scales, manual oversight is impractical, resulting in sizeable datasets that can be very noisy. Attempts to mitigate this obstacle to producing performant vision-language models have so far involved heuristics, curated reference datasets, and using pre-trained models. Here we propose a novel, bootstrapped method in which a CLIP model is...

    arxiv.org/abs/2606.23611 · PDF

  5. 05

    Polycepta: Object-Centric Appearance Estimation for Multi-Object Tracking

    Mohamed Nagy, Naoufel Werghi, Jorge Dias, Majid Khonji

    cs.CV · cs.AI

    The tracking-by-detection paradigm in multi-object tracking (MOT) typically relies on static appearance descriptors to complement motion estimation. However, these descriptors are frame-independent, limiting their robustness as visual cues. Since such descriptors are often obtained from computationally intensive pretrained backbones, real-time MOT systems frequently abandon appearance cues altogether and rely solely on motion prediction and...

    arxiv.org/abs/2606.23604 · PDF

  6. 06

    Rethinking Object-Centric Representations for Video Dynamics Modeling

    Amaury Wei, Ismail Nejjar, Olga Fink

    cs.CV · cs.AI · cs.LG

    Unsupervised video object tracking aims to decompose dynamic scenes into persistent, object-centric entities without manual annotations. Many recent approaches rely on slot-based representations, where a fixed set of latent variables ("slots") represent individual objects across frames. To preserve object identity, these models enforce temporal consistency on slot embeddings. However, when appearance and pose are entangled, this consistency...

    arxiv.org/abs/2606.23436 · PDF

  7. 07

    Changing Modalities: Adapting Remote Sensing Models to New Satellites and Sensors

    Tim G. Zhou, Anthony Fuller, Geoff Pleiss, Evan Shelhamer

    cs.CV · cs.LG

    Machine learning models for remote sensing are trained and deployed on a static set of modalities. However, as we equip newer satellites with novel sensors and retire old ones, practitioners may wish to deploy an existing model on a substitution, superset, or subset of modalities with minimal retraining given data availability or practical computational constraints. We study the setting of updating existing models to changing modalities and...

    arxiv.org/abs/2606.23356 · PDF

  8. 08

    VideoAgent: All-in-One Framework for Video Understanding and Editing

    Hengji Zhou, Lingxuan Huang, Jian Wang, Bing Zhou, Si Wu, Lianghao Xia, Chao Huang

    cs.CV · cs.AI

    Video editing has become essential in digital media creation, yet existing automated systems are restricted to short segment processing and domain-specific tasks. They face two critical limitations: i) inability to handle diverse video comprehension and editing operations, and ii) lack of long-video understanding for coherent narrative creation. We propose VideoAgent, an all-in-one agentic framework addressing these challenges through two key...

    arxiv.org/abs/2606.23327 · PDF

  9. 09

    Transfer learning-based method for automated ewaste recycling in smart cities

    Nermeen Abou Baker, Paul Szabo-Müller, Uwe Handmann

    cs.CV · cs.LG

    Sorting a huge stream of waste accurately within a short period can be done with the support of digitalization, particularly Artificial Intelligence, instead of traditional methods. The overlap of Artificial Intelligence and Circular Economy can flourish many services in the environmental technology domain, in particular smart ewaste recycling, resulting in enabling circular smart cities. We analyse the growing need for automated ewaste...

    arxiv.org/abs/2606.23286 · PDF

  10. 10

    P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture

    Felix Tristram, Stefano Gasperini, Benjamin Killeen, Marcel Walch, Christian Benz, Nassir Navab, Ghazal Ghazaei

    cs.CV · cs.AI

    The increasing maturity of embodied AI platforms has driven a growing interest in procedural video representation learning to support intelligent assistance systems for complex, multi-step tasks. Leveraging large-scale latent predictive training, video foundation models capture video dynamics, enabling downstream tasks such as activity understanding, spatiotemporal localization, and predictive control. However, procedural videos include...

    arxiv.org/abs/2606.23256 · PDF

  11. 11

    SteerVTE: Seamless Video Text Editing with Style and Glyph Control

    Kai Zeng, Moran Li, Zhengwei Wang, Yingchen Yu, Yiheng Lin, Ruichuan An, Ming Lu, Qi She, Wentao Zhang

    cs.CV · cs.AI

    Visual text editing aims to precisely modify text in images and videos while preserving stylistic consistency and visual realism. Despite significant advances in the image domain, video text editing remains largely unexplored: it is a localized task demanding stroke-level precision within small text regions, which compounds the challenges of cross-frame accuracy, temporal coherence, and stylistic fidelity. We introduce SteerVTE, a unified...

    arxiv.org/abs/2606.23254 · PDF

  12. 12

    RS-Gen: A Multi-Stage Agentic Framework for Reasoning and Search-Augmented Image Generation

    Feifei Bian, Zhimin Zheng, Wei Deng, Daiguo Zhou, Jian Luan

    cs.CV · cs.AI

    Recent years have witnessed remarkable progress in image generation and editing, particularly regarding instruction following and visual fidelity. However, when handling ambiguous intentions, logical reasoning, and Out-of-Distribution (OOD) knowledge, existing image models often yield sub-optimal results due to a lack of deep reasoning capabilities and real-time external information. Although emerging unified understanding-and-generation...

    arxiv.org/abs/2606.23221 · PDF

  13. 13

    Spectral Gating via Damped Oscillations for Adaptive Implicit Neural Representations

    Alex Costanzino, Pierluigi Zama Ramirez, Giuseppe Lisanti, Luigi Di Stefano

    cs.CV · cs.LG

    Implicit Neural Representations (INRs) have been proven successful in encoding continuous signals through coordinate-based networks, yet facing a spectral dilemma: periodic activations capture fine details but act as all-pass filters that memorise noise, while spatially compact activations regularise effectively but suffer from low-frequency bias. Existing attempts to resolve this trade-off introduce computational overhead or tuning frailty....

    arxiv.org/abs/2606.23129 · PDF

  14. 14

    Interpretable Probabilistic Medical Image Segmentation via Gaussian Process with Explicit Modelling of Annotation Bias and Variability

    Qi Li, Yuliang Huang, Shaheer U. Saeed, Qianye Yang, Vasilis Stavrinides, Zachary M. C. Baum, Dean C. Barratt, J....

    cs.CV · cs.AI

    Deep learning-based medical image segmentation models are trained using annotations that exhibit systematic bias and variability across raters. While probabilistic multi-rater approaches can emulate annotator-specific delineations, annotator characteristics are typically encoded implicitly in deep latent feature space, making direct analysis of their influence on predictive distributions less straightforward. We propose a logit-space...

    arxiv.org/abs/2606.23177 · PDF

  15. 15

    Attention-Spectrum Regularization for Replay-Free Continual Multimodal LLMs

    Chuangxin Zhao, Canran Xiao, Siyuan Ma, Mengyao Lyu, Yanbiao Ma, Jun Xia, Guiguang Ding, Yang Liu

    cs.CV · cs.AI

    Multimodal large language models (MLLMs) are increasingly required to adapt to non-stationary streams of visual domains, question types, and user instructions, yet continual fine-tuning often causes severe forgetting of previously acquired multimodal skills. Existing continual vision-language methods mainly preserve outputs, replay data or pseudo-data, regularize embedding geometry, or allocate task-specific parameters, but they provide...

    arxiv.org/abs/2606.23063 · PDF

  16. 16

    MotionHalluc: Diagnosing Kinematic Hallucinations in Fine-Grained Motion Reasoning

    Weile Guo, Shenghong He, Danying Mo, Chengdong Xu, Xuexun Liu, Chao Yu

    cs.CV · cs.AI

    Motion instruction generation in cross-video comparison aims to produce corrective feedback that describes the differences between a query and a reference motion. However, existing models often generate instructions that exhibit motion hallucinations, failing to reflect actual kinematic differences between paired videos. To systematically investigate these hallucinations, we introduce MotionHalluc, a dedicated benchmark for evaluating motion...

    arxiv.org/abs/2606.23061 · PDF

  17. 17

    Physics-Guided Spatiotemporal State Space Modeling for Lookahead Molten Pool Segmentation in Laser Wire-Feed Welding

    Sen Li, Haichao Cui, Changhao Yin, Chendong Shao, Yaqi Wang, Xinhua Tang, Fenggui Lu

    cs.CV · cs.AI

    Real-time weld-pool perception is critical for closed-loop control in laser wire-feed welding, where sensing, computation, and actuator response introduce unavoidable delay. This paper presents a physics-guided spatiotemporal state space network for lookahead weld-pool segmentation. The model uses historical coaxial grayscale images, welding process parameters, and aligned wire-state electrical signals to predict the future semantic layout of...

    arxiv.org/abs/2606.23028 · PDF

  18. 18

    ScalingAttention: Discovering Intrinsic Sparse Attention Topology for Video Diffusion Transformers

    Ruiliang Zhou, Xuecheng Wu, Kang He, Guangyun Han, Bin Liu, Qinqin Chen, Wende Xu, Qingjie Zhao, Chengru Song

    cs.CV · cs.AI

    While Diffusion Transformers (DiTs) have revolutionized high-fidelity video generation, their reliance on 3D full attention creates a quadratic computational bottleneck. Existing sparse methods face a dilemma: dynamic pruning suffers from prohibitive runtime overhead and memory fragmentation, while static heuristics fail to capture fine-grained dependencies. In this work, we propose ScalingAttention, a training-free framework grounded in a...

    arxiv.org/abs/2606.23019 · PDF

  19. 19

    From Point Estimates to Distributions: GMM Pooling for MIL in Preterm Birth Prediction

    Hussain Alasmawi, Numan Saeed, Soha Said, Mohammad Yaqub

    cs.CV · cs.LG

    Preterm birth (PTB) prediction can enable targeted surveillance and timely intervention, yet most ultrasound-based models use a single selected transvaginal ultrasound (TVUS) frame per patient despite routine exams acquiring multiple cervical images. We formulate PTB prediction as a multiple instance learning (MIL) problem, representing each patient as a variable-sized bag of TVUS images with a single outcome label. To move beyond standard...

    arxiv.org/abs/2606.23005 · PDF

  20. 20

    Subject-Level Unknown-Identity Identification from Leap Motion Controller 2 Hand Landmarks

    Bahar Moharrer, Susanna Cifani, Marco Raoul Marini, Luigi Cinque, Maria De Marsico

    cs.CV · cs.LG

    This work studies subject recognition from Leap Motion Controller 2 (LMC2) hand landmark data under a subject-level unknown-identity identification protocol on the Multi View Leap2 Hand Pose (ML2HP) dataset. Using only the landmark modality, we retain the original geometric representation and enrich it with fingertip-to-palm distances and palm-normalized inter-finger angular descriptors. Evaluation is performed under a Leave-One-Subject-Out...

    arxiv.org/abs/2606.22986 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.