cs.CV · 2026-06-25 · No. 34

Computer Vision and Pattern Recognition, 2026-06-25.

24 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

24 entries
  1. 01

    A cross-process welding penetration status prediction algorithm based on unsupervised domain adaptation in laser and TIG welding

    Sen Li, Haichao Cui, Chendong Shao, Yaqi Wang, Xinhua Tang

    cs.CV · cs.AI

    Supervised deep learning has been widely used for weld penetration state classification; however, its performance often degrades significantly under domain shift, such as when transferring models between welding processes with distinct physical mechanisms:for instance, from arc-dominated tungsten inert gas (TIG) welding to keyhole-based laser welding. To overcome this limitation, we propose an unsupervised domain adaptation (UDA) framework...

    arxiv.org/abs/2606.26078 · PDF

  2. 02

    A welding penetration prediction model for laser welding process based on self-supervised learning using physics-informed neural networks

    Sen Li, Xiaoying Liu, Xiaojian Xu, Chendong Shao, Yaqi Wang, Ling Lan, Xinhua Tang, Haichao Cui

    cs.CV · cs.AI

    The laser welding full-penetration is of critical importance, as it constitutes one of the fundamental factors in achieving defect-free welded joints. Accurate prediction of the penetration state is therefore essential for ensuring weld quality. To this end, this paper introduces SimPhysNet, a novel algorithm that achieves high classification accuracy in laser welding penetration prediction using only a limited number of labelled images. This...

    arxiv.org/abs/2606.26059 · PDF

  3. 03

    TriViewBench: Controlled Complexity Scaling for Multi-View Structural Reasoning in MLLMs

    Yu-Yang Chen, Lan-Zhe Guo

    cs.CV · cs.AI

    Multimodal Large Language Models (MLLMs) demonstrate strong performance on standard visual question answering benchmarks, yet their scalability under controlled structural complexity remains poorly understood. We introduce TriViewBench, a controlled three-view visual reasoning benchmark constructed from synthetic 3D scenes with explicitly parameterized object count and occlusion. The benchmark contains 1,923 scenes and over 14K...

    arxiv.org/abs/2606.26029 · PDF

  4. 04

    Taxonomy-aware deep learning for hierarchical marine species classification in underwater imagery

    Dan Zimmerman, Dimitris A. Pados, George Sklivanitis

    cs.CV · cs.LG

    Automated classification of marine species from underwater imagery is essential for scalable ocean biodiversity monitoring and conservation policy. Existing approaches struggle with severe domain shift across collection platforms, fine-grained visual similarity between closely related species, and uneven annotation granularity, where many specimens can only be identified to genus or a coarser taxonomic rank. We present a taxonomy-aware deep...

    arxiv.org/abs/2606.25989 · PDF

  5. 05

    Pulmonary Embolism Risk Stratification from CTPA and Medical Records: Vascular Graphs Are Not All You Need

    Nathan Painchaud, Tristan Habémont, Morgane des Ligneris, Allan Serva, Pierre Croisille, Laurent Bertoletti, Thomas...

    cs.CV · cs.AI · cs.LG · eess.IV

    Risk stratification for pulmonary embolism (PE) is critical for clinical decision-making. Stratification guidelines are based on patient medical records, parameters measured from computed tomography pulmonary angiography (CTPA), and blood tests. However, blood tests are often missing in routine practice. This work studies whether state-of-the-art models can accurately classify risk stratification from only medical records and biomarkers...

    arxiv.org/abs/2606.25956 · PDF

  6. 06

    Enhancing Brain MRI Anomaly Detection and Reasoning with ROI Rethink and Synthetic Data

    Shangkun Li, Jie Xu, Yi Guo, Zeju Li, Yuanyuan Wang

    cs.CV · cs.AI

    Medical vision-language models typically generate diagnoses through single-pass inference without indicating which image regions support their conclusions. This lack of spatial grounding limits clinical utility: outputs cannot be audited, and models may hallucinate findings on normal scans. We present BrReMark (Brain Rethink via ROI Marking), a framework that introduces explicit region marking into brain MRI diagnosis. The model first...

    arxiv.org/abs/2606.25894 · PDF

  7. 07

    Edges Before Embeddings: A Confidence-Aware Blur Gate for Vision-Language Pipelines

    Duy Tran Thanh

    cs.CV · cs.AI

    Production vision pipelines silently degrade on blurry input, wasting compute on downstream OCR, retrieval, and vision-language model (VLM) calls that cannot recover a usable output. We present MagikaDocumentFromPixel, a lightweight, CPU-friendly image quality gate that classifies a single image as sharp, blurred, or uncertain in roughly 7 ms on a single CPU core. The contributions are (i) a recipe selected from a 46-configuration, 8-sweep...

    arxiv.org/abs/2606.25838 · PDF

  8. 08

    Point Cloud Diffusion with Global and Local Reconstruction for Instance-Level 3D Anomaly Detection

    Linchun Wu, Qin Zou, Jiwen Lu, Qingquan Li

    cs.CV · cs.AI

    3D anomaly detection in point clouds is critical for high-precision industrial manufacturing. Reconstruction-based methods have laid a strong foundation by detecting 3D anomalies through comparisons between defective inputs and their reconstructed normal counterparts. However, existing methods still suffer from two challenges: 1) the foreground weak defective regions such as scratches are hard to reconstruct and detect, where the anomaly...

    arxiv.org/abs/2606.25740 · PDF

  9. 09

    Steering Vision-Language Models with Joint Sparse Autoencoders

    Huizhen Shu, Xuying Li, Hongxu Lin, Wenjie Sun, Hui Li

    cs.CV · cs.AI

    Sparse Autoencoders (SAEs) have shown promise for analyzing language models, but applying them to vision-language models (VLMs) often yields representations that are difficult to use as controllable cross-modal steering directions. We introduce the Joint Sparse Autoencoder (JSAE), which uses an explicit alignment constraint to jointly factorize sequence-pooled vision and language activations into shared, interpretable image/caption-level...

    arxiv.org/abs/2606.25657 · PDF

  10. 10

    Expresso-AI: Explainable Video-Based Deep Learning Models for Depression Diagnosis

    Felipe Moreno, Sharifa Alghowinem, Hae Won Park, Cynthia Breazeal

    cs.CV · cs.AI · cs.LG

    Given the widespread prevalence of depression and its consequential impact on individuals and society, it is crucial to obtain objective measures for early diagnosis and intervention. As a multidisciplinary topic, these objective measures should be interpretable and accessible to health care professionals, ensuring effective collaboration and treatment planning in the realm of mental health care. Even though current automated depression...

    arxiv.org/abs/2606.25606 · PDF

  11. 11

    Concept Removal for Frontier Image Generative Models

    Aditya Kumar, Pierre Joly, Adam Dziedzic, Franziska Boenisch

    cs.CV · cs.LG

    Image generative models are trained on massive, largely uncurated internet-scale datasets that contain undesirable visual concepts. Efficiently removing such concepts from the model generations without degrading the quality of output images remains challenging. We introduce a novel concept removal method for frontier diffusion and image autoregressive models, such as SD3.5, Flux, and Infinity. Our intervention replaces the internal bottleneck...

    arxiv.org/abs/2606.25548 · PDF

  12. 12

    HG-Bench: A Benchmark for Multi-Page Handwritten Answer-Region Grounding in Automated Homework Assessment

    Chuangxin Zhao, Boyan Shi, Yanling Wang, Yijian LU, Canran Xiao, Jiali Chen, Jun Xia, Yan Wang, Ji Qi, Juanzi Li

    cs.CV · cs.AI

    Automated homework assessment depends not only on recognizing student answers, but also on accurately locating where each answer and each intermediate reasoning step appears in noisy, multi-page handwritten work. This paper addresses the missing evaluation setting of page-aware, two-level answer-region grounding: given a sequence of homework page images, a model must localize complete answer regions and their ordered step-level subregions. We...

    arxiv.org/abs/2606.25491 · PDF

  13. 13

    Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models

    Kaiwen Zheng, Guande He, Min Zhao, Jintao Zhang, Huayu Chen, Jianfei Chen, Chen-Hsuan Lin, Ming-Yu Liu, Jun Zhu, Qianli Ma

    cs.CV · cs.LG

    Autoregressive video diffusion with causal diffusion transformers has emerged as a major paradigm for real-time streaming video generation and action-conditioned interactive world models. In this work, we extend rCM, an advanced diffusion distillation framework, to autoregressive video diffusion. The core philosophy of rCM lies in the complementarity between forward and reverse divergences, represented by consistency models (CMs) and...

    arxiv.org/abs/2606.25473 · PDF

  14. 14

    EchoStyle: Unlocking High-Fidelity Video Stylization with Reverse Data Synthesis

    Huaqiu Li, Jiahao Wang, Sijia Cai, Hualian Sheng, Bing Deng, Jieping Ye, Wenhan Luo

    cs.CV · cs.AI

    While image stylization has been studied extensively, video stylization remains a critical and largely unsolved challenge in the field of intelligent content creation. Existing methods, usually utilizing a reference image as the style prior, suffer from content leakage, data scarcity and limited adaptability to long videos, leading to suboptimal results with severe style drift and motion distortion. For these issues, we present EchoStyle, a...

    arxiv.org/abs/2606.25465 · PDF

  15. 15

    C3-Bench: A Context-Aware Change Captioning Benchmark

    Jae-Woo Kim, Hyeongbeom Kim, Ue-Hwan Kim

    cs.CV · cs.AI

    While Change Captioning systems have garnered substantial attention to respond to our evolving world, their true performance on diverse real-world change contexts remains largely unexplored due to the lack of comprehensive evaluation frameworks. To fill this gap, we propose C3-Bench, a comprehensive benchmark for evaluating Context-aware Change Captioning. C3-Bench features: (1) 4,996 human-labeled image pairs of 51 real-world change contexts...

    arxiv.org/abs/2606.25445 · PDF

  16. 16

    Anatomically-conditioned Latent Diffusion Model for Data-Efficient Few-Shot Cross-Domain 3D Glioma MRI Synthesis

    Salman Shaik, Truong Thanh Hung Nguyen, Hung Cao

    cs.CV · cs.AI

    Accurate classification of diffuse gliomas is often hindered by domain shifts across centers and a lack of large, annotated datasets. We propose the Anatomically-conditioned Latent Diffusion Model (ALDM), a novel framework for data-efficient, few-shot 3D volumetric MRI synthesis. ALDM utilizes a two-stage approach: a 3D variational autoencoder learns anatomical priors from a data-rich source domain, while a conditional latent diffusion model,...

    arxiv.org/abs/2606.25390 · PDF

  17. 17

    Beyond Visual Forensics: Auditing Multimodal Robustness for Synthetic Medical Image Detection

    Ching-Hao Chiu, Hao-Wei Chung, Gelei Xu, Xueyang Li, Pin-Yu Chen, John Kheir, Meysam Ghaffari, Carlos Morato, Ahmed...

    cs.CV · cs.AI

    With the rapid adoption of generative AI, synthetic medical images pose growing risks, including diagnostic deception and insurance fraud. Although prior work has explored vision-language model (VLM)-based synthetic image detection, these evaluations typically consider images in isolation. In clinical practice, however, images are interpreted alongside structured records and metadata, and VLMs are increasingly deployed under joint...

    arxiv.org/abs/2606.25375 · PDF

  18. 18

    State Space Models Meet Remote Sensing: A Survey

    Qinzhe Yang, Chenyang Liu, Jia Xu, Zhenwei Shi, Zhengxia Zou

    cs.CV · cs.LG

    State Space Models (SSMs), designed for long-range modeling, offer linear computational complexity and strong capabilities in capturing long-range dependencies. In the field of remote sensing, SSMs have gained popularity due to their effectiveness in addressing unique challenges such as dense visual predictions, multi-modal remote sensing data, and temporal remote sensing data, which have also yielded significant advancements in customized...

    arxiv.org/abs/2606.25329 · PDF

  19. 19

    REViT: Roto-reflection Equivariant Convolutional Vision Transformer

    Sheir A. Zaheer, Alexander C. Holston, Chan Y. Park

    cs.CV · cs.LG

    In this paper, we propose a discrete roto-reflection group equivariant vision transformer with convolutional attention. Roto-reflection equivariant networks preserve the rotational, flip and positional symmetry in feature maps, making them useful for tasks where orientation of the inputs is relevant to the model outputs. In image classification and object detection, most of the studies on roto-reflection equivariant models have focused on...

    arxiv.org/abs/2606.25318 · PDF

  20. 20

    ESTANet: Efficient Online Error Detection in Procedural Videos via Prediction Inconsistency

    Shih-Po Lee, Reza Ghoddoosian, Faizan Siddiqui, Enna Sachdeva, Behzad Dariush

    cs.CV · cs.AI

    An efficient and accurate system for detecting errors in procedural tasks is crucial for supporting human needs in daily life, as it can provide instant notifications and guide people to correct mistakes. In this work, we study real-time online error detection in procedural videos from a simple but overlooked perspective: the prediction behavior of action detectors themselves. Instead of designing complex architectures or specialized...

    arxiv.org/abs/2606.25317 · PDF

  21. 21

    Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation

    Atin Pothiraj, Jaemin Cho, Yue Zhang, Elias Stengel-Eskin, Mohit Bansal

    cs.CV · cs.AI

    Video generation models are increasingly capable of producing realistic videos, but they still struggle to generate videos that follow basic physical laws. Compounding this is a lack of reliable granular evaluation methods for localizing and specifying physical law violations in videos. We address this by introducing Physics Question Scene Graph (PQSG), a hierarchical question-based evaluation pipeline. PQSG evaluates generated videos by...

    arxiv.org/abs/2606.25306 · PDF

  22. 22

    Heterogeneous and Adept Snapshot Distillation for 3D Semantic Segmentation

    Xiaopei Wu, Yuenan Hou, Junkai Xu, Wenxiao Wang, Binbin Lin, Yu Li, Ping Li, Haifeng Liu, Deng Cai, Wanli Ouyang

    cs.CV · cs.AI

    Multi-modal fusion and multi-model ensembling are prevalent in enhancing the performance of 3D semantic segmentation. Despite the impressive performance, these methods either rely on auxiliary input signals or suffer from costly computational expense. To efficaciously enhance the segmentation performance without introducing intolerable costs, we propose to transfer the rich knowledge from the multi-modal model (i.e., point clouds and images)...

    arxiv.org/abs/2606.25278 · PDF

  23. 23

    Pre-Warm: Input-Conditioned Weight Initialization for Convolutional Neural Networks

    Rowan Martnishn

    cs.CV · cs.LG

    We introduce Pre-Warm, a simple yet effective zero-training-cost method for data-conditioned initialization of the first convolutional layer. Before the first forward pass, Pre-Warm extracts mean-centered local patches from a single training batch, clusters them with MiniBatchKMeans, applies inverse Manhattan spatial weighting, and uses the resulting centroids to initialize half of the first-layer filters (the remainder retain Kaiming...

    arxiv.org/abs/2606.25256 · PDF

  24. 24

    MJEPA: A Simple and Scalable Joint-Embedding Predictive Architecture for Audio-Visual Learning

    Revant Teotia, Adrien Bardes, Michael Rabbat, Sumit Chopra, Matthew J. Muckley, Nicolas Ballas

    cs.CV · cs.LG

    Self-supervised learning from large-scale video data has emerged as a dominant paradigm for visual representation learning. Since audio and visual streams naturally co-occur in video data, extending this success to jointly learn from both modalities is a natural next step, yet it remains challenging. Existing audio-visual self-supervised methods rely on modality-specific encoders and complex combinations of contrastive or reconstruction...

    arxiv.org/abs/2606.25225 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.