cs.CV · 2026-08-03 · No. 73

Computer Vision and Pattern Recognition, 2026-08-03.

22 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

22 entries
  1. 01

    TOOD: Task-Aware Out-of-Distribution Score Calibration for Continual Learners

    Mostafa ElAraby, Samer B. Nashed, Liam Paull

    cs.CV · cs.LG

    The primary challenge of continual learning (CL) systems is to learn new tasks while remaining performant on previously learned tasks. A similarly important though less well-studied aspect of CL systems is their ability to distinguish inputs that are unlikely to come from within the set of tasks the system has already encountered, often called out-of-distribution (OOD) detection. This paper presents several findings related to the dynamics of...

    arxiv.org/abs/2607.29592 · PDF

  2. 02

    TraceViT: Grounded Trace Supervision for Visual Abstract Reasoning

    Binnan Liu, Yechi Ma, Tian Xie, Wei Hua

    cs.CV · cs.AI

    The Abstraction and Reasoning Corpus (ARC) tests whether a model can infer an unseen transformation from a few input-output examples and apply it to a new grid. Looped visual reasoners refine predictions over multiple iterations, but conventional training constrains only the final output, leaving intermediate refinements unconstrained. We propose that these refinements should instead follow the transformation step by step. We introduce...

    arxiv.org/abs/2607.29586 · PDF

  3. 03

    Leveraging Transfer Learning with Class-Specific Decoders for Laparoscopic Segmentation

    Priya Tomar, Aditya Parikh, Christian Bauckhage, Rafet Sifa

    cs.CV · cs.LG

    Effective multi-organ segmentation in surgical data requires learning the intricate anatomical features and alleviating the challenge of class imbalance, which results from relatively lower proportions of small and limitedly exposed structures. Recent works on laparoscopic multi-organ segmentation focus on learning structure-specific features through class-specific decoder architectures and report favorable results. This work extends the...

    arxiv.org/abs/2607.29509 · PDF

  4. 04

    Lightweight Neural Networks for Affordance Segmentation: Enhancement of the Decoder Module

    Simone Lugani, Edoardo Ragusa, Rodolfo Zunino, Paolo Gastaldo

    cs.CV · cs.LG · cs.PF

    The deployment of deep neural networks for visual affordance segmentation on wearable robots poses may prove critical, due to some conflicting aspects of the problem. On one hand, affordance segmentation requires high-level abstraction capabilities, that typically involve large-size models. On the other hand, computing resources hosted on wearable robots prevent to run large-size models in real-time. The paper presents an analysis of the role...

    arxiv.org/abs/2607.29473 · PDF

  5. 05

    QR-Structured Thermal Triggers for Targeted Semantic Attacks on Infrared Vision-Language Models

    Xiang Chen, Yingying Zhao, Chao Li, Jiaju Han, Ben Zhang, Ang Li, Jiahuan Long, Yiwei Wei, Jiujiang Guo, Chengyin Hu

    cs.CV · cs.AI

    Infrared vision-language models (IR-VLMs) extend thermal perception to open-vocabulary classification, image captioning, and visual question answering. However, their robustness to structured thermal perturbations and the stability of cross-modal semantic alignment remain insufficiently studied. We propose QR-Structured Thermal Triggers (QR-STT), a stealthy, training-free, black-box framework for targeted semantic steering of IR-VLMs. QR-STT...

    arxiv.org/abs/2607.29445 · PDF

  6. 06

    Dense Temporal Contrast Synthesis via Conditioned Latent Transport

    Smriti Joshi, Apostolia Tsirikoglou, Daniel M. Lang, Richard Osuala, Noah Márquez Varaa, Alejandro Guzman, Grzegorz...

    cs.CV · cs.AI

    Dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) is essential for breast cancer management, but reliance on gadolinium-based contrast agents (GBCAs) restricts use in contraindicated populations, prolongs scan protocols, and presents environmental toxicity concerns. Contrast synthesis offers a non-invasive alternative; however, existing approaches struggle to balance spatial realism with temporal continuity, suffer from slow...

    arxiv.org/abs/2607.29394 · PDF

  7. 07

    DualDiT: A Conditional Dual-Output Diffusion Transformer for Joint OCT Image and Segmentation Mask Generation

    Fernando García-Torres, Rocío del Amor, Sandra Morales, Álvaro Barroso, Peter Heiduschka, Björn Kemper, Valery Naranjo

    cs.CV · cs.AI

    Background and Objective: Generating realistic medical images with anatomically accurate segmentation masks helps address the shortage of annotated data in medical imaging, particularly in optical coherence tomography (OCT) of mouse eyes, where manual retinal layer delineation is labour-intensive due to tiny structures and required expertise, resulting in scarce datasets. While diffusion models perform well in medical image synthesis, joint...

    arxiv.org/abs/2607.29337 · PDF

  8. 08

    OsteoCAD: A Human-in-the-Loop Cloud-Edge Framework for Bone Tumor Segmentation

    Maximo Rodriguez-Herrero, Dante D. Sanchez-Gallegos, Heriberto Aguirre-Meneses, Marco Antonio Núñez-Gaona, J. L....

    cs.CV · cs.AI

    Artificial Intelligence (AI) and Deep Learning (DL) have notably advanced medical image analysis, yet many health- care organizations struggle to adopt them due to limited com- putational resources and specialized expertise. To address these barriers, we introduce OsteoCAD, a modular eHealth framework that democratizes access to DL tools in clinical practice. Osteo- CAD delivers end-to-end DL capabilities-from dataset creation and...

    arxiv.org/abs/2607.29266 · PDF

  9. 09

    TAVI-TEC: An AI-Based Tool for Procedural Planning of Transcatheter Aortic Valve Implantation

    Alessandra Zerillo, Stefano Cannata, Diego Bellavia, Daniele Ciriello, Simone Manini, Salvatore Pasta, Caterina Gandolfo

    cs.CV · cs.AI · cs.LG

    Computed tomography angiography (CTA) is crucial for preprocedural TAVI planning, providing the anatomical information required for prosthesis sizing and vascular access assessment. As the volume of TAVI procedure increases, improving efficiency and standardizing annotations is becoming essential in clinical practice. This study presents TAVI-TEC, a fully automated artificial intelligence-based framework integrated into a web based DICOM...

    arxiv.org/abs/2607.29243 · PDF

  10. 10

    When Model Priors Conflict with Visual Evidence: Mitigating Commonsense-Driven Hallucinations by Selective Prior Calibration

    Kesheng Chen, Yamin Hu, Wenjian Luo

    cs.CV · cs.AI

    In vision--language models, commonsense-driven hallucination (CDH) occurs when a model's commonsense prior overrides clear visual evidence of an atypical state. For example, a model may report that a visibly six-fingered hand has five fingers. We show that these errors are systematically directed: when a model answers a question about a counterfactual (CF) image incorrectly, its answer often coincides with the candidate it prefers without...

    arxiv.org/abs/2607.29240 · PDF

  11. 11

    MoRAE: Flow-Friendly Self-Supervised Latents for Text-to-Motion Generation

    Yifei Zhu, Mingyi Shi, Yangyang Cai, Miao Cheng, Yoshifumi Kitamura, Taku Komura

    cs.CV · cs.AI

    Text-to-motion generation must produce motions that are semantically correct, temporally coherent, and physically plausible. A natural approach is to first project motion data into a structured semantic space and then train a generative model within that space. Such a paradigm has been highly successful in image generation through Representation Autoencoders (RAEs), where a frozen self-supervised encoder provides semantic features for...

    arxiv.org/abs/2607.29180 · PDF

  12. 12

    Have I Seen You? Embedding Behavior Signals Synthetic Face Dataset Membership

    Paweł Borsukiewicz, Daniele Lunghi, Wendkûuni C. Ouédraogo, Jacques Klein, Tegawendé F. Bissyandé

    cs.CV · cs.AI · cs.LG

    Synthetic face datasets are increasingly used to reduce privacy exposure and data access constraints in biometric recognition. Yet the generators that produce these datasets are trained on real faces, so synthetic data may still reveal their real source data. We study this risk through a dataset-level membership inference attack that first identifies the synthetic dataset used to train a face recognizer and then infers the real dataset used...

    arxiv.org/abs/2607.29144 · PDF

  13. 13

    SciFigPlag-Bench: A Benchmark for Provenance-Aware Scientific Figure Plagiarism Detection

    Zhiying Cui, Minghao Yang, Linlin Gao, Jie Liu, Pengyuan Li

    cs.CV · cs.LG

    Scientific figures often encode the visual evidence behind scientific findings, yet figure plagiarism remains underexplored as a benchmarked multimodal evaluation problem. We present SciFigPlag-Bench, a benchmark for provenance-aware reasoning over scientific figures in scholarly documents. Unlike general image-similarity or image-forensics benchmarks, SciFigPlag-Bench evaluates whether a suspicious figure reuses evidence from a specific...

    arxiv.org/abs/2607.29124 · PDF

  14. 14

    Adjudicated Captioning: Multi-Agent Alignment Scoring and Consensus-Distilled Beam Arbitration for Strict Zero-Shot Image Captioning

    Duy Tran Thanh, Thien-Phuc Doan, Long Nguyen-Vu, Ngo Tan Vu Khanh

    cs.CV · cs.AI · cs.CL · cs.MA

    Zero-shot image captioning (ZIC) describes images without paired image-caption supervision during captioner training, relying on text-only corpora and frozen pretrained image-text scorers. Existing retrieval-augmented methods score image-text alignment once, at retrieval, then commit the captioner's autoregressive beam under language-model probability alone, leaving the decoder without further visual grounding feedback. Progress has stalled,...

    arxiv.org/abs/2607.28986 · PDF

  15. 15

    RAID: Towards Robust AI-Generated Image Detection with Bit-Reversed Images

    Renxi Cheng, Jie Gui, Hongsong Wang

    cs.CV · cs.AI

    The rapid advancement of image generation models has made it increasingly difficult for people to distinguish AI-generated images from real ones. To prevent the potential risks associated with the misuse of fake images, AI-generated image detection has gained significant attention. Existing methods neglect the inherent differences between real and fake images, thus lacking robustness and generalization ability. In this work, we innovatively...

    arxiv.org/abs/2607.28974 · PDF

  16. 16

    Visual Distribution Anchoring for Efficient Prompt Tuning

    Pouya Parsa, Raoof Zare Moayedi, Seongjin Choi

    cs.CV · cs.LG

    Prompt tuning adapts vision--language models with few trainable parameters, but existing approaches trade off efficiency and adaptation: static textual prompts can overfit source classes, image-conditioned prompts add per-instance computation, and multimodal tuning modifies the visual branch. We propose VDA (Visual Distribution Anchoring), a training-free target adaptation framework that augments a frozen semantic classifier with class-level...

    arxiv.org/abs/2607.28967 · PDF

  17. 17

    Retrieval-Driven Training-Free AI-Generated Video Attribution

    Renxi Cheng, Chaolei Han, Jie Gui, Hongsong Wang

    cs.CV · cs.AI

    AI-generated videos are becoming increasingly realistic and difficult to distinguish from authentic ones, which facilitates malicious misuse and poses growing threats to cybersecurity and social governance. Attributing AI-generated videos to their specific generative sources is therefore of critical importance for forensic investigation and legal regulation. However, most existing visual attribution methods focus on images and particularly...

    arxiv.org/abs/2607.28955 · PDF

  18. 18

    DiffAttack: Evasion Attacks Against Face Recognition via Latent Diffusion Models

    Omid Ahmadieh, Nima Karimian

    cs.CV · cs.AI

    Facial biometric identification relies on the distinctiveness of user attributes within a high-dimensional embedding space. However, the decision boundaries of deep face recognition (FR) systems are often sufficiently narrow that they can be conflated, rendering the models vulnerable to adversarial attacks. In such scenarios, the FR system fails to distinguish between an authentic source and a meticulously crafted adversarial face. Existing...

    arxiv.org/abs/2607.28936 · PDF

  19. 19

    A Unified Benchmark of Deep Learning Models for Multi-task 3D Brain Tumor Segmentation from Magnetic Resonance Imaging

    Diego J. Torrejón, Luna Y. Hernández, Javier Sánchez

    cs.CV · cs.AI

    Automatic brain tumor segmentation from magnetic resonance imaging (MRI) has become a fundamental task in computer-assisted diagnosis, treatment planning, and disease monitoring. Although numerous deep learning architectures have recently been proposed, objective comparisons remain challenging because published studies often employ different datasets, preprocessing strategies, training protocols, and evaluation procedures. This work presents...

    arxiv.org/abs/2607.28858 · PDF

  20. 20

    WaiT for the Signal: Simple Frequency-Aware Flow-Matching

    Krunoslav Lehman Pavasovic, Théophane Vallaeys, Stéphane Mallat, Giulio Biroli, Luke Zettlemoyer, Brian Karrer, Jakob Verbeek

    cs.CV · cs.AI · cs.LG · stat.ML

    As image generation models scale to ever higher resolutions, global coherence, local detail, and texture fidelity become critical axes for generation quality. However, standard flow matching treats all spatial frequencies uniformly, ignoring the natural frequency hierarchy where high-frequency bands become indistinguishable from pure noise far earlier than coarse structures. We introduce WaiT, a Wavelet-aware image Transformer that decomposes...

    arxiv.org/abs/2607.28760 · PDF

  21. 21

    SCMA: Structure-Conditioned and Metal-Aware Flow Matching for CT Metal Artifact Reduction

    Heran Wang, Jianing Sun, Xu Jiang, Genwei Ma, Xing Zhao, Jigang Duan

    cs.CV · cs.AI

    In X-ray CT, metallic objects cause beam hardening, photon starvation, and scattering, leading to projection inconsistency, streaks, dark bands, and structural distortions that compromise clinical diagnosis and quantitative analysis. Existing metal artifact reduction (MAR) methods remain limited: optimization-based methods may leave residual artifacts or blur structures, regression networks may generalize poorly across scenarios, and...

    arxiv.org/abs/2607.28759 · PDF

  22. 22

    ReToken: One Token to Improve Vision-Language Models for Visual Retrieval

    Yao Xiao, Reuben Tan, Zhen Zhu, Yuqun Wu, Jianfeng Gao, Derek Hoiem

    cs.CV · cs.AI · cs.LG

    Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once is computationally infeasible under GPU memory constraints. We present ReToken, a single learnable embedding trained as an explicit retrieval target that selects a sparse set of query-relevant visual tokens from a pre-filled visual KV cache. Trained on only a small image-QA dataset,...

    arxiv.org/abs/2607.28627 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.