cs.CV · 2026-06-16 · No. 25

Computer Vision and Pattern Recognition, 2026-06-16.

15 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

15 entries
  1. 01

    The Importance of Phase in Neural Representations: An Internal Oppenheim-Lim Test of Image Classifiers

    Alper Yıldırım

    cs.CV · cs.AI · cs.LG

    Oppenheim and Lim (1981) showed that natural images stay recognizable when reconstructed from their Fourier phase alone, while the magnitude carries little of their identity. We ask whether trained image classifiers reproduce this asymmetry inside their hidden layers, and we test it causally: given two images, we transplant the phase of one onto the magnitude of the other at a chosen layer and record which image the prediction follows. In...

    arxiv.org/abs/2606.17037 · PDF

  2. 02

    FusionRS: A Large-Scale RGB-Infrared Remote Sensing Dataset for Dual-Modal Vision-Language Foundation Models

    Jiaju Han, Ben Zhang, Xuemeng Sun, Qike Zhang, Yuxian Dong, Chengyin Hu, Fengyu Zhang, Yiwei Wei, Jiujiang Guo

    cs.CV · cs.AI

    Remote sensing vision-language models have advanced Earth observation understanding, but most existing work remains centered on RGB imagery, leaving the complementary information in infrared data underexplored. Infrared images provide distinctive cues, including thermal intensity structures, object boundaries, and illumination-invariant scene features, which can enrich visual-language learning beyond conventional RGB observations. However, a...

    arxiv.org/abs/2606.17020 · PDF

  3. 03

    ActiveSAM: Image-Conditional Class Pruning for Fast and Accurate Open-Vocabulary Segmentation

    Tran Dinh Tien, Zhiqiang Shen

    cs.CV · cs.AI · cs.LG

    Segment Anything Model 3 (SAM 3) provides a strong frozen backbone for concept-prompted segmentation, but applying it directly to open-vocabulary semantic segmentation (OVSS) is inefficient: full-resolution decoding is typically run over the entire dataset vocabulary, whereas each image contains only a small active subset of classes. We introduce ActiveSAM, a training-free, zero-shot inference framework that turns SAM 3 into an...

    arxiv.org/abs/2606.16996 · PDF

  4. 04

    A Multi-Center Benchmark for Abdominal Disease Diagnosis and Report Generation from Non-Contrast CT

    Mariam Elbakry, Aliaa Sayed Sheha, Salma Hassan Tantawy, Aya Yassin, Concetto Spampinato, Karim Lekadir, Xiaomeng...

    cs.CV · cs.LG

    Multiphasic contrast-enhanced CT (CECT) is widely used for abdominal lesion characterization, yet it carries inherent risks of contrast-induced nephropathy, escalates acquisition burden, and heavily contributes to radiologist workload. To address these challenges, we introduce a novel multi-center benchmark for multi-organ abdominal disease diagnosis and automated radiology report generation, which learns to synthesize contrast-enhanced...

    arxiv.org/abs/2606.16991 · PDF

  5. 05

    Semantic Flip: Synthetic OOD Generation for Robust Refusal in Embodied Question Answering and Spatial Localization

    Dongbin Na, Chanwoo Kim, Giyun Choi, Dooyoung Hong

    cs.CV · cs.AI

    Detecting unanswerable user queries remains essential for the reliable deployment of real-world embodied agents. However, modern vision-language models (VLMs) often generate overly confident answers even when the available visual memory cannot support the query. Such overconfidence poses various task-dependent risks. The agent may provide misleading information to the user in Embodied Question Answering and select an arbitrary coordinate and...

    arxiv.org/abs/2606.16898 · PDF

  6. 06

    Federated Medical Image Segmentation under Real-World Label Noise: A Benchmark Suite for Noisy Label Learning Method Selection

    Markus Bujotzek, Dimitrios Bounias, Stefan Denner, Ralf Floca, Maximilian Fischer, Peter Neher, Klaus Maier-Hein

    cs.CV · cs.AI · cs.DC

    While federated learning (FL) enables collaborative medical image segmentation without centralizing sensitive data, real-world deployment is frequently complicated by cross-site label imperfections such as contour disagreement, missing or additional structures, and confused labels. Federated noisy label learning (FNLL) aims to mitigate these effects, yet remains underused in practice as existing evidence is largely based on synthetic noise,...

    arxiv.org/abs/2606.16868 · PDF

  7. 07

    Robust Spoofed Speech Detection via Temporal Pyramid Modeling

    Mahtab Masoudi Nezhad, Nima Karimian

    cs.CV · cs.AI · cs.SD

    Spoofed speech detection is increasingly challenged by realistic synthesis, voice conversion, and replay attacks, with cross-dataset generalization remaining a major limitation. This work we propose a Temporal Pyramid Adapter that utilize parallel temporal convolutions with varying receptive fields to capture multi-scale spoofing cues, ranging from local artifacts to global prosodic irregularities. We also integrated self-supervised XLS-R...

    arxiv.org/abs/2606.16837 · PDF

  8. 08

    Decoupling Semantics from Distortions: Multi-Scale Two-Stream Vision-Language Alignment for AI-Generated Image Quality Assessment

    Zijie Meng

    cs.CV · cs.AI

    Existing vision-language model (VLM)-based AI-generated image quality assessment (AIGIQA) methods suffer from a fundamental semantic-distortion dimensional conflict: monolithic representations optimized for semantic discrimination inherently entangle compositional understanding with low-level perceptual sensitivity, rendering them blind to fine-grained quality degradations. We introduce MST-CLIPIQA, a multi-scale two-stream framework that...

    arxiv.org/abs/2606.16799 · PDF

  9. 09

    Gen-VCoT: Generative Visual Chain-of-Thought Reasoning via Diffusion-Based RGB Intermediate Representations

    Zhiqiang Zhou, Junliang Dai, Xu ling

    cs.CV · cs.AI · cs.LG

    Multimodal large language models (MLLMs) excel at visual reasoning but rely on text-based chain-of-thought (CoT), lacking interpretable visual intermediates. Existing methods use opaque tokens or external tools, missing key properties. We propose Gen-VCoT, a framework using expert vision models to generate RGB images as reasoning intermediates. It has three stages: visual grounding (SAM segmentation), geometric reasoning (Marigold depth...

    arxiv.org/abs/2606.16783 · PDF

  10. 10

    Revealing Artifacts via Noise Amplification: A Novel Perspective for AI-Generated Video Detection

    Renxi Cheng, Jie Gui, Hongsong Wang

    cs.CV · cs.AI

    With the rapid advancement of video generation models, distinguishing between AI-generated and authentic videos has emerged as a challenging endeavor. The majority of existing research endeavors concentrate on the development of detectors for identifying samples generated by generative adversarial networks. Nevertheless, the detection of AI-generated videos, particularly those produced by text-to-video models, still remains an uncharted...

    arxiv.org/abs/2606.16742 · PDF

  11. 11

    DCP-Prune: Ultra-Low Token Pruning with Distribution Consistency Preservation

    Xifeng Xue, Xiaokang Wang, Zirui Li, Ming-Ming Cheng, Guolei Sun

    cs.CV · cs.AI

    Recent vision token pruning methods effectively preserve model performance under moderate token budgets but become unstable under ultra-low token budget. Our analysis shows that as the pruning budget decreases, accuracy degradation is often accompanied by larger feature distribution shifts. Critically, the degree of this distribution shift strongly correlates with performance degradation. To better characterize this phenomenon, we introduce a...

    arxiv.org/abs/2606.16633 · PDF

  12. 12

    Unified Multimodal Model for Brain MRI Imputation and Understanding

    Zhiyun Song, Che Liu, Tian Xia, Avinash Kori, Wenjia Bai

    cs.CV · cs.AI · cs.MM

    Multimodal large language models (MLLMs) hold great potential for medicine, as they inherit knowledge from LLM and allow multiple data modalities to be integrated, analysed and interpreted in natural language. However, the field of medical MLLMs is constrained by non-trivial challenges, notably the scarcity of high-quality training data and the frequent occurrence of missing data in the real-world clinical setting. Here, we propose a novel...

    arxiv.org/abs/2606.16484 · PDF

  13. 13

    Uncertainty Quality of VGGT: An Analysis on the DTU Benchmark Dataset

    Markus Hillemann, Robert Langendörfer, Steven Landgraf, Markus Ulrich

    cs.CV · cs.AI

    Visual Geometry Grounded Transformer (VGGT) has already attracted a great deal of attention in a short period of time, not least due to the Best Paper Award at CVPR-2025. Similar to DUSt3R and MASt3R, VGGT aims to bring about a paradigm shift by replacing established methods like bundle adjustment and feature matching with a simple, unified, feed-forward neural network that predicts camera poses, depth maps, and dense 3D structure directly...

    arxiv.org/abs/2606.16479 · PDF

  14. 14

    What Should a Streaming Video Model Remember?

    Haonan Ge, Yiwei Wang, Hang Wu, Yujun Cai

    cs.CV · cs.AI

    Streaming video understanding models must answer queries at any moment during an ongoing stream, using only what they have observed so far and under fixed memory and computation budgets. Existing methods address this by adding memory banks, retrieval modules, or visual token compression to preserve long-range history. However, strong recent-window baselines show that indiscriminate history injection can dilute current-scene perception,...

    arxiv.org/abs/2606.16353 · PDF

  15. 15

    Differentiable Packing of Irregular 3D Objects with Adaptive Container Estimation

    Palak Gupta, Shanmuganathan Raman

    cs.CV · cs.GR · cs.LG

    Most existing approaches either fix the container in advance or optimize only a single container dimension through an outer search loop, leaving the remaining dimensions as a manual tuning problem. We present a differentiable packing framework that jointly optimizes all 6N object pose parameters and all three container side lengths inside a single gradient-based loop. The formulation combines six physics-inspired, differentiable loss terms...

    arxiv.org/abs/2606.16333 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.