cs.CV · 2026-09-03 · No. 104

Computer Vision and Pattern Recognition, 2026-09-03.

15 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

15 entries
  1. 01

    RVSD: Retrieval Vision Sparse Decoding for Mitigating Visual Hallucinations in Large Vision-Language Models

    Canjie Liu, Jiawen Kang, Jinbo Wen, Zishao Zhong

    cs.CV · cs.AI

    Large vision-language models have achieved remarkable success in vision-language tasks. However, they remain prone to Visual Hallucinations (VHs), undermining their reliability in real-world applications. Existing solutions typically require curated datasets, additional training, or multi-round decoding, resulting in considerable computational overhead. In this paper, we propose \textbf{RVSD} (\underline{R}etrieval \underline{V}ision...

    arxiv.org/abs/2609.02731 · PDF

  2. 02

    Fine-Grained Anomaly Perception in Wild UGC-Enhanced Images: A Comprehensive Dataset and Difference-Fusion Framework

    Yan Zhong, Gefei Chen, Qiufang Ma, Zhen Wang, Zhiwei Fan, Lei Shi, Tingting Jiang

    cs.CV · cs.AI

    Image enhancement and restoration have become standard back-end operations on short-video and social media platforms to boost UGC visual experience. Yet these processes inevitably introduce visual anomalies--especially in faces, texts, and textures--that directly undermine perceptual fidelity and viewer trust. While existing IQA methods perform well on classic distortions, they target holistic quality assessment and fail to capture the...

    arxiv.org/abs/2609.02529 · PDF

  3. 03

    Orthogonal Ensembles and Tested Explanations for Performer-Independent Body-Motion Emotion Recognition

    Naoto Nishida, Yoshio Ishiguro

    cs.CV · cs.HC · cs.LG

    We study body-only, 12-class acted-emotion classification from skeleton motion under leave-performer-out (LPO) evaluation, a hard, underdetermined setting: chance is 8.3%, and a protocol-matched reproduced STGCN++ baseline reaches only 25.73 +/- 4.03% Macro-F1. We show that reliable gains come not from a new architecture but from combining eleven models with orthogonal error modes: under 10-fold LPO cross-validation on the labeled training...

    arxiv.org/abs/2609.02510 · PDF

  4. 04

    Blending Concepts: Benchmarking Visual Metaphor Generation in Text-to-Image Models

    Chuer Chen, Zichen Wang, Yi He, Zhengxi Yu, Nan Cao

    cs.CV · cs.AI

    Text-to-image (T2I) models have achieved remarkable success at faithfully rendering specified objects and attributes, yet their ability to produce visual metaphors, images that convey abstract ideas by combining elements from two distinct domains, remains largely unexamined. To bridge this gap, we introduce VMetaphor-Bench, the first benchmark for evaluating visual metaphor generation in T2I models. It comprises 1,500 visual metaphors curated...

    arxiv.org/abs/2609.02502 · PDF

  5. 05

    ORB-SVM : An Innovative Hybrid Framework for Efficient Brain Tumor Detection from MRI Scans

    Amirhosein Azarpour

    cs.CV · cs.AI

    Brain cancer remains one of the most significant challenges in modern medicine, where the accuracy of early stage diagnosis is a decisive factor in patient survival and treatment efficacy. Although Magnetic Resonance Imaging (MRI) is the established gold standard for visualizing neurological structures, the interpretation of these high dimensional scans is often complicated by subjective variability among practitioners and the inherent noise...

    arxiv.org/abs/2609.02333 · PDF

  6. 06

    VoRTeC: Taming Foundation Flow for One-step Real time Video Compression

    Yichong Xia, Qinhong Wu, Qinhong Wu, Jinpeng Wang, Zeyuan Chen, Haoqian Wang

    cs.CV · cs.AI

    Ultra-low bitrate video compression still faces critical challenges: traditional neural video compression inevitably introduces blurring artifacts, while diffusion-based generative video compression suffers from excessive decoding latency and poor temporal consistency. To address these issues, we propose $\mathtt{VoRTeC}$, a Video Compression framework built upon a foundational flow model (Wan2.1). By compactly encoding latent video...

    arxiv.org/abs/2609.02291 · PDF

  7. 07

    RouteGraph-Mona: Confusion-Aware Routing Fine-Tuning for Mineral Image Classification

    Jierui Li, Zhiyuan Qi, Hao Zhu, Yufan Liu, Jixian Liu, Shaojie Jiang, Jianda Wang, Yaqi Liu, Xiaotong Li, Wei Wang

    cs.CV · cs.AI

    Mineral image classification is important for geological exploration and resource development, but it remains challenging due to substantial intra-class variations in appearance and high inter-class visual similarity. Multi-cognitive Visual Adapter (Mona) is a vision-oriented parameter-efficient adapter that adapts pre-trained visual models by tuning only a few parameters. However, Mona statically aggregates responses from multiple scales,...

    arxiv.org/abs/2609.02282 · PDF

  8. 08

    SAUF-Net: Structure--Appearance Representation Learning with Uncertainty Feedback for Semi-Supervised Medical Image Segmentation

    Qin Lu, Zheyang Jing, Yujie Yang, Jianwang Li, Chen Yi, Shaofeng Jiang

    cs.CV · cs.AI

    Semi-supervised learning has shown great potential for reducing annotation costs in medical image segmentation. However, most existing methods mainly exploit unlabeled data through prediction-level consistency, while the reliability of internal feature representations is often overlooked. In medical images, target-related structural cues are easily entangled with unstable appearance variations, which may lead to unreliable pseudo labels and...

    arxiv.org/abs/2609.02247 · PDF

  9. 09

    InfraPatch: Cross-Task Targeted Grayscale Patch Attacks on Infrared-Adapted Vision-Language Models

    Chengyin Hu, Dingyi Lu, Jiaju Han, Xiang Chen, Weiwen Shi, Jiahuan Long, Yiwei Wei, Jiujiang Guo

    cs.CV · cs.AI

    Infrared vision-language models (IR-VLMs) have emerged as a promising paradigm for multimodal perception under low-visibility conditions, yet their robustness to targeted adversarial attacks remains poorly understood. Existing adversarial patch methods mainly study RGB-based models or a single downstream task and do not characterize whether localized perturbations can induce an intended semantic target in IR-VLMs. We propose InfraPatch, a...

    arxiv.org/abs/2609.02233 · PDF

  10. 10

    Signal or Noise? Auditing Rotation-Induced Saliency Drift in Medical and Aerial Imaging

    Khawaja Murad ul Hassan, Mehran Ebrahimi

    cs.CV · cs.AI

    Post-hoc saliency maps such as Grad-CAM are increasingly used to audit why a deployed vision model made a decision, yet the heatmap drifts when the input is rotated, even when the prediction is unchanged. In domains with no canonical orientation, such as histopathology and aerial imagery, this undermines using saliency as evidence. We ask whether that drift is faithful signal or noise introduced by the CAM operator, and answer it by measuring...

    arxiv.org/abs/2609.02224 · PDF

  11. 11

    Disease Burden over Skin Tone: Decomposing the Dermatology-AI Generalization Gap

    Nirajan Kunwor, Sanjaya Poudel, Quoc-Huy Trinh, Jahidul Arafat, Sunil Kumar Gaire

    cs.CV · cs.AI · cs.LG

    Dermatology artificial intelligence (AI) models are predominantly trained on light-skinned, cancer-focused image collections, yet they are increasingly proposed for deployment in resource-constrained settings where patients differ from training populations along two confounded axes: skin tone and disease distribution. We investigate whether poor generalization is primarily caused by skin-tone underrepresentation or disease-distribution shift....

    arxiv.org/abs/2609.02111 · PDF

  12. 12

    InstEditSeg: Instruction-Driven Image Editing for Polyp and Skin Lesion Segmentation

    Ziquan Liu, Zhewei Zhu, Xuyang Shi

    cs.CV · cs.AI

    Accurate segmentation of polyps and skin lesions is pivotal for clinical diagnosis, yet existing methods struggle with low contrast, ambiguous boundaries, and cross-domain distribution discrepancies. Discriminative networks and most diffusion-based segmentation approaches predict standalone binary masks, leaving the visual priors of large-scale pretrained generative models largely unexploited. We propose InstEditSeg, a unified generative...

    arxiv.org/abs/2609.02004 · PDF

  13. 13

    InsightSeg: Reusing Correction Insights for Guideline-Consistent Segmentation

    Vanshika Vats, Ashwani Rathee, James Davis

    cs.CV · cs.AI

    Guideline-consistent semantic segmentation requires more than category recognition, as real-world labeling policies demand fine-grained, task-specific decisions. Recent multi-agent refinement systems improve compliance with such textual guidelines by detecting and correcting errors. However, they are stateless: feedback from the critiquing agent is discarded, causing the same guideline-specific mistakes to be repeatedly rediscovered and...

    arxiv.org/abs/2609.02002 · PDF

  14. 14

    Linear Fusion MultiDiffusion for Fast Training-Free Spherical Panorama Generation

    Akio Hayakawa, Yusuke Mukuta, Tatsuya Harada

    cs.CV · cs.LG

    We propose LF-MultiDiffusion, a training-free panorama generation method that extends MultiDiffusion to support linear projections between target and reference image spaces. Our key idea is to reformulate latent aggregation as a regularized least-squares problem and solve it efficiently with a Krylov-based iterative solver inside the denoising loop. This formulation enables denser and more natural mappings than prior training-free methods,...

    arxiv.org/abs/2609.01997 · PDF

  15. 15

    Morphology signal in whole slide image foundation models can automatically triage slides

    Ayushi Sinha, Shashank Yadav, Benjamin Holmes, Pravat Das, Aaron W. Bogan, James S. Lewis, Santiago Romero-Brufau,...

    cs.CV · cs.LG

    Patient exams in the cancer diagnosis and staging process typically generate several whole slide images (WSIs). One of the initial steps in training models on WSI data is identifying one or a few slides containing tumor or other diagnostic biomarkers necessary for downstream prediction tasks such as estimating recurrence risk or progression-free survival. This step requires tedious manual curation by experienced pathologists. Many published...

    arxiv.org/abs/2609.01987 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.