cs.CV · 2026-08-21 · No. 91

Computer Vision and Pattern Recognition, 2026-08-21.

21 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

21 entries
  1. 01

    Prompt-Conditioned Channel Attention for Hierarchical Feature Modulation toward Anatomy-Agnostic Segmentation

    Mosharof Hossain, Md Rabiul Islam, Limon Halder, Erchin Serpedin, Md Kamrul Hasan

    cs.CV · cs.AI

    Anatomically plausible segmentation remains challenging because of low contrast, ambiguous boundaries, and modality-specific artifacts. Interactive segmentation has emerged as a promising strategy to guide feature extraction and improve localization, particularly in structurally ambiguous regions. However, existing methods integrate prompts through late-stage fusion and lack explicit mechanisms for prompt-driven channel-wise modulation across...

    arxiv.org/abs/2608.20229 · PDF

  2. 02

    Feature Evolution and Migration during Vision Transformer Training

    Joonas Järve, Halil Ibrahim Aysel, Tarun Khajuria, Meelis Kull

    cs.CV · cs.LG

    We present a novel view on feature evolution in Vision Transformers (ViTs) by visualizing the training process over two dimensions -- network depth (layer) and training time (epochs). We employ Sparse Autoencoders (SAEs) to extract candidate sparse features from CLS-token representations and compare their activation profiles across epoch--layer pairs. This allows us to study feature-level dynamics that are not directly visible from...

    arxiv.org/abs/2608.20134 · PDF

  3. 03

    Structured Affinity for Unsupervised Visual Class-Incremental Memory in Deep Artificial Immune Networks

    Siphesihle Sithungu

    cs.CV · cs.AI · cs.LG

    Artificial immune networks (AINs) are naturally memory-forming systems, but conventional visual AINs often rely on flattened vector affinity that ignores spatial structure. This paper studies whether structured, gradient-free immune affinity can make Deep AINs viable as replay-free visual class-incremental representation-memory learners. Visual B-cells are formalized as structured templates, including shifted-template affinity,...

    arxiv.org/abs/2608.20104 · PDF

  4. 04

    From Street View Imagery to Street Quality Indicators: Vision Language Inference for the Suburban 15-minute City

    Joan Perez, Giovanni Fusco

    cs.CV · cs.LG

    Streetscape quality has become a central concern in contemporary urban planning, particularly within the framework of the pedestrian-friendly 15-minute city, where walkability and public-space quality are increasingly recognized as key determinants of urban performance. However, assessing streetscape qualities across large suburban and peri-urban territories remains challenging due to the time and resource demands of conventional field...

    arxiv.org/abs/2608.20026 · PDF

  5. 05

    Open-Vocabulary 3D Object Detection with Co-Distillation Discovery and Dual Guidance Robust Training

    Shangbo Yuan, Jie Xu, Xiaofeng Zhu, Na Zhao

    cs.CV · cs.AI

    Recently, open-vocabulary 3D object detection (3D-OVD) has gained increasing attention for its ability to detect unseen objects in 3D scenes. Existing approaches typically adopt a two-stage pipeline that first discovers novel objects using foundation models and then trains a 3D-OVD model based on these discovered objects. Although effective, this pipeline often suffers from inaccurate localization and mismatched classification during the...

    arxiv.org/abs/2608.19973 · PDF

  6. 06

    Flow Matching Meets 3D Curvilinear Structure Segmentation in Medical Imaging

    Sidi Mohamed Sid'El Moctar, Nicolas Vitry, Hélène Bouvrais

    cs.CV · cs.LG · q-bio.QM

    Segmentation of curvilinear anatomical structures in 3D medical images remains challenging due to complex topology, severe class imbalance, weak contrast, and large variations in structure morphology. While deep learning approaches for 3D curvilinear segmentation have been proposed, they are often tailored to specific anatomies or modalities, limiting generalization across clinical settings and leaving room for improvement. Recent generative...

    arxiv.org/abs/2608.19965 · PDF

  7. 07

    Core-KAN: Continuous Vision Kernels with Kolmogorov-Arnold Networks

    Lan Guo, Mengling Li, Haoran Li, Jun Shen, Yuanbo Jiang, Qingguo Zhou, Binbin Yong

    cs.CV · cs.AI

    Conventional convolutional kernels are typically defined on fixed discrete grids, limiting their ability to accommodate heterogeneous local structures. Existing adaptive operators improve flexibility but often couple geometric scale variation with content-dependent filtering, while incurring high computational cost from per-location kernel generation. To decouple geometric scale adaptation from content-dependent filtering while avoiding...

    arxiv.org/abs/2608.19817 · PDF

  8. 08

    Far from the Crowd: Scalable Self-Supervised Learning via Geographic Isolation

    Daniele Rege Cambrin, Francesco Rossi, Mattia Varile

    cs.CV · cs.LG

    Self-supervised pretraining on remote sensing imagery typically treats all samples as equally informative, despite large variability in geographic and visual structure. We propose a curriculum learning strategy for self-supervised Earth observation that ranks samples by geographic isolation, a label-free proxy derived entirely from geolocation metadata already present in geospatial datasets, requiring no image decoding, no model feedback, and...

    arxiv.org/abs/2608.19766 · PDF

  9. 09

    Question-Guided Evidence Acquisition for Multimodal Visual Question Answering

    Alin-Ionut Popa

    cs.CV · cs.AI · cs.LG

    Multimodal LLMs can see a document, but they often can't read it reliably. Small text, tables, visual cues, and topological elements still trip them up under direct visual inference, even when the page is already sitting in the model's context. Most document-VQA systems treat perception as fixed: they encode the page once, ask the question, and answer from whatever the model happened to extract in that single fast pass. We think document VQA...

    arxiv.org/abs/2608.19739 · PDF

  10. 10

    Learning to Beat: Phenotype-Guided Latent Flow with Regional Motion Priors for Biventricular Motion Synthesis

    Xuan Yang, Xiaohan Yuan, Hao Li, Lingyu Chen, Yanan Liu, Qingya Li, Lei Li

    cs.CV · cs.AI

    Full-cycle biventricular geometry is essential for characterizing cardiac function. However, dense and temporally consistent 3D+t biventricular meshes are not routinely available, whereas end-diastolic (ED) anatomy can often be obtained reliably. We therefore investigate full-cycle biventricular motion synthesis from a single ED mesh. This task is challenging because cardiac deformation is spatially heterogeneous and phenotype dependent,...

    arxiv.org/abs/2608.19738 · PDF

  11. 11

    TempJail: Temporal Jailbreak Attack against Large Vision-Language Models via Subtitle Scheduling

    Ling Zhou, Yihao Huang, Jingling Sun, Zhiwen Tian, Yi Zeng, Qihe Liu, Shijie Zhou

    cs.CV · cs.AI · cs.CL

    Large vision-language models (LVLMs) have achieved remarkable progress in video understanding and reasoning. Despite extensive studies on text- and image-based jailbreaks, video jailbreaks against LVLMs remain largely unexplored. Existing video jailbreak methods mainly manipulate textual content embedded in videos, while overlooking how such information is organized over time. Our analysis reveals that jailbreak effectiveness depends not only...

    arxiv.org/abs/2608.19737 · PDF

  12. 12

    Scale-Separated Conditioning for Style-Encoder-Free Diffusion Stylization

    Jingtao Zhang, Haorui Gao, Youqing Liang, Zeming Liu

    cs.CV · cs.AI

    Reference-based diffusion stylization requires separating target geometry from transferable appearance. Existing tuning-based methods often rely on aligned content-style-target triplets or auxiliary visual encoders, which increases data cost and can transfer unintended scene structure from the style reference. We propose SEFS (Style-Encoder-Free Stylization), a style-encoder-free conditioning framework for diffusion transformers. SEFS forms...

    arxiv.org/abs/2608.19719 · PDF

  13. 13

    Robust Cross-Modal Foundation Model Perception for Underwater Robots under Degraded Visual Conditions

    Mohammad Arif Ul Alam

    cs.CV · cs.AI

    Reliable underwater robotic perception remains difficult because optical imagery degrades under turbidity, wavelength-dependent attenuation, low illumination, scattering, and blur. Although sonar provides complementary information that is less affected by optical visibility, prior visual-sonar research has largely focused on feature alignment and nominal detection performance. We investigate cross-modal robustness as visual reliability...

    arxiv.org/abs/2608.19710 · PDF

  14. 14

    RIPE++: Reinforced Keypoint Learning from Positive Pairs Only

    Johannes Künzel, Peter Eisert, Anna Hilsmann

    cs.CV · cs.LG

    Sparse keypoint extraction and matching underpin core tasks in geometric computer vision, including structure-from-motion, visual SLAM, augmented reality, and medical image registration. Learning robust local feature representations, however, typically requires accurate camera poses or depth supervision, which are often unavailable in real-world settings. Reinforcement learning (RL) has recently emerged as a promising alternative, requiring...

    arxiv.org/abs/2608.19693 · PDF

  15. 15

    Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning

    Haoqiang Kang, Yinpeng Chen, Luyang Liu, Jesper Sparre Andersen, Abhijit Ogale, Baochen Sun, Lichan Hong, Ed H. Chi

    cs.CV · cs.LG

    Latent reasoning has advanced multimodal reasoning through a two-stage training paradigm: (1) a helper image is encoded into latent tokens to teach visual chain-of-thought during a supervised fine-tuning (SFT) stage, and (2) these latent tokens are further refined with reward feedback during a reinforcement learning (RL) stage. In this paper, we identify two key limitations of this framework, one in each stage. First, the SFT stage typically...

    arxiv.org/abs/2608.19669 · PDF

  16. 16

    PEA-DPO: Perception-Enhanced Alignment Direct Preference Optimization for MLLMs Alignment

    Jiawei Feng, Jiancan Wu, Xingyu Zhu, Junkang Wu, Xiang Wang, Xiangnan He

    cs.CV · cs.AI · cs.CL · cs.MM

    Direct Preference Optimization (DPO) has emerged as an effective approach for aligning large language models (LLMs) with human preferences. However, its adaptation to multimodal settings remains unexplored. Through representational analysis, we identify a key limitation in multimodal preference optimization, which we term visual insensitivity: models often fail to distinguish between images and those with critical visual context removed. Our...

    arxiv.org/abs/2608.19598 · PDF

  17. 17

    VGI-BENCH: Probing Visual Intelligence in Video Generation Models

    Xuan He, Cong Wei, Yuhao Cheng, Linrui Ma, Yuxuan Zhang, Zuojun Li, Yuhao Wen, Zeyi Liu, Yuren Hao, Songcheng Cai,...

    cs.CV · cs.AI

    Recent studies suggest that video generation models can exhibit certain forms of zero-shot visual reasoning through generated frames. Yet reliable evaluation remains challenging: benchmarks should adopt inputs aligned with the visual priors of current video models, require valid evolving processes rather than only plausible final states, and calibrate task difficulty to remain challenging yet partly feasible. To this end, we introduce...

    arxiv.org/abs/2608.19583 · PDF

  18. 18

    Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models

    Yuanhao Ban, Jiaqi Feng, Hengguang Zhou, Xiaohuan Pei, Justin Cui, Cho-Jui Hsieh

    cs.CV · cs.AI

    Streaming autoregressive diffusion models enable real-time, long-horizon video generation, but their training objectives optimize local frame prediction rather than the geometry and dynamics of a coherent world: long rollouts accumulate geometric drift and degrade into static or unnatural motion. Recent bidirectional approaches address this problem using rewards signals built upon 3D Gaussian-Splatting reconstruction. However, a single rigid...

    arxiv.org/abs/2608.19556 · PDF

  19. 19

    CVSD-Reg: Cross-Modal Visual Semantic Prior Distillation for Robust LiDAR Registration

    Eunsoo Im, Junghun Suh, Gyeonggwan Lee, Seunghwan Hong

    cs.CV · cs.AI · cs.RO

    Learning-based global point cloud registration has achieved remarkable progress, yet its reliance on geometric representations makes existing methods sensitive to variations in point density, scan pattern, viewpoint, and sensor characteristics. We propose CVSD-Reg, a robust global LiDAR registration framework that distills visual semantic priors from a vision foundation model into LiDAR representations. In Stage 1, a Point Transformer V3...

    arxiv.org/abs/2608.19536 · PDF

  20. 20

    HiRA-CAM: Preserving Fine-Grained Spatial Relevance in Gradient-Based Visual Explanations

    Manasi Nerurkar, Ali A. Minai

    cs.CV · cs.AI · cs.LG · cs.NE

    Deep Learning models can include billions of parameters or more, making it difficult to explain their internal transformations and outputs. However, explainability is increasing in importance due to the use of AI in crucial applications. This paper focuses on the interpretability of convolutional neural networks (CNNs). Building on the popular gradient based method LayerCAM for extracting internal features in CNNs, we propose an improved...

    arxiv.org/abs/2608.19407 · PDF

  21. 21

    Does Marginal Coverage Guarantee Class-Conditional Safety for Zero-Shot VLMs Under Shift?

    Jai Kumar Sharma, Amartya Dutta

    cs.CV · cs.AI

    Split-conformal prediction provides marginal coverage under exchangeability and is increasingly used as an abstention layer for zero-shot vision-language models (VLMs). We audit this practice under deployment shift for CLIP, OpenCLIP, and SigLIP across ImageNet and non-ImageNet settings. Marginal coverage can remain relatively high while class-conditional tail coverage collapses: on ImageNet-Sketch, worst-class coverage falls to $\approx 0$...

    arxiv.org/abs/2608.19376 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.