cs.CV · 2026-07-23 · No. 62

Computer Vision and Pattern Recognition, 2026-07-23.

14 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

14 entries
  1. 01

    Persian Pixel: A large-scale synthetic OCR dataset for Persian language

    Pouria Mahdi, Haq Nawaz Malik

    cs.CV · cs.AI

    Optical Character Recognition (OCR) for Persian remains substantially less mature than for Latin-script languages despite Persian being spoken by more than 110 million people across multiple countries. This gap arises from two fundamental challenges: the intrinsic complexity of the Perso-Arabic writing system and the limited availability of large-scale, high-quality annotated datasets. Persian script exhibits obligatory cursive connectivity,...

    arxiv.org/abs/2607.20385 · PDF

  2. 02

    Toward Reliable RGB-D Semantic Segmentation: Handling Missing Modalities via Condition Dropout

    Xuchen Zhu, Yajuan Wei, Shuang Hao, Jiwei Jiang, Guanxiang Mao, Fang Ren

    cs.CV · cs.AI

    RGB-D semantic segmentation has achieved remarkable progress, yet most models assume that RGB and depth are always available. In practice, failures or occlusions of surveillance sensors often remove one modality. Although RGB or depth alone can contain sufficient cues, models trained only on full-modality inputs fail to exploit the remaining modality once one is missing, causing severe degradation. We tackle this issue with a simple...

    arxiv.org/abs/2607.20326 · PDF

  3. 03

    Self-supervision drives representational convergence in medical foundation models more than clinical supervision

    Soroosh Tayebi Arasteh, Sebastian Ziegelmayer, Mahshad Lotfinia, Lisa Adams, Sven Nebelung, Jakob Nikolas Kather,...

    cs.CV · cs.AI · cs.CL · cs.LG

    Medical image encoders from different groups are increasingly treated as interchangeable, on the assumption that scale and clinical supervision concentrate their representations onto a shared structure. Whether this convergence is real, what produces it, and whether it is clinically usable are untested, and the similarity measures behind such claims are fragile. We present a controlled dissection across 18 image and 7 text encoders, all...

    arxiv.org/abs/2607.20274 · PDF

  4. 04

    StreamHOI: Interaction-aware Temporal Memory Adaptation for Streaming HOI Video Generation

    Zejing Rao, Haoxian Zhang, Xiaoqiang Liu, Yiping Meng, Guoxin Zhang, Pengfei Wan, Fan Tang, Tong-Yee Lee

    cs.CV · cs.AI

    Existing human--object interaction (HOI) video generation methods are largely limited to offline short-video generation with complex driving conditions, making them unsuitable for real-time interactive applications. We present \emph{StreamHOI}, a low-latency streaming framework for long-duration HOI video generation. Instead of converting heavily conditioned HOI pipelines into streaming systems, we study how an image-to-video streaming...

    arxiv.org/abs/2607.20174 · PDF

  5. 05

    HeadCast: Casting Attention Heads for Efficient Autoregressive Video Generation

    Jinliang Shen, Lianghao Su, Zheming Li, Kang He, ZiLiang Lai, Yanbing Jiang, Chengru Song

    cs.CV · cs.LG

    Autoregressive (AR) video diffusion models have become a promising paradigm for long and streaming video synthesis, but the continuously growing Key-Value (KV) cache makes attention the dominant inference cost, especially at high resolution where each frame contributes many tokens. Existing remedies either evict the cache with coarse heuristics that cause inter-frame flickering, or require model re-training. We propose HeadCast, a...

    arxiv.org/abs/2607.20125 · PDF

  6. 06

    ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models

    Karan Goyal, Afreen Hossain, Debojyoti Das, Vishal Bhutani

    cs.CV · cs.AI · cs.CL

    Contextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, independently of whether that context is relevant, true, or even meaningful. Recently, it has been identified and given a mechanistic account in unimodal language models. Whether and how it manifests in vision-language models (VLMs) is, by contrast, largely unexamined, and the field lacks a purpose-built instrument with which to...

    arxiv.org/abs/2607.20092 · PDF

  7. 07

    A Systematic Benchmark of Intensity Normalisation Methods for 3D Knee MRI Segmentation and Cross-Domain Generalisability

    Oliver Mills, Philip Conaghan, Samuel Relton

    cs.CV · cs.AI

    Robust out-of-the-box performance is essential for the clinical deployment of deep learning models in medical imaging. An important but underexplored factor affecting model generalisability is intensity normalisation, particularly for magnetic resonance imaging (MRI), where image intensities vary across scanners and protocols. In this study, we systematically compared seven normalisation methods and their impact on the performance of a 3D...

    arxiv.org/abs/2607.20028 · PDF

  8. 08

    G-MAD: A Game-Based Data Generation Framework for Multi-View RGB-T Aerial Object Detection

    Yechan Kim, JongHyun Park, Dongho Yoon, Namhoon Jung, Moongu Jeon

    cs.CV · cs.AI

    This work introduces G-MAD, an open-source framework that uses Arma3 to generate synchronized multi-view RGB-T data for aerial object detection. G-MAD addresses key limitations of real-world aerial dataset construction, including limited viewpoint control, imperfect RGB-T alignment and high annotation cost. The framework supports structured scenario specification, controllable multi-view camera placement, simultaneous visible/thermal capture,...

    arxiv.org/abs/2607.19942 · PDF

  9. 09

    OSVE: One Step Video Editing with One Step Diffusion Models

    Habin Lim, Gyeong-Moon Park

    cs.CV · cs.AI

    Text-guided video editing with diffusion models is impractically slow, hindered by costly multi-step sampling and inversion. We present OSVE, the first framework to successfully adapt one-step Text-to-Image (T2I) models for high-quality video editing, addressing the core challenges of inversion, editability, and temporal consistency. To bypass slow iterative inversion, we train a learnable encoder that predicts the initial noise for each...

    arxiv.org/abs/2607.19895 · PDF

  10. 10

    Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos

    Penglei Sun, Yehua Huang, Zhuoli Tao, Xiang Li, Runwei Guan, Yaoxian Song, Kaiyong Zhao, Henghui Ding, Bo Han, Yang...

    cs.CV · cs.AI

    Language-guided aerial perception aims to understand user-specified tiny targets in complex unmanned aerial vehicle (UAV) scenes. In real UAV deployment, the UAV must respond while it flies, so such perception runs in an online streaming manner, where frames arrive sequentially and the model responds to each one without access to future frames. However, applying current Multimodal Large Language Models (MLLMs) to this setting raises two...

    arxiv.org/abs/2607.19857 · PDF

  11. 11

    Physics-Aware Complex-Valued State Space Model with Scattering-Prior Feature Modulation for PolSAR Image Classification

    Fangyan Zhang, Fan Zhang, Shiqi Zhou, Jun Ni, Carlos López-Martínez, Qiang Yin

    cs.CV · cs.AI

    Polarimetric synthetic aperture radar (PolSAR) image classification is a representative task for physics-aware GeoAI, where land-cover semantics are closely coupled with electromagnetic scattering mechanisms. Many existing complex-valued networks can preserve amplitude-phase information, but they are often limited in long-range spatial dependency modeling and usually incorporate polarimetric priors only as input-level or shallow auxiliary...

    arxiv.org/abs/2607.19787 · PDF

  12. 12

    PhenSPINE: A Standardized Benchmark for Spine Pathology Diagnosis

    Duong Ngoc Vu, Hai Son Nguyen, Trong-Nghia Nguyen, Bien Tran Van, Trang Mai Xuan, Huan Vu, Thien Van Luong

    cs.CV · cs.AI

    The accurate diagnosis of spinal pathologies depends heavily on radiological interpretation, yet automated systems are hindered by the lack of diverse, high-quality benchmarks. In this study, we present PhenSPINE, a Magnetic Resonance Imaging dataset comprising 16,813 images from 250 patients, curated to facilitate advanced deep learning research. We propose a robust diagnostic benchmark that integrates state-of-theart convolutional backbones...

    arxiv.org/abs/2607.19696 · PDF

  13. 13

    D3VL: Understanding Driving Scenes from 3D Time Series Data and Video with Language Models

    Heesang Han, A. Lynn Abbott, Abhijit Sarkar

    cs.CV · cs.AI

    Recent advances in Multimodal Large Language Models (MLLMs) have triggered the development of end-to-end MLLMs for autonomous driving. However, the main emphasis to date has been for MLLMs using 2D images and videos. In contrast, this paper considers MLLM effectiveness using 3D sensors, particularly LiDAR and stereo cameras. LiDAR presents unique challenges to integration within an MLLM, largely because of data sparsity and lack of a grid...

    arxiv.org/abs/2607.19528 · PDF

  14. 14

    Appearance Pointers -- Multimodal Region Control of Diffusion Transformers

    Rahul Sajnani, Yulia Gryaditskaya, Radomír Měch, Srinath Sridhar, Matheus Gadelha

    cs.CV · cs.AI · cs.GR

    Controllable image generation remains challenging for creative professionals, who often require precise regional control over materials, object identities, and spatial arrangements that cannot be reliably achieved through text prompting alone. Diffusion Transformers (DiTs) can natively ingest heterogeneous tokens stemming from texts and images, but they lack mechanisms for determining where and how these tokens should influence the output. We...

    arxiv.org/abs/2607.19344 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.