cs.CV · 2026-09-16 · No. 115

Computer Vision and Pattern Recognition, 2026-09-16.

22 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

22 entries
  1. 01

    PhysStream: Streaming Physics-Grounded Video Generation with Structured Scene Memory and Fine-Grained Motion Control

    Chuhao Chen, Peter Wonka, Chaoyang Wang, Chen Wang, Qiao Feng, Sergey Tulyakov, Lingjie Liu

    cs.CV · cs.AI · cs.GR

    Interactive control for video generation is moving from coarse prompts toward fine-grained, physically meaningful manipulation of dynamic scenes. Yet existing controllable methods either require the full control schedule before generation starts, or use pixel-space signals that dictate object positions rather than physical dynamics. To address these limitations, we propose PhysStream, an autoregressive model for physics-grounded...

    arxiv.org/abs/2609.17521 · PDF

  2. 02

    Det-LIME: Detector-Aware, Multi-Instance Local Interpretable Model-Agnostic Explanations for Automated Marine Mammal Detection

    Jiayi Zhou, David W. Johnston, Brinnae Bent

    cs.CV · cs.AI

    Despite the rapid uptake of black-box object detectors in marine mammal research and monitoring, explainability techniques are rarely integrated into conservation workflows. Furthermore, most classification-oriented explainability tools are ill-suited to detection tasks involving imagery of social organisms or those with colonial life histories, as they ignore multiple detections within a scene and produce single-instance outputs that blur...

    arxiv.org/abs/2609.17479 · PDF

  3. 03

    Tables Decoded: DELTA for Structure, TARQA for Understanding

    Jahanvi Rajput, Dhruv Kudale, Saikiran Kasturi, Utkarsh Verma, Ganesh Ramakrishnan

    cs.CV · cs.LG

    Table understanding is a core task in document intelligence, encompassing two key subtasks: table reconstruction and table visual question answering (TabVQA). While recent approaches predominantly rely on vision- language models (VLMs) operating on table images, we propose a more scalable and effective alternative based on structured textual representations. These representations are easier to process, align more naturally with LLMs, and...

    arxiv.org/abs/2609.17458 · PDF

  4. 04

    Tracking the Unseen: An Occlusion-Robust Framework for Target Tracking Under Full and Long-Term Occlusion

    Mais Mohammed, Sharifa Mohammed, Hanan Awadh, Haneen Bamaas, Raghad Bawazeer, Elham Alghamdi

    cs.CV · cs.AI

    Real-time multi-object tracking systems remain highly vulnerable to full and long-term occlusion, where targets temporarily or completely disappear from the camera's field of view. Conventional trackers may terminate trajectories prematurely, resulting in identity loss and reduced situational awareness in applications such as defense and surveillance. This work proposes an occlusion-robust target tracking framework that maintains target...

    arxiv.org/abs/2609.17427 · PDF

  5. 05

    FROD: Feature Matching Residual Denoising Oracle Bone Decipher

    Yanbin Hou, Biao Xiong, Guojun Xu, Jianwen Xiang, Cheng Tan, Yanchao Yang, Junwei Zhou

    cs.CV · cs.AI

    Oracle bone script (OBS), one of the earliest Chinese writing systems, plays an important role in the study of Chinese etymology. Traditional decipherment relies heavily on domain experts who analyze characters through semantic context and structural evolution. To assist this labor-intensive process, we formulate OBS decipherment assistance as a cross-era image translation task and propose FROD (Feature Matching Residual Denoising Oracle Bone...

    arxiv.org/abs/2609.17227 · PDF

  6. 06

    Multimodal Cultural Heritage Architectural Style Classification for Residential Buildings in the UAE Based on CLIP Embeddings and SVM

    Ahmed Ammar Kubba, Manar Abu Talib, Iman Ibrahim, Qassim Nasir

    cs.CV · cs.AI

    The analysis and classification of cultural heritage architectural styles remain challenging due to the complexity of visual images of buildings, which are highly relied on in traditional CNN-based classification approaches in comparison to textual descriptions, and the relative lack of non-western region-specific datasets. This paper addresses this gap by proposing a multimodal machine learning framework to analyze and classify Emirati...

    arxiv.org/abs/2609.17181 · PDF

  7. 07

    MUMINS: Metadata-conditioned Uncertainty-aware Medical Image Next-state Synthesis

    Anna Oliveras, Roger Marí, Rafael Redondo, Oriol Guardià, Cynthia Ifeyinwa Ugwu, Ana Tost, Bhalaji Nagarajan,...

    cs.CV · cs.AI

    Forecasting anatomical changes such as tumor growth and neurodegeneration is a challenging generative vision task. Morphological evolution is subtle relative to static anatomy, highly patient-specific, and inherently stochastic. Existing methods struggle with several issues: deterministic networks ignore biological stochasticity, while standard diffusion models require computationally prohibitive multi-pass sampling to quantify uncertainty....

    arxiv.org/abs/2609.17169 · PDF

  8. 08

    ResLRP: The Role of Residual Cancellation in Attribution Instability in Vision Transformers

    Jim Berend, Reduan Achtibat, Daniel Schäffer, Alexander Binder, Wojciech Samek, Sebastian Lapuschkin, Maximilian Dreyer

    cs.CV · cs.AI · cs.LG

    Vision Transformers (ViTs) are central to most modern vision models, yet obtaining input attributions that are fine-grained, faithful, and stable remains challenging. Layer-wise Relevance Propagation (LRP) has been adapted to transformer attention, but in ViTs it often produces noisy, unfaithful explanations. We show that the missing ingredient is the treatment of residual connections: cancellation effects in residual pathways lead to...

    arxiv.org/abs/2609.17152 · PDF

  9. 09

    From Foundation Embeddings to Cropland Maps: Label Efficiency, Temporal Transferability and Independent Human Validation

    Mohammad Ammar Mughees, Giovanni Montefoschi, Zhongxin Chen, Maria Antonia Brovelli

    cs.CV · cs.LG · eess.IV

    Geospatial foundation models provide reusable representations of satellite imagery that support downstream mapping with limited task-specific modelling. We evaluate whether annual AlphaEarth embeddings support binary cultivated-versus-non-cultivated mapping in Maine, USA, using 192 spatially separated patches and labels derived from the USDA Cropland Data Layer (CDL). Without fine-tuning the foundation model, a lightweight classifier reaches...

    arxiv.org/abs/2609.17138 · PDF

  10. 10

    Beyond In-Distribution Metrics: A Systematic Out-of-Distribution Evaluation of Congenital Heart Disease Segmentation

    Aniketh Vijesh, Shrisharanyan Vasu, Abhijit Ramesh, Clare Pomeroy-Ward, Harikrishnan Anil Maya, Sarin Xavier, Mahesh...

    cs.CV · cs.AI

    Congenital heart disease (CHD) diagnosis and surgical planning often require patient-specific 3D anatomical models, but manual segmentation is labor-intensive, particularly in complex anatomies. Although deep-learning methods can automate this process, they are typically evaluated in-distribution, despite clinically relevant shifts in scanner, protocol, institution, population, and imaging modality. We present, to our knowledge, the first...

    arxiv.org/abs/2609.17068 · PDF

  11. 11

    MedPCFM-TED: One-Step Point Cloud Flow Matching for Implant Generation via Teacher-Guided Endpoint Distillation

    Kamil Kwarciak, Marek Wodzinski

    cs.CV · cs.LG

    Cranial implant generation is an important task in medical imaging. Recent point cloud based generative methods, particularly flow matching, offer strong reconstruction quality and efficient sampling, but still require multiple neural function evaluations during inference. This limits rapid generation of multiple plausible implant candidates. We propose Teacher-guided Endpoint Distillation (TED), a simple one-step distillation framework for...

    arxiv.org/abs/2609.16934 · PDF

  12. 12

    VOR-Bench: A Human Perception-Driven Benchmark for Video Object Removal

    Haonan Huang, Tianrui Qiu, Xianghao Zang, Yinan Du, Zhixiang He, Chi Zhang, Hao Sun, Zhongjiang He, Tianwei Cao,...

    cs.CV · cs.AI

    Despite its crucial role in video object removal (VOR), existing evaluation paradigms face two critical limitations: questionable references and a misalignment between tradi- tional metrics and human preference. To address these challenges, we introduce VOR- Bench, which advances VOR evaluation through three integrated components. First, we present the VOR Dataset (VORD), the first benchmark dataset providing both paired edited videos and...

    arxiv.org/abs/2609.16878 · PDF

  13. 13

    NeuroTS-Net: Multi-Class Semantic Segmentation of Pediatric Brain Tumors in Multi-Modal MRI

    Darius Peteleaza, Razvan-Gabriel Dumitru, Bogdan Neamtu, Arpad Gellert, Mariana Sandu, Claudiu Matei

    cs.CV · cs.LG

    Pediatric brain tumors are a leading cause of cancer-related mortality in children, and their small, rare, and often low-contrast subregions make accurate manual delineation challenging. Reliable automated segmentation is therefore needed to support diagnosis, treatment planning, and response assessment. Accordingly, we introduce NeuroTS-Net, a three-dimensional encoder-decoder convolutional neural network architecture for multi-class...

    arxiv.org/abs/2609.16873 · PDF

  14. 14

    Measuring Annotation Efficiency for Handwritten Devanagari Recognition: Sample-Complexity Curves for Four Pretraining Regimes

    Manglesh Kumar Pandey, Sumit Kumar Banshal

    cs.CV · cs.LG

    To train handwritten text recognition systems we need word images and their corresponding transcriptions, and these transcriptions are produced manually. For a script that can be read by only a small number of specialists, this manual transcription is a limitation, because the trained models are supposed to save the time of those same specialists. A relevant question therefore arises: how many transcriptions are needed before a recogniser...

    arxiv.org/abs/2609.16859 · PDF

  15. 15

    RegRet: Enhancing Region-Level Retrieval in Large Multimodal Models

    Xun Liang, Honghui Yang, Weihang Pan, Ruisi Zhao, Boyuan Pan, Yao Hu, Wenxiao Wang, Binbin Lin, Deng Cai

    cs.CV · cs.AI · cs.IR

    Region-level retrieval aims to align user-specified image regions with relevant regions or textual descriptions, playing a crucial role in realworld applications such as e-commerce product search and RAG. Although recent Large Multimodal Models (LMMs) have made significant strides in multimodal retrieval, they primarily focus on global-level tasks and struggle to capture effective region-level representations. To bridge this gap, we present...

    arxiv.org/abs/2609.16847 · PDF

  16. 16

    StackTok: Accelerating VLMs Inference with Budget-Adaptive Visual Token Selection

    Zhenbin Wang, Lei Zhang, Lituan Wang, Wei Huang, Yan Wang, Zhenwei Zhang

    cs.CV · cs.AI

    Increasing image resolution produces ever-longer visual-token sequences in vision-language models (VLMs), substantially raising their inference cost. To reduce this overhead without retraining, existing methods select compact token subsets that prioritize query relevance, visual coverage, or a fixed trade-off between them. The appropriate balance, however, varies across queries and token budgets: localized questions favor relevance, whereas...

    arxiv.org/abs/2609.16841 · PDF

  17. 17

    What Breaks Local Watermarks? A Robustness Benchmark for Local Invisible Image Watermarking

    Kai Yao, Bence Szilágyi, Sebestyén Kamp, Máté Poór, Máté Szilveszter, Matyas K. Zsoldos, Marc Juarez

    cs.CV · cs.AI · cs.CR

    Local image watermarking embeds an invisible signal into selected image regions rather than spreading it across the entire image, enabling payload recovery from specific objects or regions without perceptibly altering the image. Existing studies evaluate the robustness of payload recovery and localization under image transformations, but they often focus on their own proposed method, resulting in narrow evaluations with inconsistent choices...

    arxiv.org/abs/2609.16832 · PDF

  18. 18

    Can Knowledge Transfer Parameters Be Learned? LePoKet for Efficient Robotic Vision

    Yanick C. Tchenko, Felix Mohr, Hicham Hadj-Abdelkader, Hedi Tabia

    cs.CV · cs.LG

    Efficient perception is central to robotic systems operating under constrained computation, memory, and latency budgets. Knowledge transfer from larger pretrained models offers a practical route to stronger compact perception networks, but existing approaches commonly rely on fixed distillation objectives or manually designed interaction mechanisms. Building on Hereditary Knowledge Transfer (HKT), we propose LePoKet (Learnable Parameter...

    arxiv.org/abs/2609.16637 · PDF

  19. 19

    A Vision-Language Foundation Model for Precise and Comprehensive Brain Tumor Diagnosis from Preoperative Multimodal Data

    Yinong Wang, Jianwen Chen, Zhou Chen, Shuwen Kuang, Haoning Jiang, Yanzhao Shi, Huichun Yuan, Yan-ran, Wang, Bing...

    cs.CV · cs.AI · cs.DB

    Background Non-invasive presurgical diagnosis of brain tumor types from Magnetic Resonance Imaging (MRI) is essential but challenging due to overlapping imaging features across tumor types, inter-observer variability, and the extensive training required for expertise. We aimed to develop an MRI-based Artificial Intelligence (AI) model for automatic and reliable brain tumor classification with diagnostic uncertainty quantification and...

    arxiv.org/abs/2609.16597 · PDF

  20. 20

    Efficient Text-to-Image Generation: An Adaptive Step Schedule Controller for Diffusion Models

    Kuluhan Binici, Cihan Acar, Shivam Aggarwal, Siying Liu, Tulika Mitra

    cs.CV · cs.AI

    Text-to-image diffusion models often use a fixed number of denoising steps, balancing time costs and image quality. However, the optimal number of steps depends on the complexity of the input text prompt. We propose an adaptive diffusion controller that dynamically adjusts the number of steps to generate high-quality images efficiently, without additional model training. By leveraging a mixture of step schedules with varying step sizes and...

    arxiv.org/abs/2609.16572 · PDF

  21. 21

    Vision And Text Transformer For Predicting Answerability On Visual Question Answering

    Tung Le, Huy Tien Nguyen, Le Minh Nguyen

    cs.CV · cs.AI

    Answerability on Visual Question Answering is a novel and attractive task to predict answerable scores between images and questions in multi-modal data. Existing works often utilize a binary mapping from visual question answering systems into Answerability. It does not reflect the essence of this problem. Together with our consideration of Answerability in a regression task, we propose VT-Transformer, which exploits visual and textual...

    arxiv.org/abs/2609.16565 · PDF

  22. 22

    A multimodal large language model for evidence-based autism spectrum disorder screening

    Jun Chen, Qi Zhao, Yunliang Jiang, Shuqin Cao, Yunqiang Lin, Chenglong Jia, Qiang Guo, Guang Dai, Xiongtao Zhang,...

    cs.CV · cs.HC · cs.LG

    The clinical management of autism spectrum disorder (ASD) faces a bottleneck in early screening, mainly because trained specialists are scarce and conventional assessment tools are subjective. Here, we introduce ASDchat, a multimodal large language model designed for evidence-based ASD screening, which takes video, audio, and dialogue as input. ASDchat adopts a dual-branch architecture, where the decision branch generates screening...

    arxiv.org/abs/2609.16464 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.