cs.CV · 2026-09-30 · No. 129

Computer Vision and Pattern Recognition, 2026-09-30.

24 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

24 entries
  1. 01

    Cropland PAtteRNS: Parallel Dimensional Attention Networks and Attention to Dataset Disparity for Crop Segmentation in Satellite Imagery Time Series Data

    Joseph Metcalfe, Sara Sharifzadeh, Fabio Caraffini

    cs.CV · cs.LG

    The landscape of satellite imagery time series datasets and boundary-pushing architectures for cropland segmentation has never been richer. However, in this gold rush, important truths are being missed on both fronts, as a drive for the most novel concepts or the largest datasets pushes finer details to the side. In this paper, we present our hybrid transformer-convolutional model, Cropland Parallel Attention and Refinement Network for...

    arxiv.org/abs/2609.38165 · PDF

  2. 02

    Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies

    Hui Ren, Lei Fan, Henry Pao, Han Guo, Zeeshan Zia, Ying Chen, Alexander Schwing, Gang Hua

    cs.CV · cs.AI · cs.CL · cs.IR · cs.LG

    Answering questions about long videos often requires connecting events involving the same objects across hours or days. Chronological descriptions and text-derived entities can leave physical identity unresolved: different objects may share a description, while observations of the same object remain disconnected across events. Retrieving relevant events therefore does not necessarily recover the "biography" of the particular entity a question...

    arxiv.org/abs/2609.38155 · PDF

  3. 03

    Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE

    Yu Xu, Yuxin Zhang, Xiao Yang, Haotian Yang, Yizhi Wang, Xinwei Huang, Minxuan Lin, Angtian Wang, Chongyang Ma, Fan Tang

    cs.CV · cs.AI

    Mixture-of-Experts (MoE), popularized by large language models, is a promising paradigm for scaling visual generative models. However, conventional token-wise MoE routes tokens independently within a homogeneous expert pool and regularizes expert usage toward uniformity, making it poorly matched to video data that is spatiotemporally redundant and semantically long-tailed. We show that existing visual MoEs fall into a uniformity trap:...

    arxiv.org/abs/2609.38140 · PDF

  4. 04

    From Unity Simulation to Diffusion-Based Augmentation: Quantifying Dataset Balance for Robust Object Detection

    Mohamed Benkedadra, Aissa Saoudi, Maxime Gloesener, Sidi Ahmed Mahmoudi, Matei Mancas

    cs.CV · cs.AI

    Modern computer vision models achieve high accuracy when trained on large-scale annotated datasets. In critical domains such as construction safety monitoring, data collection is costly, hazardous, and ethically constrained. This paper presents a systematic study comparing two complementary data generation paradigms, (1) Unity Simulation-based rendering and (2) Controllable Diffusion-based generation (CIA), for object detection under real...

    arxiv.org/abs/2609.38010 · PDF

  5. 05

    Does Local Video Understanding Transfer Across Encounters? The EgoGears Benchmark

    Yuedong Tan, Lei Qi, Yu Liu, Di Wen, Ruiping Liu, Xiaoye Wang, Yufan Chen, Junwei Zheng, Chengzhi Wu, Chen Zhang,...

    cs.CV · cs.AI · cs.RO

    Embodied systems must make knowledge acquired during one encounter usable in another despite changes in viewpoint, motion, and illumination. Yet aggregate cross-video accuracy conflates failures of local perception with failures to preserve observation identity, establish correspondence, and compose evidence, obscuring whether local video understanding actually transfers. We introduce EgoGears, a complementary single- and multi-video...

    arxiv.org/abs/2609.37938 · PDF

  6. 06

    Rollout-Marginal Distillation for Long-Horizon Autoregressive Video Generation

    Chenjian Gao, Zhihao Hu, Jianqi Ma, Jun Zhang, Weidong Zhang, Tianfan Xue

    cs.CV · cs.AI

    Autoregressive (AR) video diffusion enables low-latency, streamable video generation, but prediction errors often accumulate over long rollouts. Training the generator on its own rollouts exposes it to these imperfect histories. However, existing video-level distribution matching distillation (DMD) scores the whole rollout jointly. Because a chunk is evaluated together with its past and future, its correction can favor matching artifacts in...

    arxiv.org/abs/2609.37925 · PDF

  7. 07

    SYNCR: Diagnosing and Learning Cross-Video Reasoning from Simulation

    Sara Ghazanfari, Siddharth Garg, Prashanth Krishnamurthy, Farshad Khorrami

    cs.CV · cs.LG

    Reasoning across videos requires aligning events, matching identities, comparing motion, and integrating partial observations. Evaluating these capabilities and testing how to improve them requires both reliable labels and targeted supervision. We introduce SYNCR, a simulator-grounded framework that connects these two needs through shared task generators. Built on Habitat, Kubric, and CLEVRER, SYNCR derives answers from environment state and...

    arxiv.org/abs/2609.37918 · PDF

  8. 08

    ReCAP: Retrieval-Guided Capability Reuse for Multimodal Continual Instruction Tuning

    Tao Hu, Zhinuo Zhou, Xialiang Tong, De-Chuan Zhan, Da-Wei Zhou

    cs.CV · cs.LG

    Multimodal continual instruction tuning (MCIT) aims to enable multimodal large language models to acquire new capabilities from sequential tasks while preserving previously learned knowledge. Existing methods primarily mitigate catastrophic forgetting by constraining parameter updates or separating task-specific adaptations. However, continual adaptation can also benefit from external knowledge that provides domain-specific information and...

    arxiv.org/abs/2609.37889 · PDF

  9. 09

    Visual Branch is What You Need for CLIP-based Class-Incremental Learning

    Tao Hu, Zhen-Hao Xie, Jingcai Guo, De-Chuan Zhan, Da-Wei zhou

    cs.CV · cs.LG

    Class-Incremental Learning (CIL) requires models to recognize new classes over time without forgetting previously learned ones. With the rise of vision-language pre-training, CLIP has become a strong foundation for CIL. A common design in CLIP-based CIL is to construct textual classifier weights by encoding class-name templates with the CLIP text encoder, and then classify visual features by image-text cosine similarity. This design is...

    arxiv.org/abs/2609.37888 · PDF

  10. 10

    HandAnthro: Automated Hand Anthropometry from a Single Image

    Fan Zhou, Shuairan Chen, Mengying Zhang, Yulin Wu, Sadegh Jafari, Sixing Yu, Rui Li, Ali Jannesari, Guowen Song

    cs.CV · cs.AI

    Hand anthropometry supports protective-glove design, but existing measurement methods often require trained operators, specialized hardware, or manual landmarking. We present HandAnthro, which estimates 44 projected hand dimensions from a smartphone photograph of a palm-up hand on US letter-size paper. The pipeline reconstructs wrist-occluded paper boundaries for rectification, whitens non-hand pixels, and refines 41 anthropometry-specific...

    arxiv.org/abs/2609.37855 · PDF

  11. 11

    FlowMap-OPD: Rollout--Kernel Separation for On-Policy Distillation of Few-Step Flow-Map Generators

    Zhiqi Li, Bo Zhu

    cs.CV · cs.LG

    Few-step flow-map generators, including MeanFlow and consistency models, enable efficient sampling through long-range transport, yet their on-policy distillation remains underexplored. We introduce FlowMap-OPD, an on-policy distillation framework that separates student-state acquisition from teacher--student distribution comparison. A formulation based on state marginals establishes this separation, while flow--velocity consistency connects...

    arxiv.org/abs/2609.37851 · PDF

  12. 12

    Evaluation Choices Shape Biomedical ML Claims: A Pediatric Pneumonia Benchmark Case Study

    Bhanu Prakash Vangala, Sowmya Guda, Latha Peddi, Navya Vangala

    cs.CV · cs.LG

    Biomedical machine learning papers often compress model performance into one headline number. That number can look like a property of the model even when it depends strongly on how the benchmark was evaluated. We study this problem on the widely used Kermany pediatric chest radiograph dataset using nine image classifiers and a controlled evaluation protocol. Under the same protocol, the eight pretrained backbones differ by only 0.026 AUROC....

    arxiv.org/abs/2609.37848 · PDF

  13. 13

    Pixel-Level Transformers in Remote Sensing: A Canopy Height Case Study

    Sven Ligensa, Jan Pauls, Karsten Schrödter, Ibrahim Fayad, Fabian Gieseke

    cs.CV · cs.AI

    Predicting canopy height from medium-resolution satellite imagery is a common and scalable approach for assessing the condition of the world's forests, which play a crucial role in climate change mitigation. While Transformer-based architectures have shown strong performance in many domains, their straightforward application to dense (i.e., pixel-level) regression tasks often yields suboptimal results. In particular, the patch size has a...

    arxiv.org/abs/2609.37809 · PDF

  14. 14

    Planetary Feature Fields are Scalable Earth Representations

    Arjun Rao, Sebastian Loeschcke, Anthony Fuller, Isaac Corley, Nico Lang, Evan Shelhamer

    cs.CV · cs.LG

    Satellite observations, precomputed embeddings, and map products describe the same evolving Earth, yet are stored as independent, petabyte-scale data products. Their continued growth calls for compact representations of multiple products while preserving spatial and temporal detail. We introduce Planetary Feature Fields (PFFs), which exploit redundancy across data products by modeling them jointly as continuous functions of space and time at...

    arxiv.org/abs/2609.37784 · PDF

  15. 15

    A Benchmark & Dataset for Detecting AI-Manipulated Visual Evidence in the Court System

    Kelly McConvey, Sajad Ebrahimi, Nima Jamali, Jalehsadat Mahdavimoghaddam, Matina Mahdizadeh Sani, Maksym Taranukhin,...

    cs.CV · cs.AI

    Photographic evidence is becoming increasingly vulnerable to forms of alteration and fabrication that existing legal and technical workflows are not well equipped to evaluate. Surveillance frames, dashcam stills, and phone photographs may be used to establish presence, sequence, causation, damage, or identity, yet contemporary generative systems allow non-experts to alter or fabricate such images through ordinary prompt-based interfaces....

    arxiv.org/abs/2609.37783 · PDF

  16. 16

    HiRAE: Hierarchical Representation Autoencoding with Residual Budgets

    Xuanyu Zhu, Yan Bai, Yang Shi, Yihang Lou, Yuanxing Zhang, Tengfei Liu, Jing Jin, Yuan Zhou

    cs.CV · cs.AI

    Pretrained visual representations support image generation, but may not fully preserve the fine-grained details needed for faithful reconstruction. Meanwhile, intermediate encoder layers contain complementary visual details, but learning to fuse them for reconstruction can produce a latent distribution that is difficult to model. Existing fusion methods require empirical tuning of layer selection or staged optimization of fusion and decoding,...

    arxiv.org/abs/2609.37775 · PDF

  17. 17

    Multi-Site Real-World Performance of Commercial AI for Pulmonary and Incidental Pulmonary Embolism Detection

    Aawez Mansuri, Mohammadreza Chavoshi, Theodorus Dapamede, Wasif Bala, Beatrice Brown-Mulry, Rohan Isaac, Bardia...

    cs.CV · cs.AI

    Pulmonary embolism (PE) is a leading cause of cardiovascular mortality, yet the real-world performance of FDA-cleared AI detection models remains incompletely characterized. We retrospectively evaluated two FDA-cleared AI algorithms from a single commercial platform (Aidoc Medical BriefCase), one for PE triage on dedicated CT pulmonary angiography (CTPA; n = 30,678) and one for incidental PE (iPE) detection on routine contrast-enhanced CTs (n...

    arxiv.org/abs/2609.37750 · PDF

  18. 18

    The Camera Inside the Editor: Reading the Implicit Camera of Image Editors with Painted Calibration Patterns

    Sebastian Rückerl

    cs.CV · cs.LG

    Instruction-based image editors insert objects, restyle scenes and render new viewpoints, but it is unknown which camera they assume when they paint into a photograph. Asked to cover the floor with a checkerboard, an editor paints projective structure from which classical vanishing-point geometry reads pitch, roll, focal length, yaw and, on renders, the principal point, without any training. Unlike a calibrator such as GeoCalib, which...

    arxiv.org/abs/2609.37732 · PDF

  19. 19

    PolyOCR-Venus: Unified OCR Foundation Models for Text-Centric Visual Intelligence

    GuangJian Team, Kaili Huang, Yongshuo Zhang, Bingtao Fu, Changjiang Jiang, Chenfan Qu, Chenfeng Zhang, Fangming Cui,...

    cs.CV · cs.AI

    Optical Character Recognition (OCR) is evolving from plain-text transcription toward general visual intelligence, requiring models to recognize, localize, and reason over textual information in complex visual environments. However, existing OCR systems often excel at only some tasks and struggle to balance recognition, parsing, and reasoning across scenarios. In this report, we present PolyOCR, a family of unified OCR foundation models of...

    arxiv.org/abs/2609.37712 · PDF

  20. 20

    VIF-Bench: Evaluating Visual Instruction Following in Multi-Reference Image Generation

    Yuta Oshima, Masakazu Yoshimura, Masahiro Suzuki, Yutaka Matsuo, Hiroki Furuta

    cs.CV · cs.AI

    Recent multimodal image generation models can take multiple images and textual instructions as input, enabling reference-based generation guided not only by text but also by visual instructions such as layouts, arrows, and pose cues. However, existing benchmarks do not evaluate the joint setting in which multiple references must be composed under multiple and heterogeneous visual-instruction images. To address this gap, we introduce...

    arxiv.org/abs/2609.37709 · PDF

  21. 21

    Are In-Context Images Worth 10 Dimensions?

    Adhemar de Senneville, Xavier Bou, Jérémy Anger, Rafael Grompone, Gabriele Facciolo

    cs.CV · cs.LG

    There has been significant work on understanding the In-Context Learning capabilities of Large Language Models, especially on the induction circuit. For a few-shot classification task, the induction circuit leverages linear representations of each labeled example in-context in order to classify an unlabeled query. However, few works focus on how those linear representations are built in the first place. Leveraging the expressivity of the...

    arxiv.org/abs/2609.37659 · PDF

  22. 22

    TomoTransformer: Towards a Foundation Model for CT Reconstruction

    AmirEhsan Khorashadizadeh, Benjamín Béjar

    cs.CV · cs.LG

    Supervised deep learning has advanced sparse-view tomographic reconstruction. However, conventional models, which typically map filtered back-projection (FBP) images or sinograms to clean reconstructions, are brittle under distribution shifts. Because they require retraining whenever projection counts and angles, detector resolutions, or data distributions change, their deployment in real-world applications remains limited. To address this,...

    arxiv.org/abs/2609.37605 · PDF

  23. 23

    FedSocket: Recipient-Executable Knowledge Exchange for Heterogeneous Multimodal Federated Learning

    Xinyuan Zhao

    cs.CV · cs.AI

    Federated knowledge must remain usable by recipients with different modalities, private architectures, and tasks. We present FedSocket, which makes recipient execution a design requirement of the exchanged model. A shared Q combines recipient-computable inputs, task-owned outputs, and ownership-aware aggregation, connecting heterogeneous private models through a common prediction interface. Private models teach local Q copies; the returned Q...

    arxiv.org/abs/2609.37582 · PDF

  24. 24

    TReVS: Integrating Textual Relevance and Visual Saliency for Efficient Vision-Language Model Token Pruning

    Jing Wang, Zhiping Wu, Dongdong Ren, Youfang Han, Wei Zhao, Wenbin Li

    cs.CV · cs.AI

    Vision-Language Models (VLMs) excel at visual understanding and reasoning but often incur substantial inference costs due to the large number of visual tokens. Recent visual token pruning methods increasingly follow a two-stage paradigm: they first remove visually redundant tokens after the vision encoder and then discard tokens irrelevant to the textual query within the Large Language Model (LLM). However, since the first stage typically...

    arxiv.org/abs/2609.37581 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.