cs.CV · 2026-09-30 · No. 129
Computer Vision and Pattern Recognition, 2026-09-30.
24 new papers in cs.CV. Titles, authors,
abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →
01 — The papers
24 entries-
01
Cropland PAtteRNS: Parallel Dimensional Attention Networks and Attention to Dataset Disparity for Crop Segmentation in Satellite Imagery Time Series Data
Joseph Metcalfe, Sara Sharifzadeh, Fabio Caraffini
cs.CV · cs.LG
The landscape of satellite imagery time series datasets and boundary-pushing architectures for cropland segmentation has never been richer. However, in this gold rush, important truths are being missed on both fronts, as a drive for the most novel concepts or the largest datasets pushes finer details to the side. In this paper, we present our hybrid transformer-convolutional model, Cropland Parallel Attention and Refinement Network for...
-
02
Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies
Hui Ren, Lei Fan, Henry Pao, Han Guo, Zeeshan Zia, Ying Chen, Alexander Schwing, Gang Hua
cs.CV · cs.AI · cs.CL · cs.IR · cs.LG
Answering questions about long videos often requires connecting events involving the same objects across hours or days. Chronological descriptions and text-derived entities can leave physical identity unresolved: different objects may share a description, while observations of the same object remain disconnected across events. Retrieving relevant events therefore does not necessarily recover the "biography" of the particular entity a question...
-
03
Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE
Yu Xu, Yuxin Zhang, Xiao Yang, Haotian Yang, Yizhi Wang, Xinwei Huang, Minxuan Lin, Angtian Wang, Chongyang Ma, Fan Tang
cs.CV · cs.AI
Mixture-of-Experts (MoE), popularized by large language models, is a promising paradigm for scaling visual generative models. However, conventional token-wise MoE routes tokens independently within a homogeneous expert pool and regularizes expert usage toward uniformity, making it poorly matched to video data that is spatiotemporally redundant and semantically long-tailed. We show that existing visual MoEs fall into a uniformity trap:...
-
04
From Unity Simulation to Diffusion-Based Augmentation: Quantifying Dataset Balance for Robust Object Detection
Mohamed Benkedadra, Aissa Saoudi, Maxime Gloesener, Sidi Ahmed Mahmoudi, Matei Mancas
cs.CV · cs.AI
Modern computer vision models achieve high accuracy when trained on large-scale annotated datasets. In critical domains such as construction safety monitoring, data collection is costly, hazardous, and ethically constrained. This paper presents a systematic study comparing two complementary data generation paradigms, (1) Unity Simulation-based rendering and (2) Controllable Diffusion-based generation (CIA), for object detection under real...
-
05
Does Local Video Understanding Transfer Across Encounters? The EgoGears Benchmark
Yuedong Tan, Lei Qi, Yu Liu, Di Wen, Ruiping Liu, Xiaoye Wang, Yufan Chen, Junwei Zheng, Chengzhi Wu, Chen Zhang,...
cs.CV · cs.AI · cs.RO
Embodied systems must make knowledge acquired during one encounter usable in another despite changes in viewpoint, motion, and illumination. Yet aggregate cross-video accuracy conflates failures of local perception with failures to preserve observation identity, establish correspondence, and compose evidence, obscuring whether local video understanding actually transfers. We introduce EgoGears, a complementary single- and multi-video...
-
06
Rollout-Marginal Distillation for Long-Horizon Autoregressive Video Generation
Chenjian Gao, Zhihao Hu, Jianqi Ma, Jun Zhang, Weidong Zhang, Tianfan Xue
cs.CV · cs.AI
Autoregressive (AR) video diffusion enables low-latency, streamable video generation, but prediction errors often accumulate over long rollouts. Training the generator on its own rollouts exposes it to these imperfect histories. However, existing video-level distribution matching distillation (DMD) scores the whole rollout jointly. Because a chunk is evaluated together with its past and future, its correction can favor matching artifacts in...
-
07
SYNCR: Diagnosing and Learning Cross-Video Reasoning from Simulation
Sara Ghazanfari, Siddharth Garg, Prashanth Krishnamurthy, Farshad Khorrami
cs.CV · cs.LG
Reasoning across videos requires aligning events, matching identities, comparing motion, and integrating partial observations. Evaluating these capabilities and testing how to improve them requires both reliable labels and targeted supervision. We introduce SYNCR, a simulator-grounded framework that connects these two needs through shared task generators. Built on Habitat, Kubric, and CLEVRER, SYNCR derives answers from environment state and...
-
08
ReCAP: Retrieval-Guided Capability Reuse for Multimodal Continual Instruction Tuning
Tao Hu, Zhinuo Zhou, Xialiang Tong, De-Chuan Zhan, Da-Wei Zhou
cs.CV · cs.LG
Multimodal continual instruction tuning (MCIT) aims to enable multimodal large language models to acquire new capabilities from sequential tasks while preserving previously learned knowledge. Existing methods primarily mitigate catastrophic forgetting by constraining parameter updates or separating task-specific adaptations. However, continual adaptation can also benefit from external knowledge that provides domain-specific information and...
-
09
Visual Branch is What You Need for CLIP-based Class-Incremental Learning
Tao Hu, Zhen-Hao Xie, Jingcai Guo, De-Chuan Zhan, Da-Wei zhou
cs.CV · cs.LG
Class-Incremental Learning (CIL) requires models to recognize new classes over time without forgetting previously learned ones. With the rise of vision-language pre-training, CLIP has become a strong foundation for CIL. A common design in CLIP-based CIL is to construct textual classifier weights by encoding class-name templates with the CLIP text encoder, and then classify visual features by image-text cosine similarity. This design is...
-
10
HandAnthro: Automated Hand Anthropometry from a Single Image
Fan Zhou, Shuairan Chen, Mengying Zhang, Yulin Wu, Sadegh Jafari, Sixing Yu, Rui Li, Ali Jannesari, Guowen Song
cs.CV · cs.AI
Hand anthropometry supports protective-glove design, but existing measurement methods often require trained operators, specialized hardware, or manual landmarking. We present HandAnthro, which estimates 44 projected hand dimensions from a smartphone photograph of a palm-up hand on US letter-size paper. The pipeline reconstructs wrist-occluded paper boundaries for rectification, whitens non-hand pixels, and refines 41 anthropometry-specific...
-
11
FlowMap-OPD: Rollout--Kernel Separation for On-Policy Distillation of Few-Step Flow-Map Generators
Zhiqi Li, Bo Zhu
cs.CV · cs.LG
Few-step flow-map generators, including MeanFlow and consistency models, enable efficient sampling through long-range transport, yet their on-policy distillation remains underexplored. We introduce FlowMap-OPD, an on-policy distillation framework that separates student-state acquisition from teacher--student distribution comparison. A formulation based on state marginals establishes this separation, while flow--velocity consistency connects...
-
12
Evaluation Choices Shape Biomedical ML Claims: A Pediatric Pneumonia Benchmark Case Study
Bhanu Prakash Vangala, Sowmya Guda, Latha Peddi, Navya Vangala
cs.CV · cs.LG
Biomedical machine learning papers often compress model performance into one headline number. That number can look like a property of the model even when it depends strongly on how the benchmark was evaluated. We study this problem on the widely used Kermany pediatric chest radiograph dataset using nine image classifiers and a controlled evaluation protocol. Under the same protocol, the eight pretrained backbones differ by only 0.026 AUROC....
-
13
Pixel-Level Transformers in Remote Sensing: A Canopy Height Case Study
Sven Ligensa, Jan Pauls, Karsten Schrödter, Ibrahim Fayad, Fabian Gieseke
cs.CV · cs.AI
Predicting canopy height from medium-resolution satellite imagery is a common and scalable approach for assessing the condition of the world's forests, which play a crucial role in climate change mitigation. While Transformer-based architectures have shown strong performance in many domains, their straightforward application to dense (i.e., pixel-level) regression tasks often yields suboptimal results. In particular, the patch size has a...
-
14
Planetary Feature Fields are Scalable Earth Representations
Arjun Rao, Sebastian Loeschcke, Anthony Fuller, Isaac Corley, Nico Lang, Evan Shelhamer
cs.CV · cs.LG
Satellite observations, precomputed embeddings, and map products describe the same evolving Earth, yet are stored as independent, petabyte-scale data products. Their continued growth calls for compact representations of multiple products while preserving spatial and temporal detail. We introduce Planetary Feature Fields (PFFs), which exploit redundancy across data products by modeling them jointly as continuous functions of space and time at...
-
15
A Benchmark & Dataset for Detecting AI-Manipulated Visual Evidence in the Court System
Kelly McConvey, Sajad Ebrahimi, Nima Jamali, Jalehsadat Mahdavimoghaddam, Matina Mahdizadeh Sani, Maksym Taranukhin,...
cs.CV · cs.AI
Photographic evidence is becoming increasingly vulnerable to forms of alteration and fabrication that existing legal and technical workflows are not well equipped to evaluate. Surveillance frames, dashcam stills, and phone photographs may be used to establish presence, sequence, causation, damage, or identity, yet contemporary generative systems allow non-experts to alter or fabricate such images through ordinary prompt-based interfaces....
-
16
HiRAE: Hierarchical Representation Autoencoding with Residual Budgets
Xuanyu Zhu, Yan Bai, Yang Shi, Yihang Lou, Yuanxing Zhang, Tengfei Liu, Jing Jin, Yuan Zhou
cs.CV · cs.AI
Pretrained visual representations support image generation, but may not fully preserve the fine-grained details needed for faithful reconstruction. Meanwhile, intermediate encoder layers contain complementary visual details, but learning to fuse them for reconstruction can produce a latent distribution that is difficult to model. Existing fusion methods require empirical tuning of layer selection or staged optimization of fusion and decoding,...
-
17
Multi-Site Real-World Performance of Commercial AI for Pulmonary and Incidental Pulmonary Embolism Detection
Aawez Mansuri, Mohammadreza Chavoshi, Theodorus Dapamede, Wasif Bala, Beatrice Brown-Mulry, Rohan Isaac, Bardia...
cs.CV · cs.AI
Pulmonary embolism (PE) is a leading cause of cardiovascular mortality, yet the real-world performance of FDA-cleared AI detection models remains incompletely characterized. We retrospectively evaluated two FDA-cleared AI algorithms from a single commercial platform (Aidoc Medical BriefCase), one for PE triage on dedicated CT pulmonary angiography (CTPA; n = 30,678) and one for incidental PE (iPE) detection on routine contrast-enhanced CTs (n...
-
18
The Camera Inside the Editor: Reading the Implicit Camera of Image Editors with Painted Calibration Patterns
Sebastian Rückerl
cs.CV · cs.LG
Instruction-based image editors insert objects, restyle scenes and render new viewpoints, but it is unknown which camera they assume when they paint into a photograph. Asked to cover the floor with a checkerboard, an editor paints projective structure from which classical vanishing-point geometry reads pitch, roll, focal length, yaw and, on renders, the principal point, without any training. Unlike a calibrator such as GeoCalib, which...
-
19
PolyOCR-Venus: Unified OCR Foundation Models for Text-Centric Visual Intelligence
GuangJian Team, Kaili Huang, Yongshuo Zhang, Bingtao Fu, Changjiang Jiang, Chenfan Qu, Chenfeng Zhang, Fangming Cui,...
cs.CV · cs.AI
Optical Character Recognition (OCR) is evolving from plain-text transcription toward general visual intelligence, requiring models to recognize, localize, and reason over textual information in complex visual environments. However, existing OCR systems often excel at only some tasks and struggle to balance recognition, parsing, and reasoning across scenarios. In this report, we present PolyOCR, a family of unified OCR foundation models of...
-
20
VIF-Bench: Evaluating Visual Instruction Following in Multi-Reference Image Generation
Yuta Oshima, Masakazu Yoshimura, Masahiro Suzuki, Yutaka Matsuo, Hiroki Furuta
cs.CV · cs.AI
Recent multimodal image generation models can take multiple images and textual instructions as input, enabling reference-based generation guided not only by text but also by visual instructions such as layouts, arrows, and pose cues. However, existing benchmarks do not evaluate the joint setting in which multiple references must be composed under multiple and heterogeneous visual-instruction images. To address this gap, we introduce...
-
21
Are In-Context Images Worth 10 Dimensions?
Adhemar de Senneville, Xavier Bou, Jérémy Anger, Rafael Grompone, Gabriele Facciolo
cs.CV · cs.LG
There has been significant work on understanding the In-Context Learning capabilities of Large Language Models, especially on the induction circuit. For a few-shot classification task, the induction circuit leverages linear representations of each labeled example in-context in order to classify an unlabeled query. However, few works focus on how those linear representations are built in the first place. Leveraging the expressivity of the...
-
22
TomoTransformer: Towards a Foundation Model for CT Reconstruction
AmirEhsan Khorashadizadeh, Benjamín Béjar
cs.CV · cs.LG
Supervised deep learning has advanced sparse-view tomographic reconstruction. However, conventional models, which typically map filtered back-projection (FBP) images or sinograms to clean reconstructions, are brittle under distribution shifts. Because they require retraining whenever projection counts and angles, detector resolutions, or data distributions change, their deployment in real-world applications remains limited. To address this,...
-
23
FedSocket: Recipient-Executable Knowledge Exchange for Heterogeneous Multimodal Federated Learning
Xinyuan Zhao
cs.CV · cs.AI
Federated knowledge must remain usable by recipients with different modalities, private architectures, and tasks. We present FedSocket, which makes recipient execution a design requirement of the exchanged model. A shared Q combines recipient-computable inputs, task-owned outputs, and ownership-aware aggregation, connecting heterogeneous private models through a common prediction interface. Private models teach local Q copies; the returned Q...
-
24
TReVS: Integrating Textual Relevance and Visual Saliency for Efficient Vision-Language Model Token Pruning
Jing Wang, Zhiping Wu, Dongdong Ren, Youfang Han, Wei Zhao, Wenbin Li
cs.CV · cs.AI
Vision-Language Models (VLMs) excel at visual understanding and reasoning but often incur substantial inference costs due to the large number of visual tokens. Recent visual token pruning methods increasingly follow a two-stage paradigm: they first remove visually redundant tokens after the vision encoder and then discard tokens irrelevant to the textual query within the Large Language Model (LLM). However, since the first stage typically...
This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.
#D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.