cs.CV · 2026-10-01 · No. 130
Computer Vision and Pattern Recognition, 2026-10-01.
24 new papers in cs.CV. Titles, authors,
abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →
01 — The papers
24 entries-
01
ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing
Xinghao Chen, Xiangbo Gao, Jiongze Yu, Yuheng Wu, Zhengzhong Tu
cs.CV · cs.AI
Recent video generation is increasingly realistic and controllable, yet video editing remains less developed, particularly for precise local edits that must preserve the original scene dynamics. Video scene text editing replaces text on scene surfaces, such as storefront signs, whiteboards, and product labels, while preserving the surrounding content, motion, and camera dynamics. Although scene text editing is well studied for images, video...
-
02
Image Classifiers are Efficient Self-Supervised Video Representation Learners
Owais Iqbal, Sudipta Sarkar, Shyam Marjit, Omprakash Chakraborty, Anirban Chakraborty, Abir Das
cs.CV · cs.LG
We introduce VideoMSN, a Masked Siamese Network framework for efficient self-supervised spatio-temporal representation learning in videos. Instead of relying on heavy 3D architectures or reconstruction-based autoencoders for learning with unlabeled data, we repurpose standard image Vision Transformers by representing videos as super images which are grids composed of frames sampled from videos. From each super image, we construct two views:...
-
03
MatLoom: Layered Text-to-Material Generation in a Compact Program Space
Anson Y. Lam, Shuqing Li, Michael R. Lyu
cs.CV · cs.AI · cs.CL · cs.MM
Material generation should produce not only an appearance, but also the rules that construct it. We introduce MatLoom, a compact, layer-oriented language for text-to-material generation with pretrained language models. Each program composes alpha-masked layers whose shared spatial expressions define coverage and physically based rendering (PBR) channels, making dependencies between patterns, color, and relief explicit. A standalone...
-
04
Looped Diffusion Transformer
Yong Xien Chng, Tianyi Chen, Wenwen Tong, Haiwen Diao, Zhongang Cai, Lei Yang, Ziwei Liu, Lewei Lu, Dahua Lin, Gao Huang
cs.CV · cs.LG
Improving text-to-image models has traditionally relied on increasing model size or the number of denoising steps. In this work, we explore an alternative way to scale computation by repeatedly running shared Transformer blocks within each denoising step, effectively increasing computational depth while keeping the parameter count fixed. This looped computation enables iterative refinement of internal representations without explicit...
-
05
ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents
Yong Du, Tongbo Chen, Zhengxi Lu, Yizhou Liu, Bofan Chen, Tao Jiang, Wenhao Xu, Yongliang Shen
cs.CV · cs.AI
Online training enables computer-use agents (CUAs) to improve through interaction with executable environments. However, existing methods primarily rely on sparse outcome rewards, which provide no supervision for intermediate actions. On-policy self-distillation (OPSD) offers token-level learning signals through privileged rescoring, but directly applying it to CUA online training presents two challenges: fixed guidance may become misaligned...
-
06
EviRover: Reinforcing Agentic Perception Beyond a Glance
Kaixuan Fan, Kaituo Feng, Tianshuo Peng, Yilei Jiang, Manyuan Zhang, Junke Wang, Xiangyu Yue
cs.CV · cs.AI
Visual perception is conventionally formulated as a one-shot prediction from a single glance at the image, under the assumption that the image content and the model's parametric knowledge suffice to resolve the query. This assumption often fails in real-world scenarios that hinge on fine-grained visual details or require knowledge-intensive and up-to-date information. We term such cases \textit{perception under insufficient evidence} and...
-
07
Learning Skills from Historical Action Trajectories: Action Experience Dictionary for World Action Models
Qi Lyu, Jiahua Dong, Hao Shen, Xudong Wang, Hongyuan Yu, Baichen Liu, Henghui Ding, Zhi Han, Nicu Sebe, Ivan Laptev,...
cs.CV · cs.AI
World Action Models (WAMs) couple visual dynamics prediction with action generation, yet they do not explicitly support the reuse of action experience across manipulation tasks. Furthermore, existing WAMs struggle to capture underlying cross-task semantic relationships that could guide target action prediction, as redundant background elements interfere with the extraction of key visual information. To address these challenges, we develop a...
-
08
MemLife: Curating and Reasoning over Long-Term Egocentric Video Memories
Guangzhi Xiong, Xinyuan Zhang, Xiao Yang, Hyokun Yun, Kai Zhang, Shiun-Zu Kuo, Hyeonjeong Ha, Xilun Chen, Kai Sun,...
cs.CV · cs.AI · cs.CL
Long-term egocentric video enables personalized AI assistants to reason about daily life. However, as video histories grow to hundreds of hours spanning months or years, reprocessing raw clips for every query becomes computationally prohibitive. Memory systems offer a scalable alternative by compacting videos into text representations, but often fail on practical benchmarks: either the memory does not preserve key evidence, or the retriever...
-
09
GateSPINE: Gated Cross-View Fusion for Lumbar Spine MRI Report Generation
Hoang Nguyen Van, Cuong Vuong Tuan, Trang Mai Xuan, Bien Tran Van, Nam Tran Van, Thien Van Luong
cs.CV · cs.AI
Automated report generation can ease the burden radiolo gists face when interpreting multi-sequence MRI studies. Unlike CT, MRI examinations comprise multiple sequences and imaging planes, each con tributing complementary diagnostic information. Existing methods en code a study as a single volume and combine multiple acquisitions by fixed rules. Findings visible in only one plane are thus diluted and of ten missed, lowering recall on clinical...
-
10
LongEmo: Towards Emotion Understanding and Reasoning in Long Videos
Shuo Zhang, Yifan Zhou, Han Wang, Jinsong Zhang, Jingyu Li, Hongbing Li, Zhejun Zhang, Chengyi Zhao, Yuquan Hao,...
cs.CV · cs.AI
While recent Multimodal Large Language Models (MLLMs) have shown promise in affective computing, their reasoning capabilities are largely confined to short video clips with limited interactions. However, real-world emotions are not merely isolated instantaneous reactions but dynamic and cumulative processes deeply shaped by past experiences and ongoing events. To bridge this gap, we introduce LongEmoBench, a benchmark dedicated to emotion...
-
11
Less Data, Better Timing: Student-Curriculum Coupling for VLM On-Policy Distillation in Temporal Video Grounding
Jiacheng Qiu, Yunsoo Kim, Ruichen Xu, Jian Luo, Petar M. Djurić, Sima Mofakham
cs.CV · cs.AI
On-policy distillation (OPD) provides dense supervision directly on student-generated trajectories, making it an effective post-training strategy for vision-language models in temporal video grounding (TVG). However, existing pipelines typically construct the training curriculum from a fixed teacher and the initial student state, implicitly assuming that selected examples retain positive supervision value throughout optimization. We show that...
-
12
WARP: A Unified Benchmark for Invisible Image Watermarking -- Robustness and Protection Against Attacks
Khaled Abud, Aleksey Yakushev, Aleksandr Akimenkov, Irina Serzhenko, Kirill Aistov, Egor Kovalev, Dmitry Obydenkov,...
cs.CV · cs.AI · cs.MM
Digital image watermarking is increasingly critical in media contexts, as emerging regulations and industry practices require marking AI-generated content and ensuring traceable sources to prevent manipulation or misuse. Recent advances in invisible watermarking methods highlight the need to update existing benchmarking practices to reflect current techniques and evaluation criteria. We address this by introducing WARP -- a unified framework...
-
13
CoVisco: Codec-Native Vision Encoder with Native Token Compression for Unified Image-Video Understanding
Yulong Liu, Xiaotian Han, Junyuan Shang, Yuchen Ding, Zhenyu Zhang, Shuohuan Wang, Guibo Zhu, Sirui Han, Dianhai Yu
cs.CV · cs.AI
Vision-language models face a fundamental scaling bottleneck: the number of visual tokens grows with both temporal duration and spatial resolution, making long-video understanding expensive for the vision encoder and the language model. Existing methods often compress visual tokens after dense encoding, creating a mismatch between the representation used during training and the compact interface required at deployment. We present CoVisco, a...
-
14
Spherical Interpolation for Backward-Compatible Multimodal Representations
Simone Ricci, Niccolò Biondi, Federico Pernici
cs.CV · cs.LG
Contrastive vision-language models map visual and textual representations into a shared normalized embedding space, making cosine similarity the natural metric for cross-modal retrieval. A practical challenge arises during model upgrades: independently trained models generally produce incompatible representation spaces, so replacing a deployed model typically requires recomputing embeddings for the entire gallery, which is prohibitively...
-
15
Let the Carrier Carry the Attack: Preserving the Subject in Adversarial Image Generation
Linfeng Jiang, Steven McDonagh, Yuhang Chen, Xingyu Zhao, Siddartha Khastgir, Andi Zhang
cs.CV · cs.AI
Strong unrestricted adversarial attacks can distort the primary object of an image, hereafter referred to as the subject. To preserve subject integrity without compromising attack magnitude, we introduce the carrier: a secondary visual element that provides an auxiliary region to facilitate the attack under global classifier guidance. We demonstrate three key findings: 1. A carrier mitigates subject distortion by absorbing a larger share of...
-
16
When Masking Helps or Hurts Robustness in Compressed CLIP: A Pre-Deployment Diagnostic
Muhammad Zawish, Steven Davy
cs.CV · cs.AI
This paper demonstrate that whether masking-based token pruning helps or hurts worst-group robustness can be predicted before deployment, without labels or fine-tuning. A systematic study of semantic masking across 8 spurious-correlation benchmarks shows its effect on worst-group accuracy is highly unstable: it improves accuracy by up to 82.5\% relative on some datasets and degrades it by up to 100\% on others. We trace this instability to...
-
17
ShieldCLIP: Selective Safety Alignment for Harmful Content Mitigation in Multimodal Foundation Models
Tobia Poppi, Silvia Cappelletti, Samuele Poppi, Marcella Cornia, Lorenzo Baraldi, Diego Garcia-Olano, Rita Cucchiara
cs.CV · cs.AI · cs.CL · cs.MM
Multimodal encoders such as CLIP underlie many downstream systems, but their web-scale training data embed harmful associations that safety alignment must suppress without unnecessarily changing benign representations. Because ethical and practical constraints prevent collecting real unsafe content at scale, existing datasets pair safe real samples with generated counterparts, but label every generated sample unsafe, even when one modality is...
-
18
Unapologetically Distributed: A Call for Decentralized Document Analysis
Adrià Molina, Oriol Ramos Terrades, Josep Lladós
cs.CV · cs.LG
Privacy has become an increasingly important concern in the Document Analysis community, to the extent that in many environments such as archives, governmental institutions, and local businesses, the adoption of automation is restricted by legal and policy constraints. While federated learning has often been regarded as a ``necessary evil'', implying an unavoidable performance trade-off in exchange for decentralization and privacy, many prior...
-
19
SAGE: Salient Factor Discovery and Generation with Visual Foundation Representations
Shuang Liang, Lejun Liao, Shiyuan Zhang, Max C. Zhang, Xiaolong Luo, Han Wang, Stefano Anzellotti, Yuan Yuan
cs.CV · cs.LG
Given a target dataset, such as faces with eyeglasses, and a background dataset, such as faces without, contrastive analysis separates \textit{salient} factors specific to the target from \textit{common} content shared by both. We aim for salient representations that capture target-specific detail in each image, such as the shape, color, and position of the glasses, so that they reveal subtypes without subtype labels and guide the generation...
-
20
D-Scope: Decomposing and Steering Diffusion Transformers with Sparse Autoencoders
Xinyue Xu, Jiahao Zhang, Lijie Hu, Peter Hase, Hao Wang
cs.CV · cs.AI
Sparse autoencoders (SAEs) reveal visual structure in diffusion transformers (DiTs), but interpreting a feature does not establish whether it can be used to control generation. We introduce D-Scope (Diffusion Scope), a framework that connects feature interpretation to generation control through shared visual evidence. D-Scope aggregates SigLIP~2 embeddings of highly activating image patches into visual centroids. Matching target text...
-
21
GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang,...
cs.CV · cs.AI · cs.LG · cs.RO
Precise grounding matters. It specifies which object is the target and where that object is, even in clutter and for tiny objects, and it has to be fast enough for closed-loop control. Yet vision-language-action (VLA) and world-action models (WAMs) take perception from general-purpose vision-language and video-generation backbones, which still fail in these settings. We introduce GroundingPI, a 4B grounding foundation model that generates...
-
22
GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed
Qize Yu, Lianrui Fan, Bowen Ping, Xini Ding, Zetian Song, Junbo Niu, Kaixuan Wang, Tianxing Chen, Yue Chen, Minghua...
cs.CV · cs.AI · cs.LG · cs.RO
Autoregressive (AR) grounding models serialize spatial predictions, introducing sequential latency and imposing a causal order on output tokens. We view grounding as visual evidence extraction: objects, locations, and spatial relations are jointly constrained by the image and query, yet their dependencies do not imply an intrinsic left-to-right generation order. This distinction makes bidirectional diffusion a natural fit, allowing spatial...
-
23
KilometerVision: A New Frontier for Large-Scale Spatial Intelligence in VLMs
Aravindh Mahendran, Michael King, Matthew Koichi Grimes, Antoine Yang, Tyler Zhu, Joseph Heyward, Tengda Han, Shiry...
cs.CV · cs.LG
We push the frontier of large-scale spatial intelligence in Vision-Language Models (VLMs) and introduce the first benchmark that probes geographical layout understanding from real-world videos, spanning up to 1km distances. Inspired by the cognitive science literature, we evaluate models against the hierarchical stages of human spatial awareness: anchoring via landmarks, connecting them through routes, and integrating these into global mental...
-
24
Steering Fields: Adaptive Vector Fields for Safe Image Generation and Beyond
Simone Facchiano, Jan Eric Lenssen, Bernt Schiele, Wolfgang Stammer, Fabio Galasso, Jonas Fischer
cs.CV · cs.AI
As state-of-the-art text-to-image flow models achieve near-photorealistic quality, controlling their outputs, e.g., suppressing harmful content while promoting benign alternatives, has become a central challenge. The current steering paradigm consists of adding a global steering vector to selected activations. While functional, a fixed and example-agnostic vector applied uniformly along the entire trajectory cannot adapt to the changing state...
This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.
#D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.