cs.CV · 2026-10-01 · No. 130

Computer Vision and Pattern Recognition, 2026-10-01.

24 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

24 entries
  1. 01

    ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing

    Xinghao Chen, Xiangbo Gao, Jiongze Yu, Yuheng Wu, Zhengzhong Tu

    cs.CV · cs.AI

    Recent video generation is increasingly realistic and controllable, yet video editing remains less developed, particularly for precise local edits that must preserve the original scene dynamics. Video scene text editing replaces text on scene surfaces, such as storefront signs, whiteboards, and product labels, while preserving the surrounding content, motion, and camera dynamics. Although scene text editing is well studied for images, video...

    arxiv.org/abs/2609.40356 · PDF

  2. 02

    Image Classifiers are Efficient Self-Supervised Video Representation Learners

    Owais Iqbal, Sudipta Sarkar, Shyam Marjit, Omprakash Chakraborty, Anirban Chakraborty, Abir Das

    cs.CV · cs.LG

    We introduce VideoMSN, a Masked Siamese Network framework for efficient self-supervised spatio-temporal representation learning in videos. Instead of relying on heavy 3D architectures or reconstruction-based autoencoders for learning with unlabeled data, we repurpose standard image Vision Transformers by representing videos as super images which are grids composed of frames sampled from videos. From each super image, we construct two views:...

    arxiv.org/abs/2609.40347 · PDF

  3. 03

    MatLoom: Layered Text-to-Material Generation in a Compact Program Space

    Anson Y. Lam, Shuqing Li, Michael R. Lyu

    cs.CV · cs.AI · cs.CL · cs.MM

    Material generation should produce not only an appearance, but also the rules that construct it. We introduce MatLoom, a compact, layer-oriented language for text-to-material generation with pretrained language models. Each program composes alpha-masked layers whose shared spatial expressions define coverage and physically based rendering (PBR) channels, making dependencies between patterns, color, and relief explicit. A standalone...

    arxiv.org/abs/2609.40322 · PDF

  4. 04

    Looped Diffusion Transformer

    Yong Xien Chng, Tianyi Chen, Wenwen Tong, Haiwen Diao, Zhongang Cai, Lei Yang, Ziwei Liu, Lewei Lu, Dahua Lin, Gao Huang

    cs.CV · cs.LG

    Improving text-to-image models has traditionally relied on increasing model size or the number of denoising steps. In this work, we explore an alternative way to scale computation by repeatedly running shared Transformer blocks within each denoising step, effectively increasing computational depth while keeping the parameter count fixed. This looped computation enables iterative refinement of internal representations without explicit...

    arxiv.org/abs/2609.40305 · PDF

  5. 05

    ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents

    Yong Du, Tongbo Chen, Zhengxi Lu, Yizhou Liu, Bofan Chen, Tao Jiang, Wenhao Xu, Yongliang Shen

    cs.CV · cs.AI

    Online training enables computer-use agents (CUAs) to improve through interaction with executable environments. However, existing methods primarily rely on sparse outcome rewards, which provide no supervision for intermediate actions. On-policy self-distillation (OPSD) offers token-level learning signals through privileged rescoring, but directly applying it to CUA online training presents two challenges: fixed guidance may become misaligned...

    arxiv.org/abs/2609.40253 · PDF

  6. 06

    EviRover: Reinforcing Agentic Perception Beyond a Glance

    Kaixuan Fan, Kaituo Feng, Tianshuo Peng, Yilei Jiang, Manyuan Zhang, Junke Wang, Xiangyu Yue

    cs.CV · cs.AI

    Visual perception is conventionally formulated as a one-shot prediction from a single glance at the image, under the assumption that the image content and the model's parametric knowledge suffice to resolve the query. This assumption often fails in real-world scenarios that hinge on fine-grained visual details or require knowledge-intensive and up-to-date information. We term such cases \textit{perception under insufficient evidence} and...

    arxiv.org/abs/2609.40230 · PDF

  7. 07

    Learning Skills from Historical Action Trajectories: Action Experience Dictionary for World Action Models

    Qi Lyu, Jiahua Dong, Hao Shen, Xudong Wang, Hongyuan Yu, Baichen Liu, Henghui Ding, Zhi Han, Nicu Sebe, Ivan Laptev,...

    cs.CV · cs.AI

    World Action Models (WAMs) couple visual dynamics prediction with action generation, yet they do not explicitly support the reuse of action experience across manipulation tasks. Furthermore, existing WAMs struggle to capture underlying cross-task semantic relationships that could guide target action prediction, as redundant background elements interfere with the extraction of key visual information. To address these challenges, we develop a...

    arxiv.org/abs/2609.40219 · PDF

  8. 08

    MemLife: Curating and Reasoning over Long-Term Egocentric Video Memories

    Guangzhi Xiong, Xinyuan Zhang, Xiao Yang, Hyokun Yun, Kai Zhang, Shiun-Zu Kuo, Hyeonjeong Ha, Xilun Chen, Kai Sun,...

    cs.CV · cs.AI · cs.CL

    Long-term egocentric video enables personalized AI assistants to reason about daily life. However, as video histories grow to hundreds of hours spanning months or years, reprocessing raw clips for every query becomes computationally prohibitive. Memory systems offer a scalable alternative by compacting videos into text representations, but often fail on practical benchmarks: either the memory does not preserve key evidence, or the retriever...

    arxiv.org/abs/2609.40195 · PDF

  9. 09

    GateSPINE: Gated Cross-View Fusion for Lumbar Spine MRI Report Generation

    Hoang Nguyen Van, Cuong Vuong Tuan, Trang Mai Xuan, Bien Tran Van, Nam Tran Van, Thien Van Luong

    cs.CV · cs.AI

    Automated report generation can ease the burden radiolo gists face when interpreting multi-sequence MRI studies. Unlike CT, MRI examinations comprise multiple sequences and imaging planes, each con tributing complementary diagnostic information. Existing methods en code a study as a single volume and combine multiple acquisitions by fixed rules. Findings visible in only one plane are thus diluted and of ten missed, lowering recall on clinical...

    arxiv.org/abs/2609.40091 · PDF

  10. 10

    LongEmo: Towards Emotion Understanding and Reasoning in Long Videos

    Shuo Zhang, Yifan Zhou, Han Wang, Jinsong Zhang, Jingyu Li, Hongbing Li, Zhejun Zhang, Chengyi Zhao, Yuquan Hao,...

    cs.CV · cs.AI

    While recent Multimodal Large Language Models (MLLMs) have shown promise in affective computing, their reasoning capabilities are largely confined to short video clips with limited interactions. However, real-world emotions are not merely isolated instantaneous reactions but dynamic and cumulative processes deeply shaped by past experiences and ongoing events. To bridge this gap, we introduce LongEmoBench, a benchmark dedicated to emotion...

    arxiv.org/abs/2609.40079 · PDF

  11. 11

    Less Data, Better Timing: Student-Curriculum Coupling for VLM On-Policy Distillation in Temporal Video Grounding

    Jiacheng Qiu, Yunsoo Kim, Ruichen Xu, Jian Luo, Petar M. Djurić, Sima Mofakham

    cs.CV · cs.AI

    On-policy distillation (OPD) provides dense supervision directly on student-generated trajectories, making it an effective post-training strategy for vision-language models in temporal video grounding (TVG). However, existing pipelines typically construct the training curriculum from a fixed teacher and the initial student state, implicitly assuming that selected examples retain positive supervision value throughout optimization. We show that...

    arxiv.org/abs/2609.40055 · PDF

  12. 12

    WARP: A Unified Benchmark for Invisible Image Watermarking -- Robustness and Protection Against Attacks

    Khaled Abud, Aleksey Yakushev, Aleksandr Akimenkov, Irina Serzhenko, Kirill Aistov, Egor Kovalev, Dmitry Obydenkov,...

    cs.CV · cs.AI · cs.MM

    Digital image watermarking is increasingly critical in media contexts, as emerging regulations and industry practices require marking AI-generated content and ensuring traceable sources to prevent manipulation or misuse. Recent advances in invisible watermarking methods highlight the need to update existing benchmarking practices to reflect current techniques and evaluation criteria. We address this by introducing WARP -- a unified framework...

    arxiv.org/abs/2609.40031 · PDF

  13. 13

    CoVisco: Codec-Native Vision Encoder with Native Token Compression for Unified Image-Video Understanding

    Yulong Liu, Xiaotian Han, Junyuan Shang, Yuchen Ding, Zhenyu Zhang, Shuohuan Wang, Guibo Zhu, Sirui Han, Dianhai Yu

    cs.CV · cs.AI

    Vision-language models face a fundamental scaling bottleneck: the number of visual tokens grows with both temporal duration and spatial resolution, making long-video understanding expensive for the vision encoder and the language model. Existing methods often compress visual tokens after dense encoding, creating a mismatch between the representation used during training and the compact interface required at deployment. We present CoVisco, a...

    arxiv.org/abs/2609.39924 · PDF

  14. 14

    Spherical Interpolation for Backward-Compatible Multimodal Representations

    Simone Ricci, Niccolò Biondi, Federico Pernici

    cs.CV · cs.LG

    Contrastive vision-language models map visual and textual representations into a shared normalized embedding space, making cosine similarity the natural metric for cross-modal retrieval. A practical challenge arises during model upgrades: independently trained models generally produce incompatible representation spaces, so replacing a deployed model typically requires recomputing embeddings for the entire gallery, which is prohibitively...

    arxiv.org/abs/2609.39836 · PDF

  15. 15

    Let the Carrier Carry the Attack: Preserving the Subject in Adversarial Image Generation

    Linfeng Jiang, Steven McDonagh, Yuhang Chen, Xingyu Zhao, Siddartha Khastgir, Andi Zhang

    cs.CV · cs.AI

    Strong unrestricted adversarial attacks can distort the primary object of an image, hereafter referred to as the subject. To preserve subject integrity without compromising attack magnitude, we introduce the carrier: a secondary visual element that provides an auxiliary region to facilitate the attack under global classifier guidance. We demonstrate three key findings: 1. A carrier mitigates subject distortion by absorbing a larger share of...

    arxiv.org/abs/2609.39723 · PDF

  16. 16

    When Masking Helps or Hurts Robustness in Compressed CLIP: A Pre-Deployment Diagnostic

    Muhammad Zawish, Steven Davy

    cs.CV · cs.AI

    This paper demonstrate that whether masking-based token pruning helps or hurts worst-group robustness can be predicted before deployment, without labels or fine-tuning. A systematic study of semantic masking across 8 spurious-correlation benchmarks shows its effect on worst-group accuracy is highly unstable: it improves accuracy by up to 82.5\% relative on some datasets and degrades it by up to 100\% on others. We trace this instability to...

    arxiv.org/abs/2609.39704 · PDF

  17. 17

    ShieldCLIP: Selective Safety Alignment for Harmful Content Mitigation in Multimodal Foundation Models

    Tobia Poppi, Silvia Cappelletti, Samuele Poppi, Marcella Cornia, Lorenzo Baraldi, Diego Garcia-Olano, Rita Cucchiara

    cs.CV · cs.AI · cs.CL · cs.MM

    Multimodal encoders such as CLIP underlie many downstream systems, but their web-scale training data embed harmful associations that safety alignment must suppress without unnecessarily changing benign representations. Because ethical and practical constraints prevent collecting real unsafe content at scale, existing datasets pair safe real samples with generated counterparts, but label every generated sample unsafe, even when one modality is...

    arxiv.org/abs/2609.39688 · PDF

  18. 18

    Unapologetically Distributed: A Call for Decentralized Document Analysis

    Adrià Molina, Oriol Ramos Terrades, Josep Lladós

    cs.CV · cs.LG

    Privacy has become an increasingly important concern in the Document Analysis community, to the extent that in many environments such as archives, governmental institutions, and local businesses, the adoption of automation is restricted by legal and policy constraints. While federated learning has often been regarded as a ``necessary evil'', implying an unavoidable performance trade-off in exchange for decentralization and privacy, many prior...

    arxiv.org/abs/2609.39684 · PDF

  19. 19

    SAGE: Salient Factor Discovery and Generation with Visual Foundation Representations

    Shuang Liang, Lejun Liao, Shiyuan Zhang, Max C. Zhang, Xiaolong Luo, Han Wang, Stefano Anzellotti, Yuan Yuan

    cs.CV · cs.LG

    Given a target dataset, such as faces with eyeglasses, and a background dataset, such as faces without, contrastive analysis separates \textit{salient} factors specific to the target from \textit{common} content shared by both. We aim for salient representations that capture target-specific detail in each image, such as the shape, color, and position of the glasses, so that they reveal subtypes without subtype labels and guide the generation...

    arxiv.org/abs/2609.39635 · PDF

  20. 20

    D-Scope: Decomposing and Steering Diffusion Transformers with Sparse Autoencoders

    Xinyue Xu, Jiahao Zhang, Lijie Hu, Peter Hase, Hao Wang

    cs.CV · cs.AI

    Sparse autoencoders (SAEs) reveal visual structure in diffusion transformers (DiTs), but interpreting a feature does not establish whether it can be used to control generation. We introduce D-Scope (Diffusion Scope), a framework that connects feature interpretation to generation control through shared visual evidence. D-Scope aggregates SigLIP~2 embeddings of highly activating image patches into visual centroids. Matching target text...

    arxiv.org/abs/2609.39625 · PDF

  21. 21

    GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

    Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang,...

    cs.CV · cs.AI · cs.LG · cs.RO

    Precise grounding matters. It specifies which object is the target and where that object is, even in clutter and for tiny objects, and it has to be fast enough for closed-loop control. Yet vision-language-action (VLA) and world-action models (WAMs) take perception from general-purpose vision-language and video-generation backbones, which still fail in these settings. We introduce GroundingPI, a 4B grounding foundation model that generates...

    arxiv.org/abs/2609.39601 · PDF

  22. 22

    GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed

    Qize Yu, Lianrui Fan, Bowen Ping, Xini Ding, Zetian Song, Junbo Niu, Kaixuan Wang, Tianxing Chen, Yue Chen, Minghua...

    cs.CV · cs.AI · cs.LG · cs.RO

    Autoregressive (AR) grounding models serialize spatial predictions, introducing sequential latency and imposing a causal order on output tokens. We view grounding as visual evidence extraction: objects, locations, and spatial relations are jointly constrained by the image and query, yet their dependencies do not imply an intrinsic left-to-right generation order. This distinction makes bidirectional diffusion a natural fit, allowing spatial...

    arxiv.org/abs/2609.39600 · PDF

  23. 23

    KilometerVision: A New Frontier for Large-Scale Spatial Intelligence in VLMs

    Aravindh Mahendran, Michael King, Matthew Koichi Grimes, Antoine Yang, Tyler Zhu, Joseph Heyward, Tengda Han, Shiry...

    cs.CV · cs.LG

    We push the frontier of large-scale spatial intelligence in Vision-Language Models (VLMs) and introduce the first benchmark that probes geographical layout understanding from real-world videos, spanning up to 1km distances. Inspired by the cognitive science literature, we evaluate models against the hierarchical stages of human spatial awareness: anchoring via landmarks, connecting them through routes, and integrating these into global mental...

    arxiv.org/abs/2609.39588 · PDF

  24. 24

    Steering Fields: Adaptive Vector Fields for Safe Image Generation and Beyond

    Simone Facchiano, Jan Eric Lenssen, Bernt Schiele, Wolfgang Stammer, Fabio Galasso, Jonas Fischer

    cs.CV · cs.AI

    As state-of-the-art text-to-image flow models achieve near-photorealistic quality, controlling their outputs, e.g., suppressing harmful content while promoting benign alternatives, has become a central challenge. The current steering paradigm consists of adding a global steering vector to selected activations. While functional, a fixed and example-agnostic vector applied uniformly along the entire trajectory cannot adapt to the changing state...

    arxiv.org/abs/2609.39573 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.