cs.CV · 2026-10-05 · No. 134

Computer Vision and Pattern Recognition, 2026-10-05.

18 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

18 entries
  1. 01

    Less Decoder is More Encoder: Geometric Representation Learning from Novel View Synthesis

    Keerthi Kaashyap, Dennis Anthony, Akshay Krishnan, Nhi Ngoc Nguyen, Jeremy Collins, James Hays, Shreyas Kousik, Animesh Garg

    cs.CV · cs.AI · cs.RO

    This paper examines the role of Novel View Synthesis (NVS) in geometric representation learning. In principle, NVS should reason about 3D scene structure, thereby enabling transferable multi-view geometric representations. Yet, existing encoder-based NVS methods yield poor representations. This is not because of a lack of supervisory signal, but rather due to inconspicuous architectural choices: \textit{spatially expressive decoders} that...

    arxiv.org/abs/2610.03717 · PDF

  2. 02

    4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes

    Ruihong Shen, Žiga Kovačič, Peter Kulits, Xingrui Wang, Zizhang Li, Joshua B. Tenenbaum, Alan Yuille, Jieneng Chen, Jiajun Wu

    cs.CV · cs.AI · cs.GR

    We introduce 4DCodeBench, a benchmark for 4D inverse graphics through code generation, in which agents reconstruct dynamic scenes from video as executable graphics programs. To accomplish this, agents must translate visual observations into compact representations of scene structure and dynamics, by implementing abstractions such as physical simulations to reproduce complex behavior. To evaluate this capability, we curate a set of real-world...

    arxiv.org/abs/2610.03715 · PDF

  3. 03

    On-Board Anomaly Detection for Efficient Marine Environmental Monitoring

    Thomas Goudemant, Clotilde Szywala, Benjamin Francesconi, Michelle Aubrun, Yves Bobichon, Marjorie Bellizzi, Adrien Girard

    cs.CV · cs.AI · cs.LG

    Marine ecosystems are impacted by various threats such as oil spills, algal blooms, and sediment floods, which disrupt habitats, wildlife, and human activities. Advances in satellite imagery and Artificial Intelligence (AI) have enhanced our capabilities for early detection and mitigation of such hazards. In this paper, we propose a marine event detection pipeline for Earth observation satellites equipped with multi- or hyperspectral sensors....

    arxiv.org/abs/2610.03649 · PDF

  4. 04

    LoGo: Local-Global Rewards for Consistent Long-Horizon Video Generation

    Ziqi Ma, Shreya Sharma, Mohamed El Banani, Katja Schwarz, Chongjie Ye, Chao-Yuan Wu, Li Fei-Fei, Ben Mildenhall,...

    cs.CV · cs.AI

    Camera-controlled video models are rapidly advancing toward long generation horizons and complex camera control. A key failure mode is 3D inconsistency: as the camera moves, objects lose permanence and scene structures shift. Existing post-training techniques, which assign a single scalar reward to the entire generation, are poorly suited to correcting these inconsistencies over long horizons. We introduce LoGo, which blends global and...

    arxiv.org/abs/2610.03636 · PDF

  5. 05

    Rethinking What to Cache in Few-Step Diffusion Transformers: Solver-Aware Target Selection

    Shuo Yang, Lihao Fang, Yi Zhang, Haixiang Wang, Xincheng Ye, Shufan Chen, Jipeng Guo, Youqing Wang

    cs.CV · cs.AI

    Diffusion Transformers (DiTs) can generate high-quality images and videos, but generating each sample requires multiple costly DiT forward passes. Two common ways to accelerate DiT sampling are step distillation, which reduces the number of sampling steps, and caching, which skips some DiT evaluations by reusing a tensor computed at an earlier step. Most caching methods decide in advance which tensor to reuse. After distillation, adjacent...

    arxiv.org/abs/2610.03577 · PDF

  6. 06

    Weave Forcing: Compositional Memory Routing for Interactive Long Video Generation

    Ziyi Wang, Junchi Yao, Heqian Qiu, Wenbo Shi, Chengjiu Wang, Jinyang He, Binkai Hong, Hongliang Li

    cs.CV · cs.AI

    Recent advances in autoregressive video generation have improved temporal consistency over extended durations, yet interactive storytelling requires more than continuous scene extension: a new shot may combine characters and backgrounds from different historical shots. Whole prompt retrieval can overlook the distinct reference needs of individual components, while directly combining all historical memories may introduce unrelated visual...

    arxiv.org/abs/2610.03510 · PDF

  7. 07

    Preserving Anatomical Continuity: Three-Stage Pipeline for Colon Segmentation in 3D Abdominal CT Scans

    Deshan Kalupahana, Sonit Singh, Praveen Ravindran, Arcot Sowmya

    cs.CV · cs.AI

    Accurate colon segmentation from CT images is essential for colorectal disease analysis, yet deep learning based methods often produce disconnected predictions due to complex anatomy. This study introduces a three-stage, topology-preserving segmentation pipeline to address this issue. The first stage performs initial deep learning-based segmentation, followed by centreline bridging to reconnect disjoint regions and a reconstruction stage to...

    arxiv.org/abs/2610.03467 · PDF

  8. 08

    Corrupted but Correct: Why Vision-Language Models Lie to Themselves Internally

    Arun Josephraj Arokiaraj, Zekun Wu, Adriano Koshiyama

    cs.CV · cs.AI

    A targeted adversarial perturbation can drive a vision-language model's (VLM's) teacher-forced training loss for a fixed target caption to near zero, yet the same model, allowed to generate freely, produces the original, correct description with no trace of the target. We call this dissociation the train/inference gap, and give it a precise mechanistic account on Qwen2.5-VL-7B-Instruct using a controlled two-stage PGD attack on 200 held-out...

    arxiv.org/abs/2610.03445 · PDF

  9. 09

    ForestQuery: Boundary-Aware and Spatially Anchored Query Learning for Unified Forest Point Cloud Segmentation

    Zhihao Zhan, Le Tao, Yifei Tian, Xin Liu, Jie Yuan

    cs.CV · cs.AI

    Forest point cloud segmentation is fundamental for fine-grained 3D forest scene understanding, yet remains challenging due to irregular tree structures, severe occlusions, density variations, and ambiguous instance boundaries. Recent query-based forest segmentation methods have shown promise for unified semantic and instance prediction, but they still insufficiently exploit forest-specific spatial structure and account for boundary...

    arxiv.org/abs/2610.03403 · PDF

  10. 10

    From Patching to Pruning Visual Computation in Vision Language Models

    Rahul Chowdhury, Timothy A Rupprecht, Xuan Shen, Shaoyi Huang, Pu Zhao, Yanzhi Wang

    cs.CV · cs.LG

    Vision language models (VLMs) incur substantial inference cost because every visual token is processed by the attention and MLP projections of every decoder layer, even when token-specific visual computation is unnecessary at many depths. We introduce Patch-to-Prune (P2P), inspired by Mechanistic Interpretability, a training-free framework that converts activation patching from a diagnostic tool into an inference-time computation bypass. P2P...

    arxiv.org/abs/2610.03389 · PDF

  11. 11

    Consecutive Posterior Fusion for Diffusive Recovery of Unobservable Image Structures

    Elena Morotti, Davide Evangelista, Elena Loli Piccolomini

    cs.CV · cs.AI

    Solving severely ill-posed imaging inverse problems requires recovering image structures that are unobservable or weakly constrained by the measurements. Diffusion models provide expressive learned priors for inferring such missing information, while posterior sampling incorporates measurement consistency along the reverse process. Standard diffusion posterior samplers, however, rely on instantaneous measurement-aware estimates, without...

    arxiv.org/abs/2610.03261 · PDF

  12. 12

    Uncertainty as a Proxy for Semantic Correctness in Diffusion-Based Medical Image Synthesis

    Yuxuan Ou, Konstantinos Kamnitsas, OxAAA Study, AICT Consortium, Regent Lee, Vicente Grau

    cs.CV · cs.AI

    Diffusion models can synthesise contrast-enhanced CT (CECT) from non-contrast CT (NCCT), avoiding contrast administration and its environmental and patient-access costs. However, visually realistic images are not necessarily anatomically correct, and the pixel-intensity and feature-space similarity metrics used to assess generation quality do not directly measure anatomical correctness. In this work, we investigate whether uncertainty can...

    arxiv.org/abs/2610.03224 · PDF

  13. 13

    Contextual Flow Matching: Adaptive Step Selection in Flow Models for Efficient Visual Generation

    Divya Jyoti Bajpai, Arun Verma, Manjesh Kumar Hanawal

    cs.CV · cs.AI

    Flow Matching enables high-quality visual generation via continuous-time dynamics, but inference remains costly due to multiple sequential function evaluations. Existing acceleration methods reduce the number of function evaluations but often introduce additional training overhead, degrade quality, or fail to account for input-dependent variability. We propose COFLOW, an inference-time method that adaptively selects the step counts each...

    arxiv.org/abs/2610.03202 · PDF

  14. 14

    Does Physics Live in the Activations? Localizing Physical Quantities in Video Diffusion Models

    Jonas Kneifl, Jakub Skalski, Bartłomiej Twardowski, Kamil Deja

    cs.CV · cs.LG

    Video generation models produce strikingly realistic sequences and are increasingly proposed as world models, yet recent benchmarks reveal pronounced deficits in their physical reasoning. This raises the question of whether these models internalize physical principles or merely reproduce familiar motion patterns. We address this by probing internal representations of video Diffusion Transformers (DiTs) for simulator-derived ground-truth...

    arxiv.org/abs/2610.03154 · PDF

  15. 15

    Foresight: planning future perception in streaming VLMs without retraining

    Ashok Prasad Neupane, Dipan Bartaula, Ankit Belbase, Saugat Adhikari, Samip Ghimire, Saroj Poudel, Binod Bhattarai,...

    cs.CV · cs.AI

    Existing streaming vision-language models (VLMs) continuously perceive and reason over visual streams, but their computational pathways remain fixed throughout inference. Consequently, they cannot adapt computation to evolving scene dynamics, where different future events demand different levels and forms of perception. We show that streaming VLMs inherently possess the ability to anticipate the immediate future, and leverage this capability...

    arxiv.org/abs/2610.03123 · PDF

  16. 16

    Beyond Single Videos: Benchmarking and Active Evidence Seeking for E-Commerce Cross-Video Reasoning

    Jinghan Zhao, Yiman Hu, Liang Wu, Jian Xu, Bo Zheng

    cs.CV · cs.AI

    E-commerce videos are information-dense and frequently compared by consumers evaluating products and merchants assessing marketing strategies. However, existing multimodal models mainly focus on single-video understanding and have limited ability to compare information across videos. We introduce AdsCVR, the first e-commerce cross-video reasoning benchmark, containing 2,483 videos and 6,110 question-answer pairs across six reasoning...

    arxiv.org/abs/2610.03099 · PDF

  17. 17

    NegT2IBench: When Negation Changes the Picture. A Polarity Benchmark for Text-to-Image Models

    Omar Elfatairy, Maria A. Bravo, Jessica Bader, Zeynep Akata

    cs.CV · cs.AI · cs.LG

    Text-to-image (T2I) models are judged by benchmarks that measure whether requested content appears, but these benchmarks largely overlook the complementary ability to satisfy negated constraints, for example, generating "a non-red cup." Measuring negation raises challenges not faced by affirmation-based benchmarks and requires careful prompt and evaluation design. We introduce NegT2IBench, a benchmark of 4,800 prompts covering two attribute...

    arxiv.org/abs/2610.03084 · PDF

  18. 18

    OmniAct3D: Leveraging Foundation Geometry and Evidence-Grounded Reasoning for Panoramic 3D Detection

    Runtong Wu, Fei Teng, Di Wen, Guoqiang Zhao, Kunyu Peng, Kailun Yang

    cs.CV · cs.AI · cs.RO

    Accurate 3D detection is essential for mobile embodied agents, while Vision Foundation Models (VFMs) offer transferable visual and geometric priors. Yet existing VFM-based 3D detectors rely on narrow-view monocular images or discrete perspective views, limiting coherent surround perception; equirectangular projection (ERP) instead encodes a continuous 360 scene in a single image. Direct transfer remains difficult because ERP organizes...

    arxiv.org/abs/2610.03015 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.