cs.CV · 2026-10-02 · No. 131

Computer Vision and Pattern Recognition, 2026-10-02.

22 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

22 entries
  1. 01

    One Basis to Animate Them All: Gaussian Blendshape Distillation for Real-Time Avatars

    Ramazan Fazylov, Stamatis Lefkimmiatis, Ivan Laptev

    cs.CV · cs.AI · cs.HC · cs.LG

    3D Gaussian avatars support fast rendering, however, their real-time animation is often challenged by the costly neural inference. We address this bottleneck and show that the animation of pretrained avatar models can be closely approximated by a linear combination of identity-independent blendshapes. Building on this finding, we introduce GALA (Gaussian Animation via Linear Approximation), a distillation method that replaces per-frame heavy...

    arxiv.org/abs/2610.02207 · PDF

  2. 02

    Embedding Prediction Helps Image Generation

    Sihan Xu, Ji Xie, Zilin Wang, Hui Shen, Stella X. Yu

    cs.CV · cs.LG

    In diffusion transformers, a class label or a text prompt is embedded once, and the same condition is reused at every denoising step. We ask whether predicted embeddings can serve as this condition instead. Next-Embedding Predictive Autoregression (NEPA) trains a Transformer to predict the next continuous embedding in a sequence. In generation, the clean image follows the noisy image, so its embeddings are the next embeddings after the...

    arxiv.org/abs/2610.02203 · PDF

  3. 03

    SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation

    Tianjiao Yu, Xinzhuo Li, Yifan Shen, Ying Shen, Kiet A. Nguyen, Adheesh Sunil Juvekar, Ismini Lourentzou

    cs.CV · cs.AI · cs.LG

    High-resolution 3D generation increasingly relies on voxel latents and multi-stage pipelines that first predict active structure and then synthesize local geometry. While effective, this design fragments continuous surfaces into many local tokens, inflates generation cost, and often weakens topological consistency for thin or highly connected shapes. We introduce SILSA, a topology-aware 3D generation framework that represents shapes with...

    arxiv.org/abs/2610.02201 · PDF

  4. 04

    DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation

    Zhengming Yu, Junkun Yuan, Haotian Yang, Gordon Guocheng Qian, Yizhi Wang, Angtian Wang, Yiding Yang, Bo Liu, Xin...

    cs.CV · cs.AI

    Distribution Matching Distillation (DMD) trains a few-step student from the difference between separately estimated target and student scores, so it must keep an auxiliary diffusion model fitted to the student's evolving distribution at extra memory and computation cost. We introduce DMAD, Distribution Matching as Adversarial Distillation, which recasts distribution matching as classification and learns the required log-density ratios...

    arxiv.org/abs/2610.02188 · PDF

  5. 05

    Generative Cinematographer: Composing Camera and Object Motion in 3D

    Jiahan Zhang, Chaohao Yang, Namitha Guruprasad, Vivekjyoti Banerjee, Trong-Tung Nguyen, Alan Yuille, Anand Bhattad

    cs.CV · cs.AI

    Current controllable video generation systems often rely on 2D motion trajectories or sparse drag signals for object motion. These controls are ambiguous because the same 2D trajectory can correspond to different 3D motions, especially when the camera and objects move simultaneously. We present Generative Cinematographer (GenCine), a system that lifts a single image into an editable 3D scene scaffold where artists jointly author camera and...

    arxiv.org/abs/2610.02180 · PDF

  6. 06

    MIRTO: a registration-gated, multiverse-tested evaluation protocol for unsupervised anomaly segmentation in brain MRI

    Negin Kafee Hernashki, Soumick Chatterjee

    cs.CV · cs.AI · eess.IV · physics.med-ph

    Unsupervised anomaly detection (UAD) methods for brain MRI are ranked by a single score, yet that score rests on choices that are rarely reported: how each anomaly map is aligned with the reference, how and on which data the threshold is set, and which false-positive budget, metric, aggregation and lesion definition are used. We present MIRTO, an evaluation protocol that makes these choices explicit and measures their effect. It gates the...

    arxiv.org/abs/2610.02136 · PDF

  7. 07

    Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes

    Sophia Sirko-Galouchenko, Monika Wysoczanska, Andrei Bursuc, Nicolas Thome, Spyros Gidaris

    cs.CV · cs.AI · cs.CL · cs.LG

    On-policy self-distillation has recently emerged as an effective approach for improving language-model reasoning by supervising students with a frozen or EMA version of themselves that receives privileged information. Its application to multimodal large language models (MLLMs), however, remains largely unexplored. Recent approaches use privileged visual information, such as image crops corresponding to a question, to improve fine-grained...

    arxiv.org/abs/2610.02117 · PDF

  8. 08

    GeoLatent: Geometry-Guided Latent Structuring with Routed Optimization for 3D Reasoning

    Yakun Zhu, Yi Bin, Yujuan Ding, Zheng Wang, Pengpeng Zeng, Duo Peng, Jingkuan Song, Heng Tao Shen

    cs.CV · cs.AI

    Despite progress in vision-language models, 3D spatial reasoning from 2D images remains challenging. Text-based methods describe intermediate geometry with discrete tokens, limiting fidelity for continuous spatial relations. Continuous latents offer richer representations, but a single latent type does not explicitly separate the cues needed across spatial tasks. Decomposed spatial latents address this by representing position, direction, and...

    arxiv.org/abs/2610.02091 · PDF

  9. 09

    Task-Adaptive Grounded 3D-Programmers Using 2D VLMs

    Arman Raayatsanati, Sombit Dey, Anna-Maria Halacheva, Jan-Nico Zaech, Luc Van Gool, Danda Pani Paudel

    cs.CV · cs.AI

    Recent vision-language models (VLMs) exhibit remarkable generalization and reasoning abilities, yet 3D understanding in these models is limited by data scale, training diversity, and reasoning capacity. Instead of naively extending these models into 3D, we take a different approach: we enable powerful 2D VLMs to operate reliably in 3D by introducing 3D grounding and iterative feedback loops with two novel concepts: Canonical Coordinate...

    arxiv.org/abs/2610.02021 · PDF

  10. 10

    Exploring Weaknesses of Generative Image Watermarks against Latent Frequency Masking

    Kirill Aistov, Khaled Abud, Irina Serzhenko, Egor Kovalev, Aleksey Yakushev, Aleksandr Akimenkov, Dmitry Obydenkov,...

    cs.CV · cs.AI · cs.MM

    Invisible watermarking has become a central tool for tracing AI-generated images, but its robustness against adaptive removal attacks remains an open security question. We introduce Latent Frequency Masking, an attack that erases watermark evidence by replacing selected Fourier coefficients in the latent representation of a watermarked image. The replacement can be sampled from Gaussian noise for efficiency or derived from diffusion...

    arxiv.org/abs/2610.02010 · PDF

  11. 11

    Weather-Aware Domain Adaptation for Street-View Weather Recognition

    Hossein Maghsoumi, George Atia, Yaser P. Fallah

    cs.CV · cs.LG

    Adverse conditions such as rain, snow, fog, and dust remain challenging for camera-based perception in autonomous driving. We study multi-class weather recognition from street-view images under domain shift, where most available training data come from non-street-view sources that differ markedly from real driving scenes. We propose Weather-Aware Adversarial Discriminative Domain Adaptation (WA-ADDA), which conditions the domain discriminator...

    arxiv.org/abs/2610.02000 · PDF

  12. 12

    Comparing a gradient boosting algorithm to the GOES FDC for wildfire detection

    Asaf Vanunu, Boaz Nadler, Arnon Karnieli

    cs.CV · cs.LG

    Wildfires pose severe risks to human life, ecosystems, and property. This study presents a machine learning approach for wildfire detection from GOES ABI imagery. A CatBoost model was trained on a large dataset with thousands of ABI images and over 300,000 matching VIIRS fire detections. An evaluation on a separate dataset across five regions showed that the learned CatBoost model outperformed the operational GOES Fire Detection and...

    arxiv.org/abs/2610.01994 · PDF

  13. 13

    MoLE: Mixture of Latent Experts for Complementary Visual Reasoning

    Yingcheng Liu, Tianyi Jiang, Yujuan Ding, jiangbo Ai, Xun Jiang, Guoqing Wang, Wei Ye, Yi Bin

    cs.CV · cs.AI · cs.CL

    Latent visual reasoning equips vision--language models with continuous intermediate states that can process visual evidence without explicit textual reasoning traces or repeated image operations. However, existing methods often allow multiple latent tokens to access the same visual evidence through shared value projections, providing no mechanism for them to extract complementary visual information; simply increasing the latent budget can...

    arxiv.org/abs/2610.01917 · PDF

  14. 14

    Unsupervised Domain Adaptation for Enhanced Radiometer Image Precipitation Estimation using Conditional Flow Matching

    Victor Enescu, Assaad Zeghina, Matthieu Meignin, Nicolas Viltard, Cécile Mallet

    cs.CV · cs.AI

    Deep generative networks have recently achieved unprecedented performance in precise image and video editing using sophisticated textual prompts. However, the effectiveness of such models heavily depends on access to very large supervised and annotated image datasets, which can be very difficult to obtain. This is particularly true for satellite instruments, which very rarely overlap with labelled data, and suffer from domain shifts in the...

    arxiv.org/abs/2610.01890 · PDF

  15. 15

    PhaseAT: Fourier Phase Adversarial Training for Medical Image Domain Generalization

    Ahmed Sharshar, Asif Hanif, Naveen Kumar Kummari, Mohammad Yaqub, Mohsen Guizan

    cs.CV · cs.LG

    Reliable clinical deployment of deep medical image models is hindered by distribution shifts across scanners, sites, and acquisition protocols. Existing domain generalization (DG) methods often focus on style or intensity diversification, but they can still leave networks dependent on domain-specific texture correlations. Inspired by evidence that Fourier phase encodes semantic structure, we introduce PhaseAT, a phase-aware adversarial...

    arxiv.org/abs/2610.01807 · PDF

  16. 16

    VETO: Video Efficient Token Optimization for Vision Language Models

    Gueter Josmy Faure, Hao Ping Wang, Min-Hung Chen, Winston H. Hsu

    cs.CV · cs.AI · cs.CL

    Processing long videos with Vision-Language Models (VLMs) is bottlenecked by the quadratic cost of visual tokens, making long-form inference prohibitively expensive. While single-axis compression methods mitigate this, they hit a hard efficiency floor because they treat spatial and temporal redundancy independently. We present VETO (Video Efficient Token Optimization for Vision-Language Models), a training-optional plug-in that eliminates...

    arxiv.org/abs/2610.01785 · PDF

  17. 17

    Cog-VADU: A Training-Free Cognitive Reasoning Framework for Video Anomaly Detection and Understanding

    Mohd Ubaid Wani, Sara Atito, Josef Kittler, Muhammad Awais

    cs.CV · cs.AI

    Video Anomaly Detection (VAD) aims to temporally localize abnormal events in videos. Most existing approaches rely on dataset-specific training and curated annotations, limiting generalization in open-set scenarios. Recent zero-shot methods based on Large Vision- Language Models (LVLMs) alleviate this dependency but often lack temporal continuity and structured reasoning. We propose Cog-VADU, a fully training-free framework that reformulates...

    arxiv.org/abs/2610.01754 · PDF

  18. 18

    Architectural Sampling: Test-Time Scaling via Computational Diversity in Frozen Vision-Language Models

    Akshit Singh, Shyam Marjit, Wei Lin, Leonid Karlinsky, M. Jehanzeb Mirza

    cs.CV · cs.AI

    Test-time scaling often seeks better answers by sampling multiple responses from a frozen model, yet conventional temperature sampling generates every candidate along the same fixed computation path. We introduce architectural sampling, a training-free method that generates candidates through distinct forward computations by reusing selected blocks of decoder layers. Varying the block location and repetition count introduces computational...

    arxiv.org/abs/2610.01687 · PDF

  19. 19

    Do MLLM Judges Judge the Edit? Auditing Bias in Image Editing Evaluation with Verified Quality Preservation

    Yuan Huang, Zirui Song, Xiuying Chen

    cs.CV · cs.LG

    Multimodal large language models (MLLMs) are increasingly used as automated judges for instruction-based image editing and as reward signals for model training. However, systematically auditing whether these judges are influenced by cues irrelevant to editing quality is challenging because visual interventions may themselves alter the quality being evaluated. A judgment shift can therefore be attributed to bias only when the intervention is...

    arxiv.org/abs/2610.01670 · PDF

  20. 20

    Not All Error Yields to Scale: Where Scaling Stops in Vision-Language Inference

    Xinye Zhao, Yunkai Dang, Yunchen Wu, Wenbin Li

    cs.CV · cs.AI

    Vision-language models (VLMs) face a fixed-budget trade-off between processing more visual information for fine-grained perception and using a larger language backbone for complex reasoning. Existing studies do not tell us which combination of backbone size and input resolution to deploy, especially in high-resolution deployments. To address this gap, we propose the Separable Law that describes how VLM performance changes with language...

    arxiv.org/abs/2610.01640 · PDF

  21. 21

    Hob-VL: A Benchmark for Visually Grounded Boolean Reasoning

    Yuzhou Wang, Emile Anand, Ijay Narang

    cs.CV · cs.AI · cs.LG · cs.LO

    Reliable visual reasoning requires composing multiple visual observations and returning consistent answers to logically equivalent questions. We introduce Hob-VL, a benchmark for visually grounded Boolean reasoning. Hob-VL comprises two tasks: (1) evaluating whether a Boolean rule holds in an image, and (2) identifying the (unique) object satisfying a Boolean description. Hob-VL contains 6,000 human-verified balanced Yes/No questions, each...

    arxiv.org/abs/2610.01605 · PDF

  22. 22

    Before It Fades: Reinforcing Temporal Representations at Inference Time in VideoLLMs

    Youngwoo Shin, Yusung Ro, Minseo Kim, Junmo Kim

    cs.CV · cs.AI

    Video Large Language Models (VideoLLMs) receive frames in sequential order and interpret how visual content evolves along the temporal axis, yet temporal reasoning remains a persistent weakness across architectures. Reversing the frame order of a video, a transformation that should invert temporal answers, often leaves the final prediction unchanged. We investigate where this failure originates by defining the temporal divergence vector...

    arxiv.org/abs/2610.01595 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.