cs.CV · 2026-09-07 · No. 108

Computer Vision and Pattern Recognition, 2026-09-07.

16 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

16 entries
  1. 01

    UniMate: One Unified Model to Animate Diverse Skeletons

    Linzhan Mou, Jiahui Lei, Zhiyang Dou, Chenyue Cai, Chaoyue Song, Adam Finkelstein, Szymon Rusinkiewicz

    cs.CV · cs.GR · cs.LG

    Recent advances in automatic rigging now deliver animation-ready 3D assets at scale, yet generating the motion to drive them remains a bottleneck. Existing learned animators are topology-constrained: they rely on category-specific templates or require per-skeleton fine-tuning and reference motions at inference. We present UniMate, a unified foundation model that synthesizes articulated motion for arbitrary skeletons from a rigged 3D asset and...

    arxiv.org/abs/2609.05415 · PDF

  2. 02

    Reflection-aware Generative Novel View Synthesis

    GeonU Kim, Shin Dong-Yeon, Tae-Hyun Oh

    cs.CV · cs.AI

    We propose Ref-GeNVS, a training-free, reflection-aware method for generative novel view synthesis (NVS) in mirror scenes. Existing multi-view diffusion models often fail to recognize the mirror in the scene and cannot exploit reflected content for scene generation. To fix this issue without additional training, our key idea is to treat a mirror image as two complementary views. From input images, we estimate the mirror plane and reflect...

    arxiv.org/abs/2609.05382 · PDF

  3. 03

    Lightweight Vision Transformer Compression for On-Device Plant Disease Detection in Resource-Constrained Agricultural Field Conditions

    Mahadev Sunil Kumar, Bhavika Gondi, Desaisetty Venkata Satya Sai Swapnith, Gangireddy Rahul Jogi, Sudheesh Manalil,...

    cs.CV · cs.AI · cs.LG

    Chilli (Capsicum annuum) is one of India's most economically significant crops, yet its productivity is persistently threatened by diseases that are difficult to identify without expert intervention. While Vision Transformers (ViTs) have achieved high classification accuracy, their large computational footprint makes deployment on resource constrained devices challenging. Existing compression approaches typically address pruning,...

    arxiv.org/abs/2609.05334 · PDF

  4. 04

    Adaptive Gated Deepfake Detection for Low-Resolution and Resource-Constrained Environments

    Vaishnavi Sen, Cody Laurie, Rashida Hasan

    cs.CV · cs.LG

    Deepfake detection models often rely on high-quality inputs, fixed inference paths, and computationally expensive architectures, limiting their use in low-resolution and resource-constrained settings. This paper proposes AdaGate-DF, an adaptive gated deepfake detection framework that uses image-quality cues to route samples through a dual multi-exit system so high-quality images can exit earlier and save compute. We evaluated AdaGate-DF...

    arxiv.org/abs/2609.05320 · PDF

  5. 05

    SMILE: Self-Explainable Multimodal Information Bottleneck for Medical Diagnosis

    Yuqing Yang, Alexander Schmatz, Zhaozhao Ma, Changkyu Choi, Robert Jenssen, Shujian Yu

    cs.CV · cs.LG

    Explainability is increasingly seen as a crucial requirement in AI-based medical diagnosis, particularly in safety-critical clinical decision-making. Most existing explainability methods in healthcare operate in a post-hoc manner and are predominantly designed for unimodal data, which limits their applicability in increasingly prevalent multimodal diagnostic settings. This paper addresses the problem of self-explainable multimodal diagnosis...

    arxiv.org/abs/2609.05174 · PDF

  6. 06

    Adaptive Multi-Granularity Temporal Modeling for Weakly Supervised Video Anomaly Detection

    Changyi Li, Yu Xiao

    cs.CV · cs.AI

    As the scale of video surveillance data outpaces manual annotation capacities, weakly supervised video anomaly detection (WSVAD) has emerged as a critical research frontier. Most existing approaches formulate WSVAD within a Multiple Instance Learning (MIL) framework that relies on rigid, hand-crafted temporal priors to supervise anomaly scoring. However, such formulations exhibit limited adaptability to the wide variation in anomaly durations...

    arxiv.org/abs/2609.05066 · PDF

  7. 07

    VICAL: Vicinal Consistency Alignment for Long-Tailed Visual Recognition

    Jiangang Zhu, Zheng Wang, Bin Zhu, Yi-Ping Phoebe Chen, Jingjing Chen

    cs.CV · cs.AI

    Multi-expert models have become the dominant paradigm for long-tailed learning, largely attributed to their presumed ability to benefit from expert diversity. However, we revisit this central assumption and reveal that diversity induced by logit adjustment or explicit regularizers does not guarantee better ensemble accuracy. Our work suggests that multi-expert models benefit more from variance reduction than diversity maximization. We...

    arxiv.org/abs/2609.04948 · PDF

  8. 08

    MCPO: Modality-Contrastive Preference Optimization for Multimodal Chain-of-Thought Compression

    Guangheng Yang, Zhenliang Ni, Zhenkai Wu, Han Shu, Juan Feng, Wenming Yang, Jie Hu

    cs.CV · cs.AI

    Recently, multimodal large-scale reasoning models have demonstrated remarkable capabilities in solving complex tasks through long Chains-of-Thought (M-CoT). However, excessively long reasoning trajectories incur substantial computational costs and significant KV-cache pressure. Existing CoT compression and alignment paradigms mainly rely on static rules or single-dimensional preferences, lacking fine-grained cross-modal constraints; as a...

    arxiv.org/abs/2609.04947 · PDF

  9. 09

    One Diffusion Model, Two Roles: Guided Trajectory Planning and Safety-Critical Scenario Generation in Closed-Loop Simulation

    Arka Pal, Rajesh Kumar, Hannes Eriksson, Rémi Lacombe, Arvid Laveno Ling, Ankit Gupta, Maciej Wozniak

    cs.CV · cs.AI · cs.LG · cs.RO

    Diffusion probabilistic models can capture the multi-modal, interaction-rich distribution of joint future trajectories in driving scenes. We show that a single pretrained diffusion traffic model can serve two complementary roles in the autonomous driving development loop: as an ego motion planner, and as a controllable generator of safety-critical scenarios for stress-testing the planners. On the planning side, we introduce a Single-Stream...

    arxiv.org/abs/2609.04921 · PDF

  10. 10

    Methane Detection On Board Satellites from Unorthorectified Imagery

    Luca Marini, Maggie Chen, Hala Lamdouar, Laura Martínez-Ferrer, Dr C. P. Bridges, Giacomo Acciarini

    cs.CV · cs.AI · cs.LG

    As a potent greenhouse gas, methane is a major driver of climate change. Its effective mitigation relies on timely detection. Conventional detection methods rely on orthorectification to correct geometric distortions and matched filters to enhance plume signals, which are steps designed for ground processing and poorly suited to onboard execution. We introduce UnorthoDOS, a dataset and approach for training machine learning models directly on...

    arxiv.org/abs/2609.04906 · PDF

  11. 11

    Sound-based Multi-Person 3D Pose Estimation

    Yusuke Oumi, Yuto Shibata, Go Irie, Akisato Kimura, Yoshimitsu Aoki, Mariko Isogawa

    cs.CV · cs.AI · cs.LG · cs.RO · cs.SD

    Can we recover the 3D poses of multiple people using only sound? This paper presents the first attempt to estimate multi-person 3D poses solely from acoustic signals. Estimating the poses of multiple individuals using acoustic signals is inherently challenging due to the superposition of motion-dependent signal variations. Unlike single-person scenarios, the presence of multiple subjects leads to overlapping acoustic signatures, making it...

    arxiv.org/abs/2609.04902 · PDF

  12. 12

    SimFuse3D: Source-Guided Target Simulation and Confidence-Guided Multi-Stage Localization Reweighting for Cross-Platform 3D Object Detection

    Yongchun Lin, Xinliang Zhang, Yun Zou, Zhixuan Xiao, Liang Lei, Jianya Guo, Yuqiang Zhai, Xiaofeng Wang, HaiKuo Xu, Haoang Li

    cs.CV · cs.AI

    Changes in sensor height and viewpoint alter object-level point distributions, making cross-platform LiDAR unsupervised domain adaptation (UDA) difficult. Self-training uses labeled source scans and unlabeled target scans, yet a retained prediction may provide a useful target location while enclosing sparse foreground returns, background clutter, or points inconsistent with the predicted box. We refer to this mismatch as box-point...

    arxiv.org/abs/2609.04886 · PDF

  13. 13

    Mitigating Performance Discrepancy in Cross-Domain 3D Class-Incremental Learning

    Jinge Ma, Gautham Vinod, Bruce Coburn, Jui-Feng Chi, Siddeshwar Raghavan, Fengqing Zhu

    cs.CV · cs.AI

    3D perception plays a crucial role in real-world applications such as autonomous driving, robotics, and AR/VR. In practical scenarios, 3D perception models need to continually adapt to newly emerging 3D object categories, making class-incremental learning (CIL) particularly important. However, unlike 2D images, 3D point clouds are inherently heterogeneous: objects from the same class may not only come from the clean CAD domain, but also from...

    arxiv.org/abs/2609.04860 · PDF

  14. 14

    Linguistic Trajectory Encoding for Efficient Long-Horizon Spatial Memory in Embodied Agents

    Tianyidan Xie, Shenyi Wang, Qiang Tang, Mingjie Wang, Zhicheng Qiu, Xuanfu Li, Zhan Xu, Jian Yang, Lanjun Wang, Zili Yi

    cs.CV · cs.AI

    Embodied agents performing long-horizon tasks require a memory representation in which the state transitions of dynamic objects remain queryable in natural language across hours-to-days observation horizons. Existing systems either drop fine-grained motion (clip-level video-language embeddings), keep it only as raw coordinates (geometric SLAM), or organise it around immediate task context (agent working memories). None of them gives the agent...

    arxiv.org/abs/2609.04802 · PDF

  15. 15

    LookThere! Sparse Vision by Reinforced Selection

    Sreehari Rammohan, Yousef Yassin, Anthony Fuller, Junfeng Wen, Carl Vondrick, Evan Shelhamer

    cs.CV · cs.LG

    Vision transformers typically treat every image token as equally important, yet for most tasks in computer vision only a fraction are needed. Adaptive computation methods accelerate inference by choosing which tokens to process, but existing methods struggle at extreme sparsity and require heuristics that may not generalize like token diversity and attention scores. We address these limitations with LookThere, achieving a new pareto frontier...

    arxiv.org/abs/2609.04698 · PDF

  16. 16

    Enhancing Multimodal Emotion Recognition via Multi-Feature Encoding and Attention-Based Fusion

    Xu Lin, Ke Wang, Hui Kang, Xinying Wang

    cs.CV · cs.AI · cs.LG

    Multimodal emotion recognition has attracted growing interest due to its importance in human-computer interaction, remote education, and healthcare. This paper proposes a novel multimodal emotion recognition framework that integrates rich audio and visual feature extraction with an attention-based fusion strategy. For audio, we extract three complementary feature types: semantic embeddings from Wav2Vec2, MFCC features, and statistical...

    arxiv.org/abs/2609.04690 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.