cs.CV · 2026-09-22 · No. 121

Computer Vision and Pattern Recognition, 2026-09-22.

17 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

17 entries
  1. 01

    GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay

    Yiran Wang, Xingyilang Yin, Junfu Pu, Guangzhi Wang, Kaifeng Li, Mingyu Ouyang, Huiqiang Sun, Lingen Li, Cheng...

    cs.CV · cs.AI

    Modern video games provide a measurable testbed for AI models, combining abilities of visual understanding, instruction decomposition, goal planning, and precise action control over multiple temporal horizons. Existing datasets and benchmarks, however, either cover a narrow range of games, lack language instructions, or rely on high-variance online rollouts. To address these challenges, we introduce GameHorizon, a unified data and evaluation...

    arxiv.org/abs/2609.25001 · PDF

  2. 02

    WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory

    Wangbo Yu, Kunhao Liu, Wenbo Hu, Shenghai Yuan, Chaoran Feng, Haiyang Zhou, Yukun Huang, Yiran Wang, Wang Zhao,...

    cs.CV · cs.AI · cs.GR

    Video world models enable interactive exploration of dynamic environments, yet struggle to respect prior observations over long horizons and across viewpoints. We present WorldCrafter, a video world model that learns a camera-queryable implicit 3D-aware memory for this purpose. The key insight is to let the requested viewpoint shape how multi-view evidence is compressed into the video generator's limited token budget. Trained jointly with the...

    arxiv.org/abs/2609.24984 · PDF

  3. 03

    Mobile Imaging Solutions for Medical Diagnosis: Trends and Applications

    Syed Muhammad Ibne Zulfiker, Tanzima Hashem, Fariha Tabassum Islam, Md Sultanul Arifin, Khandker Aftarul Islam,...

    cs.CV · cs.AI · cs.LG

    Advances in processing power, camera technologies, and mobile image analysis have made smartphones and other mobile devices, such as laptops, increasingly suitable for medical diagnosis and healthcare applications. Researchers have developed low-cost solutions for the early detection and monitoring of various health conditions, including eye and ENT diseases, malnutrition, heart rate variability, skin and oral conditions, and injuries, using...

    arxiv.org/abs/2609.24814 · PDF

  4. 04

    PrismGPT: Proxy-Guided Learning for Region-Aware Photo Editing with Self-Synthesized Reasoning

    Ke Zhao, Hue Nguyen, Abhijith Punnappurath, Zhongling Wang, Iqbal Mohomed, Michael S. Brown

    cs.CV · cs.AI

    Professional photo finishing relies on both global adjustments and region-specific local edits guided by semantic masks, yet current automated methods handle this workflow only partially. We present PrismGPT, a Vision-Language Model (VLM) framework that produces structured, region-aware editing plans from a single input image without relying on commercial black-box tools. Training a VLM to simultaneously diagnose aesthetic deficiencies at...

    arxiv.org/abs/2609.24768 · PDF

  5. 05

    MiTHras: Task-specific Hierarchical Semi-supervised Contrastive Masked Autoencoder for Mitotic Figure Analysis

    Trinh T. L. Vuong, Simon Graham, Quoc Dang Vu, Phat T. H. Ho, Jeewoo Lim, Mostafa Jahanifar, Nasir Rajpoot, Jin T. Kwak

    cs.CV · cs.LG

    Mitotic figure (MF) analysis supports tumor grading and prognostic assessment, but automated models remain sensitive to differences in tissue type and image acquisition. We present MiTHras, a task-specific pretraining framework that combines pseudo-label-guided image- and token-level contrastive learning with masked reconstruction. We construct TCGA-MF-Pseudo, a corpus of 1.8 million cell-centered images from 14 TCGA cohorts spanning 11 organ...

    arxiv.org/abs/2609.24736 · PDF

  6. 06

    What Makes a Good Medical Image Tokenizer? Rethinking Reconstruction and Generation in Medical Image Tokenization

    Niklas Bubeck, Yundi Zhang, Vasiliki Sideri-Lampretsa, Julian McGinnis, Jiancheng Yang, Daniel Rueckert, Jiazhen Pan

    cs.CV · cs.AI

    Latent diffusion models now dominate medical image generation, and every such pipeline rests on a \emph{tokenizer} that compresses images into the latent codes for image generation to operate on. Thereby, the tokenizer choice bounds every downstream task from reconstruction fidelity and generation quality to the representations available for downstream analysis. Yet, medical imaging pipelines routinely utilize tokenizers from natural imaging...

    arxiv.org/abs/2609.24691 · PDF

  7. 07

    FedMust: Semi-supervised Multi-task Student-Teacher Federated Learning for Multi-organ CT Segmentation

    Ashkan Moradi, Bendik Skarre Abrahamsen, Mattijs Elschot

    cs.CV · cs.AI

    Multi-organ segmentation using deep learning requires large amounts of annotated patient data; however, institutions often lack sufficiently large and diverse annotated datasets. Privacy constraints further prevent institutions from sharing patient data to overcome this limitation. Moreover, due to the labor-intensive nature of annotation and the scarcity of diverse expertise, institutions typically have labels for only a small portion of...

    arxiv.org/abs/2609.24627 · PDF

  8. 08

    AgentSTAR: Agentic Shape Tracking and Reconstruction from Monocular Videos

    Kirill Mazur, Nikita Karaev, Matthew Chang, Jitendra Malik, Nur Muhammad "Mahi'' Shafiullah

    cs.CV · cs.AI

    In this work, we present a method for shape reconstruction and tracking from video via agentic analysis-by-synthesis. Unlike prior methods which first estimate dense pixel correspondences and then recover object motion from them, our method infers a structured 3D object model, including its geometry and kinematic structure, and uses this model to optimise object track estimates over time. In our optimisation loop, a Vision-Language Model...

    arxiv.org/abs/2609.24487 · PDF

  9. 09

    VPRune: Efficient Training-free Pre-LLM Visual Token Pruning

    Guangchuan Lv, Dianxing Shi, Dingjie FU

    cs.CV · cs.AI

    Visual token pruning is a promising approach to reducing the inference cost of large vision-language models (LVLMs), yet aggressive token reduction often causes substantial performance degradation. We identify three key factors behind this degradation: text-guided selection bias, information loss from discarded tokens, and positional distortion caused by sequence compaction. Based on these observations, we propose \textbf{VPRune}, a...

    arxiv.org/abs/2609.24485 · PDF

  10. 10

    MECAIL: Communication-Aware Incremental Learning for Object Detection with 14.6 KB Spatiotemporal Experts

    Matthias Neuwirth-Trapp, Maarten Bieshaar, Danda Paudel, Konrad Schindler, Luc Van Gool, Christos Sakaridis

    cs.CV · cs.LG

    Intelligent transportation systems require Incremental Learning (IL) to continually improve their overall performance in dynamic environments. However, most edge devices lack the computational resources to support on-device IL, requiring updates to be transmitted from centralized servers. We propose using this setup to obtain dense, specialized module coverage that adapts a fixed base model to specific spatiotemporal contexts, such as parking...

    arxiv.org/abs/2609.24455 · PDF

  11. 11

    Do LiDAR Language Models Really Understand Spatio-temporal Relationships?

    Runyi Yang, Murat Akkoyun, Di Wen, Ruiping Liu, Yufan Chen, Junwei Zheng, Xiaoye Wang, Kailun Yang, Danda Pani...

    cs.CV · cs.AI · cs.RO

    Recent 4D LiDAR language models aim to reason about objects and their evolving spatial relationships. Yet, in our evaluation, always selecting the same option nearly matches the multiple-choice accuracy of two B4DL-derived configurations. We introduce LiDAR-Hallu, a geometry-referenced benchmark and diagnostic protocol with 10,000 questions across 150 nuScenes scenes. It covers object existence, ego-relative position, distance ordering,...

    arxiv.org/abs/2609.24452 · PDF

  12. 12

    Estimating Accurate Hand Pose in Camera Space with Vision Transformer

    Kaiwen Ren, Yiran Jiang, Yongjing Ye, Shihong Xia

    cs.CV · cs.AI · cs.GR

    Monocular RGB-based hand pose estimation has emerged as a critical research frontier in computer vision. The local hand pose estimation methods predict hand poses relative to the wrist, while global hand pose estimation also requires estimating the wrist's position in the camera coordinate system. However, this camera-space estimation confronts two fundamental challenges: (1) depth ambiguity in monocular settings, and (2) the coupling effect...

    arxiv.org/abs/2609.24424 · PDF

  13. 13

    A Lightweight Convolutional Neural Network for Real-Time Recognition of Hand-Drawn Geometric Shapes

    Shahir Abdullah

    cs.CV · cs.LG

    Recognizing hand-drawn geometric shapes is a foundational sub-problem of sketch recognition, with applications in education, human-computer interaction, and diagram digitization. This paper presents the design, implementation, and evaluation of a desktop application that recognizes four basic hand-drawn geometric shapes, circle, square, rectangle, and triangle using a compact Convolutional Neural Network (CNN). A dataset of 2,000 labeled...

    arxiv.org/abs/2609.24384 · PDF

  14. 14

    Topographic Training Concentrates Causal Circuits Without Improving Neuron Monosemanticity

    Gautam Ranka, Shubham Santosh Pandere, Aiden Dsouza

    cs.CV · cs.LG

    Mechanistic interpretability of vision transformers seeks to decompose model computation into human-readable units, but learned representations entangle many concepts in each neuron. Feature superposition is widely treated as the central obstacle to this decomposition, yet most mitigations (sparse autoencoders, dictionary learning) are post-hoc and leave the underlying network unchanged. We ask whether a spatial-locality training loss...

    arxiv.org/abs/2609.24379 · PDF

  15. 15

    Dissecting Agentic Forensics: The Role of Triage, Prompting, and Evidence Arbitration in Open-World Fake Image Detection

    Xianlong Li, Pietro Bongini, Niccoló Pancino, Marco Blanchini, Benedetta Tondi, Mauro Barni

    cs.CV · cs.AI · cs.CR

    Image forensics is increasingly an open-world problem: manipulations range from fully synthetic images to localized edits, splicing and swapping, while most forensic detectors remain specialized to a single manipulation family. Agentic AI has recently emerged as a promising solution. In principle, such systems can assess the reliability of individual detectors, identify out-of-scope evidence, and arbitrate conflicting reports. However, it...

    arxiv.org/abs/2609.24359 · PDF

  16. 16

    STAR: Scene- and Task-Aware 4D Radar Preprocessing Towards End-to-End Cognitive Radar

    Seung-Hyun Song, Dong-Hee Paek, Seung-Hyun Kong

    cs.CV · cs.AI

    Four-dimensional (4D) Radar has emerged as a key sensor for environmental perception, providing range, azimuth, elevation, and Doppler measurements while remaining robust to illumination changes and adverse weather conditions. However, conventional Radar preprocessing methods, such as constant false alarm rate (CFAR) detection, select measurements primarily based on signal-level criteria and may therefore discard information valuable for...

    arxiv.org/abs/2609.24151 · PDF

  17. 17

    Action-Slot: Structured Action-Centric Representation Learning for Multi-Agent Atomic Activity Understanding

    Yu-Ho Chang, Chi-Hsi Kung, Yi-Hsuan Tsai, Yi-Ting Chen

    cs.CV · cs.AI

    Atomic activity understanding aims to recognize and localize structured traffic behaviors that jointly encode motion patterns and their grounding in road topology. Unlike conventional action recognition, atomic activities are multi-agent, multi-label, and topology-aware: multiple activities co-occur while many agents remain inactive. We introduce Action-Slot, a structured action-centric representation learning framework. Slot attention is...

    arxiv.org/abs/2609.24127 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.