cs.CV · 2026-09-25 · No. 124

Computer Vision and Pattern Recognition, 2026-09-25.

20 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

20 entries
  1. 01

    TrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations

    Ayush Jain, Sreeharsha Paruchuri, Ishita Gupta, Fan Zhang, Tanner Schmidt, Jakob Engel, Katerina Fragkiadaki, Adam W. Harley

    cs.CV · cs.AI · cs.RO

    Existing point tracking models face a fundamental tradeoff: they can either track a sparse set of query points over long horizons, or track all points across only short clips. We introduce TrackEverything, a 3D point tracker that breaks this trade-off by representing videos as persistent 3D scene tracks in world coordinates. Grounded in the insight that videos are 2D projections of an underlying 3D world, TrackEverything decouples model...

    arxiv.org/abs/2609.30222 · PDF

  2. 02

    The Alignment Illusion in Multimodal Large Language Models

    Hong-Han Wang, Yuntao Wang, Hu Ding

    cs.CV · cs.LG

    Layer-wise visual-text similarity in Multimodal Large Language Models (MLLMs) is widely interpreted as evidence that the language model progressively integrates visual content into a shared representation space. This reading rests on the assumption that scalar alignment scores reflect content-level cross-modal interaction. To test this assumption, we apply controlled interventions to the visual stream. Across 13 MLLMs from five families...

    arxiv.org/abs/2609.30210 · PDF

  3. 03

    Accelerating Video Diffusion via Training-Free Trajectory Routing

    Mustafa Munir, Huy Vu, Shreyas Misra, Rohit Jena, Sajad Norouzi, Ali Taghibakhshi, Anis Ahmad, Anjul Patney, Pavlo...

    cs.CV · cs.AI

    Video diffusion is computationally expensive, as it requires executing a large model across many denoising steps. Even with step-distillation, inference remains expensive because every distilled step still requires a costly model evaluation. We present TRACK: TRajectory-Aware Capacity routing via top-K selection, a heterogeneous denoising strategy that switches between compatible large and small models at selected steps, reducing the average...

    arxiv.org/abs/2609.30096 · PDF

  4. 04

    AERIAL: Adversarial Evaluation of Robustness in Accuracy-Preserving Low-Precision EEG Decoders

    Saim Rehman, Muhammad Shafique

    cs.CV · cs.CR · cs.LG · eess.SP

    Deployment-oriented compression is attractive for resource-constrained brain--computer interfaces (BCIs), but whether it changes adversarial vulnerability remains unclear. On BCI Competition IV-2a, we compare 32-bit floating-point (FP32) EEGNet and ShallowConvNet models with global magnitude pruning and simulated INT8 post training quantization (PTQ) and quantization-aware training (QAT) across nine subjects and three seeds. Simulation...

    arxiv.org/abs/2609.30037 · PDF

  5. 05

    GHOST-Q: Towards Studying Grounding Hallucinations Overlooked Under Same-score TradeOffs in Quantized VLMS

    Saim Rehman, Muhammad Shafique

    cs.CV · cs.AI · cs.LG · cs.MM

    Post-training quantization of vision--language models (VLMs) is typically assessed through aggregate task accuracy and memory savings, but preserving a headline score does not guarantee preservation of visual grounding behavior. We present GHOST-Q, a cross-precision controlled evaluation of three 8B VLM families under FP16, INT8, and NF4 across utility and hallucination-sensitive benchmarks. Rather than comparing only aggregate accuracy, we...

    arxiv.org/abs/2609.29999 · PDF

  6. 06

    ADATEX4D: adaptive texture capacity allocation for 4D gaussian splatting

    De Jiang, Peiqiang Wang, Kehong Yuan, Shaohua Ma

    cs.CV · cs.AI

    Textured Gaussians improve local appearance capacity, but assigning the same texture resolution to every primitive wastes storage on low-detail or weakly visible regions. We introduce AdaTex4D, an adaptive texture-capacity module for deformation-based 4D Gaussian Splatting. Each Gaussian carries packed RGBA triplanes whose two axes grow independently according to visibility normalized screen-space gradients and deformed local scales....

    arxiv.org/abs/2609.29963 · PDF

  7. 07

    Not All Confusion Is Equal: A Source-Aware Uncertainty Diagnosis for Fine-Grained Aircraft Detection

    Hai Huang, Helmut Mayer

    cs.CV · cs.LG

    Fine-grained object detectors are commonly evaluated with confusion matrices, which show where the model is confused but not why, nor whether the confusion can be reduced. We argue that confusion can be attributed to distinct, separable sources, each quantitatively measurable, turning a passive measurement into actionable guidance. We present $A^2E^2$, a diagnostic tool that decomposes the sources of confusion along two axes, $\{$aleatoric,...

    arxiv.org/abs/2609.29959 · PDF

  8. 08

    Mind What Matters for Reasoning: Aligning Cross-Modal Attention via Selective Probability Mass Concentration

    Jiaqi Deng, Zonghan Wu, Zhan Heng, Xiaoshui Huang, Huan Huo, Guandong Xu

    cs.CV · cs.AI

    Multimodal large language models (MLLMs) achieve strong performance on visual reasoning tasks, yet remain prone to hallucinations and over-reliance on language priors, often generating answers without adequately using task-relevant visual evidence. Existing approaches primarily improve reasoning through reasoning-oriented supervision or inference-time strategies. In this work, we study a complementary question: can multimodal reasoning be...

    arxiv.org/abs/2609.29940 · PDF

  9. 09

    Efficient Continuous DEM Reconstruction under Limited Target-Resolution Supervision

    Zekai Shi, Meng Zhang, Haokun Zhang, Bo Zhang

    cs.CV · cs.LG · eess.IV

    High-resolution digital elevation models (DEMs) support Earth observation applications, but paired training references are often available only at coarser output resolutions. Reconstructing finer terrain grids therefore requires both effective transfer beyond the supervised scale and control of dense-query computation. To address this problem, SCOPE learns a continuous terrain representation from coarser-resolution pairs. It predicts a latent...

    arxiv.org/abs/2609.29864 · PDF

  10. 10

    S2Planner: Multi-Scale Semantic Planner for End-to-End Autonomous Driving

    Zhaowei Lu, Liguo Zhou, Yujie Guo, Lei Yu, Alois Knoll

    cs.CV · cs.AI

    We present S2Planner, a trajectory planner that combines three front-facing cameras with ego-motion history and the current driving command. A fine-tuned DINOv3 backbone and a Spatial Tuning Adapter produce multi-scale image features; a coarse-to-fine decoder then uses trajectory self-attention and camera-projected cross-attention to refine candidate waypoints. The contribution is the integration of ego-conditioned trajectory initialization...

    arxiv.org/abs/2609.29813 · PDF

  11. 11

    AgriCountDINO: Parameter-Efficient Exemplar-Guided Counting and Localization in Agriculture

    Shengjie Guo, Xin Li, Borjana Arsova, Hanno Scharr, Silvio Salvi

    cs.CV · cs.AI

    Accurate counting and localization of plants and their organs support phenotyping and yield estimation, yet target appearance, scale, and density vary widely across species and imaging conditions. Exemplar boxes specify the target without category-specific retraining, and point predictions identify the individual instances contributing to the count. We introduce AgriCountDINO, a parameter-efficient exemplar-guided framework for joint counting...

    arxiv.org/abs/2609.29460 · PDF

  12. 12

    Frame-to-Panorama Localization and Context-Aware Sampling for Scene-Specific Ship Detection in a Smart Marina Testbed

    Ignat Romanov, Andreas Hadjipieris, Neofytos Dimitriou

    cs.CV · cs.AI · cs.CG

    Smart maritime infrastructures provide continuous access to heterogeneous sensing streams, enabling repeated experimentation, digital-twin development, and AI-based maritime services. However, sensing hardware alone is not sufficient for scene-specific model development: historical video streams must also be spatially indexed, contextualized, and reduced to informative subsets for annotation. This paper presents a frame-to-panorama...

    arxiv.org/abs/2609.29447 · PDF

  13. 13

    Detecting Glaucoma Across Multi-ethnic Myopic and Non-Myopic Populations Using an Uncertainty-Aware Vision Transformer: A Multicentre Model Development and Validation Study

    Raghavan Lavanya, Yangqin Feng, Ten Cheer Quek, Quan V. Hoang, Linda Yi-Chieh Poon, Jost B. Jonas, Ya Xing Wang,...

    cs.CV · cs.AI

    Background: Artificial intelligence (AI)-based glaucoma detection from colour fundus photographs (CFP) offers scalable screening, but performance may decline on external datasets because of differences in ground-truth definitions, populations, and coexisting conditions such as high myopia (HM). We developed and validated a Vision Transformer-based deep learning (DL) model for glaucoma detection across multi-ethnic cohorts with and without HM....

    arxiv.org/abs/2609.29433 · PDF

  14. 14

    Segment-Level Risk Discovery in Online Handwriting for Alzheimer's Disease Detection

    Changqing Gong, Huafeng Qin, Mounîm A. El-Yacoubi

    cs.CV · cs.AI

    Online handwriting provides a non-invasive and low-cost behavioral biomarker for Alzheimer's disease (AD) detection, as it reflects both cognitive planning and fine motor control. Existing handwriting-based AD detection methods usually rely on global trajectory features or whole-sample representations, which can be strongly affected by individual writing style, task-specific variation, and acquisition noise. In this paper, we propose...

    arxiv.org/abs/2609.29384 · PDF

  15. 15

    Domain Recentering and Confidence-Weighted Prior Calibration for Vision-Language Models

    Youngeun Seol, Jimin Shin, Heeseo Yoon, Uiwon Hwang

    cs.CV · cs.AI

    Vision-language models such as CLIP achieve strong zero-shot classification, yet under distribution shift, visual embeddings drift from fixed text embeddings. Training-free calibration avoids the per-sample optimization of prompt learning, but prior feature calibration gives each image the full bias of one hard cluster. We propose Domain Recentering with Confidence Calibration (DRC), a training-free method adapting CLIP from a set of...

    arxiv.org/abs/2609.29358 · PDF

  16. 16

    Learning a Flow to Self-Supervised Representations

    Yuling Jiao, Wensen Ma, Houduo Qi, Defeng Sun

    cs.CV · cs.LG · stat.ML

    Explicit geometric references offer a direct way to structure self-supervised representations. Existing adversarial distribution-matching formulations, however, require costly encoder-critic optimization. We introduce Flow-Based Distribution Matching (FBDM), a non-adversarial framework that learns this reference-directed geometry through spherical conditional velocity regression. An ETF-inspired reference allows its number of components K' to...

    arxiv.org/abs/2609.29350 · PDF

  17. 17

    Hyperbolic Multimodal Continual Learning: A Closest-Admissible Solution

    Jiahong Liu, Ming Shen, Xiaohao Liu, Rex Ying, Menglin Yang, Tat-Seng Chua, Irwin King

    cs.CV · cs.AI

    Existing continual-learning methods protect parameters, replayed examples, or Euclidean feature subspaces. When applied to hyperbolic multimodal models, they do not explicitly preserve the Lorentz geometry that jointly encodes within-modality similarity, cross-modal correspondence, and semantic hierarchy; sequential updates can therefore retain task scores while still distorting previously learned relations. We address this gap with...

    arxiv.org/abs/2609.29329 · PDF

  18. 18

    Deep learning of longitudinal visual fields predicts glaucoma progression rate and identifies fast progressors

    Taiabur Rahman, Siddiqur Rahman, Muhammad Moniruzzaman, Ummay Kawsar, Sayedatunnessa Ratna, Shadman Siddique, Rafsan...

    cs.CV · cs.AI

    Glaucoma is the leading cause of irreversible blindness, and timely identification of fast progressors is essential to prevent disability. Current practice estimates progression by ordinary least-squares regression of mean deviation (MD) on time, requiring 6--10 visual field (VF) tests over several years to obtain a reliable slope. We present GLAM (Glaucoma Longitudinal Analysis Model), a deep learning framework that ingests longitudinal...

    arxiv.org/abs/2609.29256 · PDF

  19. 19

    TOLA: Text-aware One-Step Latent Adaptation for Diffusion-based Text Image Super-Resolution

    Yike Xu, Yue Shi, Yong Guo, Jiezhang Cao

    cs.CV · cs.AI

    Text image super-resolution (TSR) aims to recover visually faithful and readable text under unknown degradations. Existing diffusion-based methods typically rely on multi-step prediction of either the high-resolution image or its text prior, resulting in prohibitive computational cost and inference latency. More critically, an erroneous text prior may be repeatedly injected into the denoising process, causing image and text predictions to...

    arxiv.org/abs/2609.29240 · PDF

  20. 20

    SARFusion: Scene-Aware Routing Fusion for Robust Camera-LiDAR 3D Object Detection

    Yuting Zhao, Ziyi Zheng, Shuxiao Li

    cs.CV · cs.AI

    Camera-LiDAR fusion has become a prevailing paradigm for 3D object detection in autonomous driving. However, existing fusion detectors often establish strong inter-modality dependencies by decoding object queries from tightly coupled multimodal representations. Under corrupted driving conditions, such dependencies make the detector vulnerable to unreliable modalities, where degraded observations may interfere with reliable modality-specific...

    arxiv.org/abs/2609.29235 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.