cs.CV · 2026-08-24 · No. 94

Computer Vision and Pattern Recognition, 2026-08-24.

22 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

22 entries
  1. 01

    Re$^3$Cap: Retrieval-Guided Refinement for Image Captioning Enhancement via Reinforcement Learning

    Haonan Jia, Shichao Dong, Zenghui Sun, Jiawen Zheng, Ziqi Miao, Gege Shi, Qiuyu Zhao, Jinsong Lan, Xiaoyong Zhu, Bo Zheng

    cs.CV · cs.AI

    Reinforcement Learning (RL) has demonstrated significant gains in image captioning, yet it is still limited in encouraging Large Vision-Language Models (LVLMs) to explore novel reasoning strategies. This limitation leads to a performance gap between RL and Supervised Fine-Tuning (SFT). In this paper, we argue that multi-modal retrieval can serve as an effective reasoning signal for caption refinement. Based on this insight, we present the...

    arxiv.org/abs/2608.21305 · PDF

  2. 02

    On the Transferability of Agricultural Weed Detection Under Cross-Field Distribution Shift

    Nikhilesh Prabhakar, Pranuthi Tenali, Wilfredo Abudeye Fernandez, Shekhar Borah, Athresh Karanam, Erik Blasch,...

    cs.CV · cs.LG

    Accurate agricultural weed detection in real-world field conditions is essential for precision agriculture, enabling targeted intervention and reducing yield loss. Recent work has reported strong detection performance from UAV-based imagery across a range of crops, yet existing approaches evaluate within a single crop and field, leaving practitioners with little evidence that a model trained on one crop will generalize to a new field or crop...

    arxiv.org/abs/2608.21254 · PDF

  3. 03

    Towards Investigating Residual Hearing Loss: Quantification of Fibrosis in a Novel Cochlear OCT Dataset

    Julia Dietlmeier, Benjamin Greenberg, Wenxuan He, Teresa Wilson, Rubing Xing, Jordan Hill, Adrienne Fettig, Madeline...

    cs.CV · cs.AI

    Objective: Cochlear implants (CIs) are bionic prostheses that restores hearing via electrical stimulation of the auditory nerve. Hybrid CIs, which use electroacoustic stimulation (EAS), combine residual low-frequency acoustic hearing with CI electrical stimulation. Intracochlear fibrosis, which forms in response to the presence of the implant, may impede residual hearing function and gradually reduce the efficacy of EAS. It is therefore a...

    arxiv.org/abs/2608.21189 · PDF

  4. 04

    Is Visual Prompting All You Need? Studying VLM Spatial Reasoning under Progressive Visual Scaffolds

    Lars Benedikt Kaesberg, Tianyu Yang, Florian Valentin Wunderlich, Terry Ruas, Jan Philip Wahle, Daniel Kurzawe, Bela Gipp

    cs.CV · cs.AI

    Vision-language models (VLMs) have advanced rapidly in multimodal reasoning, yet recent work shows that their failures often reflect an interaction between visual grounding and downstream reasoning. What remains less clear is how the visual presentation of a task shapes model performance and failure modes when the underlying reasoning problem is unchanged. We study this question in SPaRC, a benchmark for grid-based visual spatial planning, by...

    arxiv.org/abs/2608.21170 · PDF

  5. 05

    Human-JEPA: A Human-Centric Vision Model that Perceives and Anticipates

    Hui Wei, Licai Sun, Guoying Zhao

    cs.CV · cs.LG

    Machines that understand humans should perceive the present and anticipate the future. Existing human-centric vision model are pretrained on human images, set the state of the art in static dense perception, so motion and anticipation are out of reach. Here we present Human-JEPA, a human-centric vision model trained on video by anchored forecasting: dense targets are pinned to a frozen copy of the initialization, preventing a silent collapse...

    arxiv.org/abs/2608.21160 · PDF

  6. 06

    A Modular Agent for Reliable and Auditable Spatial Relation Verification in CT Scans

    Simon Vincent Abel, Heiko Hillenhagen, Michael Götz, Timo Ropinski, Ayhan Can Erdur, Daniel Santak Wolf

    cs.CV · cs.AI

    Reliable spatial understanding is an important prerequisite for future medical vision-language systems that aim to support radiological report generation and structured image understanding. While modern vision-language models (VLMs) show promising performance on many medical imaging tasks, recent evidence suggests they remain weak in controlled spatial reasoning and often fail to reliably ground spatial relations in image evidence. Given that...

    arxiv.org/abs/2608.21140 · PDF

  7. 07

    Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs

    Luka Ribar, Jeevan Bhoot, Douglas Orr

    cs.CV · cs.LG

    Deploying vision-language models (VLMs) on mobile devices is challenging due to their significant memory and compute requirements. We present a framework for quantizing VLMs for efficient inference on resource-constrained hardware. Our approach combines a quantization pipeline that uses the model itself to generate training data and does not require access to the training setup, with a novel 2.7-bit-per-parameter format supporting efficient...

    arxiv.org/abs/2608.21134 · PDF

  8. 08

    CIVA: Critic-Induced Value-Subspace Attacks on Visual World-Model Agents

    Jiancheng Wang, Mingli Zhu, Tong Zhang, Jiaqi Ruan, Wei Wang, Siyuan Liang, Dacheng Tao

    cs.CV · cs.AI

    Visual world-model agents such as DreamerV3 act through a recurrent latent state rather than a single observation, which weakens frame-wise observation attacks and makes their perturbations vary sharply over time under a strict per-frame perturbation constraint. We study white-box, causal, online attacks on such agents and propose Critic-Induced Value-Subspace Attacks (\textbf{CIVA}). Our key observation is that, along a rollout,...

    arxiv.org/abs/2608.21114 · PDF

  9. 09

    A2DINOv3: Rethinking Multi-Modal Object Detection via Socialized Collaboration

    Jiekang Feng, Zhihe Fan, Yunqi Zhu, Xinjie Yao, Yueying Zhang, Yike Gao, Ranxin Li, Guanzuo Chen

    cs.CV · cs.AI

    Multi-modal object detection is essential for robust scene understanding in challenging conditions, including low-light and adverse environments. Recent vision foundation models (e.g., DINOv3) have exhibited strong representation capabilities, yet adapting them to multi-modal scenarios remains challenging. Existing dense cross-modal fusion strategies often force heterogeneous modalities to interact indiscriminately, which may introduce...

    arxiv.org/abs/2608.21099 · PDF

  10. 10

    AT-ViT: Area-Targeted Multi-View Vision Transformer with Cross-Attention and Multi-Scale Patching for Plant Trait Recognition in Herbarium Images

    Amani Sedrat, Takieddine Chehhat, Youcef Sklab, Hanane Ariouat, Abderrazak Sebaa, Eric Chenin, Jean-Daniel Zucker,...

    cs.CV · cs.AI

    Automated plant traits recognition from herbarium images is essential for plant sciences, yet remains challenging because background elements (e.g., textual labels, mounting artifacts, and color charts) can introduce shortcut learning, leading models to rely on spurious non-plant cues rather than plant morphology. This bias degrades both generalization and interpretability. In this paper, we introduce AT-ViT, a dual-branch Vision Transformer...

    arxiv.org/abs/2608.21067 · PDF

  11. 11

    CoAnchor: Robust Collaborative Perception under Spatio-Temporal Misalignment via Object-Level Anchors

    Chi Li, Rui Lin, Aobo Ji, Dongzhu Xu

    cs.CV · cs.AI

    Collaborative perception extends the sensing range of a single vehicle by fusing observations from nearby agents, which improves the robustness of autonomous driving. In realistic deployments, however, the received collaborator messages are often affected by both communication delay and relative-pose noise, which jointly cause stale observations, spatial misalignment, and unstable feature fusion. Existing methods usually address these issues...

    arxiv.org/abs/2608.21055 · PDF

  12. 12

    CoST: Semantic-Aware Urban Understanding via Spatial-Temporal Alignment

    Yutian Jiang, Jiabo Liu, Xixuan Hao, Yuxuan Liang

    cs.CV · cs.AI

    Geospatial representation learning from satellite imagery is a fundamental problem for large-scale urban analysis and real-world applications. Despite recent advances, current methods struggle with cross-region generalization and semantic interpretability due to their reliance on region-specific auxiliary data and the neglect of semantic alignment within multi-temporal urban imagery. Therefore, we present CoST, a novel...

    arxiv.org/abs/2608.21041 · PDF

  13. 13

    COMET: Contrastive Motion-Enhanced Temporal Reasoning for Video Multimodal Large Language Models

    Chenghua Zhu, Zhaolu Kang, Qifan Shi, Siyan Wu, Kehan Jiang, Lei Wei, Lianyu Hu, Guangyuan Dong, Mingbo Yang, Rui...

    cs.CV · cs.CL · cs.LG

    Video multimodal large language models have advanced significantly, yet fine-grained motion-temporal understanding remains fragile. The core bottleneck is not only sparse frame sampling, but also the lack of a complete temporal modeling pipeline for explicitly representing frame-to-frame change, enabling appearance-motion interaction, and optimizing temporal direction sensitivity. We propose COMET, a temporally grounded framework that...

    arxiv.org/abs/2608.21030 · PDF

  14. 14

    WA-JEPA: Rethinking the Video JEPA Paradigm for World-Action Modeling in Autonomous Driving

    Xinlin Wang, Yujiao Xiang, Yuheng Zhou, Jingqi Wang, Minqing Huang, Jiajie Huang, Dongxu Wei, Tingguang Zhou, Xiyang...

    cs.CV · cs.AI

    Video Joint Embedding Predictive Architecture (V-JEPA) learns powerful spatiotemporal representations from video through self-supervised latent feature prediction. However, V-JEPA is built around random-mask completion and deterministic regression, making it fundamentally ill-suited for autonomous driving planning that demands future-directed prediction tightly coupled with action. To address this, we rethink the V-JEPA paradigm and present...

    arxiv.org/abs/2608.20974 · PDF

  15. 15

    Explainable Deepfake Detection with Feature-robust Augmentation and Evidence-grounded Explanation Optimization

    Zhu Xu, Jiaqi Tang, Pokai Chen, Yuxin Peng, Yang Liu

    cs.CV · cs.AI

    Explainable deepfake detection extends binary classification by requiring models to not only predict authenticity but also provide interpretable justifications. This expanded scope is critical in practice, where users like forensic analysts need insight into the rationale behind the detection. Despite advancements, current approaches suffer from two critical deficiencies: (1)vulnerability to image quality degradation: detection accuracy...

    arxiv.org/abs/2608.20913 · PDF

  16. 16

    EviRank: Structured Relevance Evidence for Multimodal Image Re-ranking

    Enjun Du, Siyi Liu, Zirong Chen, Xinyu Zuo, Jinwen Luo, Ruiwen Tao, Lisheng Duan, Haijin Liang, Jin Ma, Junfu Pu,...

    cs.CV · cs.LG

    Real-world image search queries are multimodal and compositional: ``find this shirt in pink'' specifies an entity to retain, an attribute to modify, and context to ignore. Yet existing re-rankers either compress such multifaceted relevance into an opaque embedding or rely on free-form chain-of-thought that easily omits or hallucinates fine-grained constraints. Drawing on rubric- and checklist-based evaluation from NLP, we recast multimodal...

    arxiv.org/abs/2608.20886 · PDF

  17. 17

    TRACE: Training-time Report-guided and Clinically Ordered Concept Editing

    Wentao Yue, Tianyou Lai, Jiayu Luo, Qingyu Mao, Ziying Wang, Zhenyuan Ning, Qilei Li

    cs.CV · cs.AI

    Breast ultrasound diagnosis relies on clinically meaningful semantic concepts, yet most deep learning methods adopt end-to-end image-to-label paradigms that lack interpretability and robustness. While concept-based approaches offer a promising alternative, they often assume complete annotations or require multimodal inputs at inference, which significantly limits their real-world applicability. To tackle these issues, we propose Training-time...

    arxiv.org/abs/2608.20809 · PDF

  18. 18

    CertVLA: Certified Defense against Physical Visual Attacks for Vision-Language-Action Models

    Hui Lu, Zhijie Peng, Yuqi Lin, Zaijia Yang, Jiaming He, Shuhan Ye, Yi Yu, Hanwei Zhu, Bingquan Shen, Alex Kot, Xudong Jiang

    cs.CV · cs.AI

    Vision-Language-Action (VLA) policies are vulnerable to localized physical perturbations, yet existing certified patch defenses target discrete labels and cannot directly certify continuous, temporally correlated actions. We introduce CertVLA, a certified defense for closed-loop VLA control under bounded patch and texture attacks. CertVLA proposes a calibrated region of behaviorally consistent actions, while deterministic covering masks...

    arxiv.org/abs/2608.20791 · PDF

  19. 19

    CARD: Diagnosing Belief to Action Routing Failures in Vision Language Models

    Souptik Kumar Majumdar, Fabian Kögel, Andreas Bulling

    cs.CV · cs.AI

    Linear probes and activation steering have uncovered that vision-language models (VLMs) internally represent mental states such as agents' beliefs, knowledge, and intentions. However, it is unclear whether and how these representations are used by downstream predictions along these axes. To close this gap, we introduce Cross-Axis Routing Diagnostic (CARD), which steers activations along one axis while measuring the response of a different...

    arxiv.org/abs/2608.20763 · PDF

  20. 20

    Vis-Poison: Poisoning Visual Knowledge in Multimodal Retrieval-Augmented Generation

    Rujin Liang, Zhongpu Chen, Yuhao Lei, Xin Miao

    cs.CV · cs.AI

    While multimodal retrieval-augmented generation (RAG) systems increasingly rely on images as external knowledge sources, the introduction of poisoned visual evidence can severely compromise multimodal large language model (MLLM) generation. Unlike prior attacks that rely on altering textual metadata, we introduce Vis-Poison, a novel visual knowledge poisoning attack where the poisoned image itself is the attacker-controlled payload, without...

    arxiv.org/abs/2608.20756 · PDF

  21. 21

    Identity-Aware Human-Object Interaction Motion Captioning

    Yiming Wang, Yonghao Dang, Huilai Li, Jiawei Tu, Jianqin Yin

    cs.CV · cs.AI

    Existing human-object interaction (HOI) motion captioning methods typically describe what happens while referring to the subject using generic terms such as "a person" or "someone", without grounding the caption in subject identity. To address this limitation, we introduce Identity-Aware Human-Object Interaction Motion Captioning task. This task requires each generated caption to specify both the subject identity and the corresponding HOI...

    arxiv.org/abs/2608.20690 · PDF

  22. 22

    Aggregate, Don't Adapt: Subject-Level Posterior Aggregation and Transductive Calibration for Cross-Site Parkinsonian Gait Severity

    Junlong Shen

    cs.CV · cs.AI

    We describe the winning entry to the MoCha 2026 Benchmark and Challenge on Parkinsonian Gait, which predicts MDS-UPDRS gait severity from canonicalized SMPL motion recorded at clinical sites unseen during training. The system reaches 0.6945 macro-F1 on the hidden test and ranked first of 58 entries, ahead of the runner-up at 0.5807 and the organizers' baseline at 0.4289, on a frozen public motion encoder with a single $4\times512$ linear...

    arxiv.org/abs/2608.20587 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.