cs.CV · 2026-08-05 · No. 75

Computer Vision and Pattern Recognition, 2026-08-05.

15 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

15 entries
  1. 01

    Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent

    Zhen Fang, Yu Zeng, Wenxuan Huang, Yiming Zhao, Shiting Huang, Tianfei Ren, Qi Lu, Qingnan Ren, Qisheng Su, Lionel...

    cs.CV · cs.AI

    We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video streams, a setting that demands dense spatiotemporal grounding coupled with open-web exploration. Preliminary evaluations reveal two critical bottlenecks in current models: (1) modality bias, where agents bypass visual tools in favor of textual search, and (2) parametric knowledge leakage, where models rely on internal memory rather...

    arxiv.org/abs/2608.03979 · PDF

  2. 02

    When and Where to Look: Adaptive Visual Evidence Scheduling for Efficient Long Video Understanding

    Ke Li, Jiayu Chen, Maoliang Li, Zihao Zheng, Hailong Zou, Hengyi Zhang, Xuanzhe Liu, Xiang Chen

    cs.CV · cs.AI

    Efficient long-video understanding requires vision--language models (VLMs) to reason over a small number of frames selected as sparse visual evidence. Existing relevance-based methods rely on static one-shot selection with fixed frame budgets and candidate pools, while agent-based schedulers achieve adaptivity through costly multi-round reasoning and interactive search. We propose EcoFrame, a training-free framework for low-overhead...

    arxiv.org/abs/2608.03918 · PDF

  3. 03

    CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

    Mercy Prasanna Ranjit, Anirban Porya, Sathvik Joel, Niharika Vadlamudi, Nikhilesh Chowdary Eathamukkala, Prasanth V,...

    cs.CV · cs.AI

    A clinically useful chest X-ray system must go beyond fluent report generation: it should classify findings with tunable decision thresholds, localize them spatially, and derive the anatomical measurements upon which many diagnoses depend. Today's Vision-Language Models (VLMs) treat these as separate problems, if they address them at all, leaving a gap between what radiologists need and what generative models provide. We introduce CARE-X, a...

    arxiv.org/abs/2608.03890 · PDF

  4. 04

    Geo-Embed: Towards Unified Multimodal Embeddings for Urban Understanding

    Jiapeng Li, Yong Li, Junjie Zhou, Fan Zhang, Yu Liu

    cs.CV · cs.LG

    Geospatial and urban applications increasingly require models to compare heterogeneous evidence across street-view imagery, remote-sensing observations, text descriptions, region proposals, and temporal change cues. However, existing multimodal embedding models and benchmarks are still largely designed and evaluated around general-purpose image-text matching, leaving unclear whether unified embedding space can support heterogeneous geospatial...

    arxiv.org/abs/2608.03826 · PDF

  5. 05

    FlowForm: Synergizing Fluid Physics with Topological Consistency for Satellite Flood Synthesis

    Zhang Weihui, Wang Ruizhi, Xu Hongye, Wang Huiqiong, Sun Li, Song Mingli

    cs.CV · cs.AI

    Developing robust flood assessment models requires high-quality paired satellite imagery, yet such data remain scarce for flood-specific image generation. Although generative models provide a promising means of data augmentation, existing methods often yield implausible spatial layouts of flooded regions and distort scene structures. We propose FlowForm, a framework for satellite flood synthesis that integrates SWE-inspired latent...

    arxiv.org/abs/2608.03822 · PDF

  6. 06

    UHP Detection: LVLMs have their Unique Hallucination Pattern in the Consistency Space

    Amir Mohammad Ezzati, Kiyan Rezaee, Bardiya Kariminia, Mohamad Amin Yousefi, Asal Mohammadjafari Mamaqani, Behrad...

    cs.CV · cs.AI

    Large vision--language models (LVLMs) demonstrate strong multimodal reasoning capabilities but remain prone to hallucination, where model predictions are not grounded in visual evidence. Existing black-box hallucination detection methods estimate uncertainty through a single consistency metric, implicitly assuming that model uncertainty can be adequately characterized by a single measure. However, hallucinations exhibit diverse manifestations...

    arxiv.org/abs/2608.03817 · PDF

  7. 07

    Attention is Case-Sensitive

    Maximilian Dillitzer, Tin Stribor Sohn, Jason J. Corso, Michael Auerbach

    cs.CV · cs.CL · cs.LG

    In human visual perception, uppercase lettering serves as a natural salience cue that captures attention within lowercase text. In this paper, we present a systematic empirical characterization study revealing that Large Language Models (LLMs) exhibit an analogous property: letter casing modulates internal attention allocation. Through analysis across 13 models, nine LLMs and four Vision-Language Models (VLMs), with diverse tokenization...

    arxiv.org/abs/2608.03711 · PDF

  8. 08

    Test-Time Augmentation for Tabular-to-Image Classifiers under Distribution Shifts

    Malena Loza, Felipe Grijalva, Eva Milara, Luis Bote-Curiel, Francisco J. Lara-Abelenda, David Chushig-Muzo

    cs.CV · cs.LG

    Tabular-to-image methods that convert tabular data into visual representations have emerged as a novel paradigm for leveraging the high performance of deep learning models. Despite their advantages, the robustness of these methods under distribution shifts remains under explored. Test-Time Augmentation (TTA) is an effective approach in image classification to improve model generalization and robustness, where predictions over multiple...

    arxiv.org/abs/2608.03557 · PDF

  9. 09

    How Many Labels Are Enough? ALDA: Active Learning Deployment Advisor for Medical Image Classification

    Julia Machnio, Mads Nielsen, Mostafa Mehdipour Ghazi

    cs.CV · cs.AI · cs.LG

    Active learning (AL) promises to reduce the cost of medical imaging projects by lowering the number of clinical labels required. However, practical deployment requires committing to a sampling strategy before the full annotation budget is spent, and choosing the wrong strategy can increase rather than decrease costs. We propose Active-Learning Deployment Advisor (ALDA), a deployment-oriented framework for AL method selection under clinical...

    arxiv.org/abs/2608.03511 · PDF

  10. 10

    Dual-domain U-Nets with embedded back projection operators for motion-resolved 4D CBCT reconstruction

    Ivo Herzig, Pascal Paysan, Daniel Barco, Marc André Stadelmann, Frank-Peter Schilling, Igor Peterlik, Michal...

    cs.CV · cs.LG

    Four-dimensional cone beam CT (4D CBCT) is important for image-guided radiation therapy of thoracic cancers, but its use is limited by long scan times, causing high patient dose and motion/sparse-sampling artifacts. We propose a deep learning method for motion-resolved 4D CBCT reconstruction from conventional free-breathing scans, without a respiratory signal or explicit projection binning. Our CNN takes free-breathing 3D CBCT projections as...

    arxiv.org/abs/2608.03430 · PDF

  11. 11

    OliveGemma: A 3 Billion Visual Language Model for Recognising the Mediterranean & European Diet

    Dimitrios I. Zaridis, Traianos Tsiokris, Vasileios C. Pezoulas, Daphni Plati, Eugenia Mylona, Eleni Georga, Nikos...

    cs.CV · cs.AI

    Image based dietary assessment offers a scalable alternative to self reported food diaries, yet fine-grained food recognition remains challenging due to high intra-class variability and visually similar dishes. This study presents OliveGemma, a vision language model for recognising and reasoning about Mediterranean and European cuisine. Built on the open-weight PaliGemma-2-3B architecture, OliveGemma is fine-tuned with LoRA on a unified...

    arxiv.org/abs/2608.03428 · PDF

  12. 12

    Distilled Roads: Generalisable Road Network Extraction Across Sensors, Resolutions, and Region

    Sanayya, Rakshith Sathish, Ashwathi Nambiar

    cs.CV · cs.AI · cs.LG

    Road network segmentation from satellite imagery remains challenging due to large geographic variation in road appearance, occlusions, and domain shifts introduced by differing resolutions and sensors. Existing models, typically trained under narrow resolution--region combinations, generalise poorly to unseen environments such as rural settings, regions with distinct road materials, or imagery from new satellite platforms, often producing...

    arxiv.org/abs/2608.03407 · PDF

  13. 13

    SRAP: SVD-Refined Adversarial Perturbations for Imperceptible Face-Swap Defense

    Sungwon Cho, Kwanghyun Ko, Myungjoo Kang

    cs.CV · cs.LG

    Deepfake technologies pose increasing threats to facial privacy and identity security, motivating proactive defenses that protect facial images before misuse. Although adversarial perturbations generated by projected gradient descent (PGD) can disrupt the identity representations used by face-swapping models, their visual quality is degraded by two characteristics: perturbations are distributed broadly over the image, including...

    arxiv.org/abs/2608.03395 · PDF

  14. 14

    When Oracle Conditioning Misleads Deployment: Conditioning-Availability Bias in Echocardiographic Segmentation

    Dang P. M. Cao, Hieu D. Pham, Hieu Pham

    cs.CV · cs.AI

    Conditional segmentation models may be trained and evaluated with auxiliary signals cleaner than those available at deployment. We study this protocol-level manifestation of shortcut learning and auxiliary-variable shift in phase-conditioned echocardiographic segmentation. The complementary gap pair measures loss on the deployable oracle-estimated pathway and probes sensitivity on the oracle-random pathway. On held-out CAMUS data, one...

    arxiv.org/abs/2608.03342 · PDF

  15. 15

    Test-Time Scaling for Safe Text-Guided Image Generation via Intermediate Clean Estimates

    Jinya Sakurai, Shueicheng Yan, Xun Xu

    cs.CV · cs.AI

    Ensuring safety and policy compliance in text-to-image diffusion models remains a critical challenge, as benign or adversarial prompts can often elicit prohibited content, e.g. nudity and protected intellectual property. While training-based unlearning methods are effective, they are computationally expensive and prone to catastrophic interference with general capabilities. Conversely, existing test-time defenses are primarily prompt-centric,...

    arxiv.org/abs/2608.03284 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.