cs.CV · 2026-07-09 · No. 48

Computer Vision and Pattern Recognition, 2026-07-09.

21 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

21 entries
  1. 01

    MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models

    Hyunjae Kim, Dain Kim, Pan Xiao, Serina S. Applebaum, Younjoon Chung, Xuguang Ai, Yu Yin, Roy Jiang, Yuexi Du, Yawen...

    cs.CV · cs.LG

    Medicine is inherently multimodal, requiring clinicians to synthesize information across diverse data streams. Yet the development of multimodal foundation models is constrained by limited access to large-scale, high-quality clinical data. Although PubMed Central (PMC) offers a complementary source of expert-authored image-text data, existing PMC-derived resources remain limited in fidelity, reproducibility, and clinical validation. We...

    arxiv.org/abs/2607.07673 · PDF

  2. 02

    HIVE: Understanding Post-Hallucination Reasoning in Vision Language Models

    Feng He, Zhenting Wang, Qifan Wang, Qiang Guan, Dongfang Liu, Ruixiang Tang, Qiankun Li

    cs.CV · cs.AI

    Hallucinations in vision language models (VLMs) are commonly treated as semantic errors, yet they often arise from partial or ambiguous visual evidence. Prior work mainly focuses on detecting or suppressing hallucinations at generation time, leaving the subsequent reasoning stage largely unexplored. In this work, we study Post Hallucination Reasoning (PHR), the stage in which hallucinated semantics enter the model's inference context and...

    arxiv.org/abs/2607.07507 · PDF

  3. 03

    Heterogeneity-Adaptive Diffusion Schrodinger Bridge for PET-Guided Whole-Body MRI Translation

    Chengbo Wang, Jiacheng Yu, Linjie Bian, Ming Qi, Xiaosheng Liu, Tongtong Che, Jichang Zhang, Shuyu Li, Shaoli Song,...

    cs.CV · cs.AI · cs.LG

    While whole-body multimodal medical imaging scanners have been increasingly recognized for more effective medical applications, the excessive long acquisition time in PET-MR scanning is a major obstacle in more efficient clinical practice. Deep learning-based MRI translation provides a potential solution to reduce scan duration. However, current models often focus on specific anatomical regions and face challenges for whole-body scans that...

    arxiv.org/abs/2607.07401 · PDF

  4. 04

    When Prompts Ignore Structure: Graph-Based Attribute Reasoning for Calibrated VLMs

    Tanay Sodha, Aditya Sharma, Ramya Hebbalaguppe, Vinti Agarwal, Pranav Murthy Yeluripaty

    cs.CV · cs.AI · cs.LG

    Reliable confidence estimation remains a key limitation of test-time adaptation in vision-language models (VLMs), where prompt tuning improves zero-shot accuracy but often degrades calibration due to entropy-driven overconfidence. Prior approaches mitigate this using LLM-derived class attributes and contrastive regularization, yet treat attributes independently, ignoring their relational structure. We propose ARGTCA, which represents (class,...

    arxiv.org/abs/2607.07395 · PDF

  5. 05

    HAJJv2-CrowdCount: Zero-Shot Benchmark for Dense Crowd Counting

    Reem AlYabis, Fares AlTuwaim, AlJawharh AlOtaibi, Mohamed Eltahir

    cs.CV · cs.AI

    Automated crowd counting in Hajj video is difficult not because current models lack capacity, but because the footage violates the assumptions those models were built on: cameras observe the crowd from steep, near-vertical angles, individuals occlude one another extensively, and a single frame can contain well over a thousand people. Benchmarks that test crowd counting in such an environment are either private or not detailed per second. We...

    arxiv.org/abs/2607.07322 · PDF

  6. 06

    CarbonCLIP: Enhance Carbon Prediction from Satellite Imagery via Integrated Street-View Semantics and Temporal Context Training

    Zeru Yang, Fang-Ying Gong, Steve H. L. Yim, Chau Yuen

    cs.CV · cs.AI

    Accurately estimating urban carbon emissions is critical for sustainable urban planning, yet many existing approaches remain difficult to apply consistently across cities due to data-source heterogeneity and the lack of fine-grained semantic-temporal context in remote sensing data. We propose CarbonCLIP, a task-oriented multimodal distillation framework that improves satellite-based carbon emission prediction by transferring contextual...

    arxiv.org/abs/2607.07292 · PDF

  7. 07

    Vision Foundation Models in Radiology: A Scoping Review of Data, Methodology, Evaluation and Clinical Translation

    Alejandro Vergara-Richart, Xavier Rafael-Palou, Almudena Fuster-Matanzo, Ignacio Iborra Roncales, Ángel...

    cs.CV · cs.AI · cs.LG

    Vision foundation models (VFMs) are increasingly being developed for radiological imaging, yet their definition, development and evaluation remain heterogeneous. We conducted a PRISMAScR scoping review of peer-reviewed studies published between January 2017 and March 2026 describing foundation models trained exclusively on radiological imaging data. Sixty-seven studies were included and mapped across three pillars: data scale and...

    arxiv.org/abs/2607.07219 · PDF

  8. 08

    Comparative Study of Domain-adapted VLMs for General Document Visual Question Answering

    Miguel Lopez-Duran, Elena Marrero, Julian Fierrez, Marta Robledo-Moreno, Ruben Vera-Rodriguez, Daniel DeAlcala,...

    cs.CV · cs.LG

    Document Visual Question Answering (DocVQA) presents a complex multimodal challenge, requiring models to exploit visual, textual, and layout information from documents. Although Vision-Language Models (VLMs) have shown remarkable performance in text-vision tasks, their robustness and transferability to different document domains remains underexplored. In this study, we present a comprehensive evaluation of 8 open-source pretrained VLMs on...

    arxiv.org/abs/2607.07179 · PDF

  9. 09

    Tree-of-Thoughts Reasoning for Text-to-Image In-Context Learning

    Stepanida Alekseeva, Jenifer Kalafatovich, Seong-Whan Lee

    cs.CV · cs.AI

    In text-to-image in-context learning (T2I-ICL), a model has to infer a latent compositional pattern from fewshot demonstrations for generating a query image. Recent studies show that state-of-the-art multimodal large language models struggle with this setting, particularly due to limited compositional reasoning and sensitivity to prompt construction. In this work, we propose a Tree-of-Thoughts (ToT) reasoning framework for T2I-ICL that...

    arxiv.org/abs/2607.07117 · PDF

  10. 10

    AT-Attn: Temporal-Aware Cross-Attention for Longitudinal Multimodal Alzheimer's Disease Diagnosis

    Xinyue Du, Yibo Liu, Zhenglei Zhou, Xuancheng Yao, Weimin Zhong, Qiuhui Chen

    cs.CV · cs.AI

    In longitudinal Alzheimer's disease (AD) diagnosis support, clinical and imaging information is often collected at irregular visits. Integrating these multimodal observations may improve diagnostic assessment, but naive fusion can degrade performance when MRI is noisy or intermittently unavailable. We propose AT-Attn, a temporal-aware multimodal framework that combines Change-and-Time encoding, time-biased asymmetric cross-attention, and...

    arxiv.org/abs/2607.07091 · PDF

  11. 11

    Navigating Hierarchy: Hyperbolic Learning on Brain Graphs for Disorder Diagnosis

    Yapeng Li, Bo Jiang, Ziyan Zhang, Dongdong Chen, Zhengzheng Tu

    cs.CV · cs.AI

    Functional brain networks exhibit a hierarchical organization across ROI, community, and whole-brain levels, supporting local processing, inter-community coordination, and global integration. Recent studies have demonstrated that brain community-aware modeling is beneficial for both diagnosis and biomarker identification of brain networks. However, existing brain graph modeling methods often struggle to model ROI-community interactions,...

    arxiv.org/abs/2607.07077 · PDF

  12. 12

    Making Implicit Preservation Intent Explicit in Conversational Image Editing

    Soomin Han, Jihyung Ahn, Bumsoo Kim, Buru Chang

    cs.CV · cs.AI

    Conversational image editing requires preserving not only visible content, but also content that temporarily disappears across turns. When newly added or modified content occludes a previously visible region, that region should reappear if it was never semantically changed. However, existing systems often fail to recover such occluded-but-unchanged content, producing inconsistent or hallucinated results. We introduce OCCUR-Bench, a diagnostic...

    arxiv.org/abs/2607.07051 · PDF

  13. 13

    AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning

    Kyuan Oh, Bumsoo Kim

    cs.CV · cs.AI

    Large vision-language models incur substantial inference costs because high-resolution inputs introduce thousands of visual tokens, many of which are redundant for a given query. Existing pruning methods often combine query relevance and token diversity, yet these objectives can conflict under aggressive compression: relevance-driven selection may overconcentrate the budget on correlated local evidence, while diversity-driven selection may...

    arxiv.org/abs/2607.07033 · PDF

  14. 14

    EdgeCompress: Coupling Multidimensional Model Compression and Dynamic Inference for EdgeAI

    Hao Kong, Di Liu, Shuo Huai, Xiangzhong Luo, Ravi Subramaniam, Christian Makaya, Qian Lin, Weichen Liu

    cs.CV · cs.AR · cs.LG

    Convolutional neural networks (CNNs) have demonstrated encouraging results in image classification tasks. However, the prohibitive computational cost of CNNs hinders the deployment of CNNs onto resource-constrained embedded devices. To address this issue, we propose EdgeCompress, a comprehensive compression framework to reduce the computational overhead of CNNs. In EdgeCompress, we first introduce dynamic image cropping (DIC), where we design...

    arxiv.org/abs/2607.06982 · PDF

  15. 15

    Self-Supervised Pretraining Improves Cross-Site and Cross-Scale Robustness of Point Cloud Leaf-Wood Segmentation

    Heeju Mun, Tackang Yang, Yunsoo Nam, Changhyun Choi

    cs.CV · cs.AI

    The accuracy of existing leaf-wood segmentation methods for tree point clouds varies across forest types and sites. Self-supervised learning (SSL) on point clouds has improved the generalization of deep learning models for forestry point cloud tasks, including biomass regression and individual tree segmentation, but its applicability to leaf-wood segmentation remains untested. In this study, we pretrained Point-M2AE, a widely used SSL...

    arxiv.org/abs/2607.06948 · PDF

  16. 16

    Compass: Prostate Cancer Detection Needs Multi-View Context

    Paul F. R. Wilson, Mohamed Harmanani, Zhuoxin Guo, Obed K. Dzikunu, Hannes Cash, Adam Kinnaird, Brian Wodlinger,...

    cs.CV · cs.LG

    Artificial intelligence (AI) analysis of micro-ultrasound ($μ$US) has shown promise for prostate cancer (PCa) detection. However, most existing AI methods focus on the analysis of single $μ$US images in isolation. By contrast, expert $μ$US readers typically assess a full recorded video study, which provides three-dimensional context, to improve PCa detection compared to single-frame analysis. Inspired by this clinical workflow, we propose...

    arxiv.org/abs/2607.06919 · PDF

  17. 17

    LoCA: Spatially-Aware Low-Rank Convolutional Adaptation of Vision Foundation Models

    Sojung An, Junha Lee, Sujeong You, Nam Ik Cho, Donghyun Kim

    cs.CV · cs.AI · cs.LG

    Pre-trained Vision Foundation Models (VFMs) provide strong visual representations for diverse downstream tasks. The key challenge of VFM adaptation stems from the prohibitive costs of full fine-tuning and catastrophic forgetting. To address this, Low-Rank Adaptation (LoRA) has emerged as the prevailing paradigm for Parameter-Efficient Fine-Tuning (PEFT). However, LoRA is typically designed for transformer self-attention layers parameterized...

    arxiv.org/abs/2607.06918 · PDF

  18. 18

    Smart Scissor: Coupling Spatial Redundancy Reduction and CNN Compression for Embedded Hardware

    Hao Kong, Di Liu, Shuo Huai, Xiangzhong Luo, Weichen Liu, Ravi Subramaniam, Christian Makaya, Qian Lin

    cs.CV · cs.AR · cs.LG

    Scaling down the resolution of input images can greatly reduce the computational overhead of convolutional neural networks (CNNs), which is promising for edge AI. However, as an image usually contains much spatial redundancy, e.g., background pixels, directly shrinking the whole image will lose important features of the foreground object and lead to severe accuracy degradation. In this paper, we propose a dynamic image cropping framework to...

    arxiv.org/abs/2607.06915 · PDF

  19. 19

    ReMoDEx: A Local-to-Global Relevance-Based Model Decision Explainability Framework for large-Scale Image Datasets

    Abhay Kumar Pathak, Mrityunjay Chaubey, Manjari Gupta

    cs.CV · cs.AI

    Deep learning image classifiers achieve strong predictive performance yet remain opaque in how decisions are formed. A model may predict correctly while relying on irrelevant cues, shortcut associations, peripheral structures, or device level artifacts instead of task relevant regions. On large scale datasets this opacity is especially problematic, since inspecting heatmaps one sample at a time cannot scale to thousands of predictions. We...

    arxiv.org/abs/2607.06889 · PDF

  20. 20

    Video2Reaction: Mapping Video to Audience Reaction Distribution in the Wild

    Trang Nguyen, Sidong Zhang, Shiv Shankar, Gauri Jagatap, Deepak Chandran, Andrea Fanelli, Madalina Fiterau

    cs.CV · cs.LG

    Understanding and forecasting audience reactions to video content are crucial for improving content creation, recommendation systems, and media analysis. To enable audience reaction prediction and other content engagement applications, we introduce $\textbf{Video2Reaction}$, a multimodal dataset that maps short movie segments to a distribution of $\textit{induced emotions}$ of viewers in the wild, as expressed through social media....

    arxiv.org/abs/2607.06875 · PDF

  21. 21

    Gen4U: Unifying Video Generation and Understanding via Diffusion

    Michael King, Aravindh Mahendran, Matthew Koichi Grimes, Fedor Kitashov, Adham Elarabawy, Pedro Velez, Maks...

    cs.CV · cs.LG

    Prior work suggests that diffusion representations capture low-level geometry but struggle with high-level semantics. We demonstrate that state-of-the-art video diffusion models overcome this limitation. By systematically probing their intermediate activations using recent mutual-kNN alignment metrics, we reveal a highly structured latent space where visual representations evolve across both network depth and noise levels. We show that while...

    arxiv.org/abs/2607.06856 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.