cs.CV · 2026-08-26 · No. 96

Computer Vision and Pattern Recognition, 2026-08-26.

17 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

17 entries
  1. 01

    LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training

    Andreas Hochlehnert, Marianna Nezhurina, Mehdi Cherti, Andrej Radonjic, Thaddäus Wiedemer, Christoph Schuhmann,...

    cs.CV · cs.AI · cs.LG

    We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl. From these, we download 80M videos with a total duration of 10 million hours. The dataset is designed for multimodal pre-training across the video, audio, and image modalities. Using content-aware scene detection, we extract clips for which we synthetically generate video and audio...

    arxiv.org/abs/2608.24845 · PDF

  2. 02

    Ensemble of Convolutional Neural Networks for StrokePrediction: Towards Improved Diagnostic Accuracy

    Md Shahriar Sajid

    cs.CV · cs.AI

    Brain stroke, known for its high mortality and incidence rates, poses significant health risks and requires rapid intervention for survival. Early diagnosis and preventive measures can greatly reduce life loss and disabilities. Recent advancements in deep learning have led to novel computer-aided diagnostic techniques for early stroke detection. This study proposes an intelligent system that predicts potential strokes using eleven features,...

    arxiv.org/abs/2608.24771 · PDF

  3. 03

    MoTE: Mixture of Task Experts for Multi-Task Video Understanding

    Muhammad Asad Ali, Umar Khan, Nadia Robertini, Didier Stricker

    cs.CV · cs.LG

    Procedural video-language models must solve heterogeneous tasks from the same visual evidence, including action recognition, forecasting, and procedure prediction. Dense transformer decoders share the same feed-forward networks across tasks, which can entangle task behavior and make controlled capability expansion difficult. Sparse Mixture-of-Experts (MoE) decoders provide conditional computation, but token-level learned routing is not...

    arxiv.org/abs/2608.24763 · PDF

  4. 04

    Weakly Supervised Seafloor Segmentation for Seagrass Habitat Mapping in Side-Scan Sonar Imagery

    Hayat Rajani, Nuno Gracias, Rafael Garcia

    cs.CV · cs.LG

    Seagrass meadows are crucial blue-carbon habitats, and mapping their extent is a prerequisite for coastal management and carbon inventory. Optical satellite sensors cover large areas but cannot reach deep or turbid water, whereas side-scan sonar (SSS) images the seabed at high resolution and at any depth. Interpreting SSS, however, still relies on dense manual annotation, which is slow and costly. We address this by adapting a weakly...

    arxiv.org/abs/2608.24756 · PDF

  5. 05

    Deep Learning Super Resolution for Satellite Cloud Mask Downscaling

    Angelos Georgakis, Valentina Kanaki, Giorgos Giannopoulos, Stella Girtsou, Ioannis Kontogiorgakis, Charalampos...

    cs.CV · cs.AI

    A vast amount of optical satellite data is being transmitted to Earth-based servers every day, and more than half of this data is affected by haze or clouds. Additionally, this data suffers from the fundamental trade-off between spatial and temporal resolution, which remains largely unresolved, making the acquisition of continuous high-resolution satellite observations of clouds an ongoing challenge. This work addresses this challenge by...

    arxiv.org/abs/2608.24715 · PDF

  6. 06

    Low-Rank Ternary Adaptation for Fine-Tuning Transformers

    Alexandru-Dragos Manolache, Yunqiang Li, Jan van Gemert

    cs.CV · cs.LG

    Ternary transformers offer extreme memory and compute efficiency, but existing low-bit LoRA-based methods cannot directly fine-tune ternary weights. Current approaches either require dequantization, restoring low-bit base weights to higher precision to merge with adaptation weight, or update only quantization parameters, preventing a merged model that remains ternary. We propose ternary multiplicative adaptation, which represents discrete...

    arxiv.org/abs/2608.24469 · PDF

  7. 07

    Markerless Pose Estimation for Resistance Training Technique Assessment

    Joseph Turner, Jeff Clark, Nawid Keshtmand

    cs.CV · cs.AI

    Resistance training can be a high risk activity, and safe form is essential to avoiding injury. Laboratory-based movement analysis provides quantitive technique assessment, yet is not easily accessible. Markerless pose estimation infers body landmarks from images or video without physical markers and could offer a feasible alternative for technique assessment. We present a pose estimation framework to evaluate resistance-training technique...

    arxiv.org/abs/2608.24384 · PDF

  8. 08

    Metadata-Aware Adaptation of a Generative Foundation Model for Conditional CMR Synthesis

    Marc Rodríguez, Grzegorz Skorupko, Nay Aung, Steffen E Petersen, Karim Lekadir, Polyxeni Gkontra

    cs.CV · cs.AI

    Synthetic image generation is a promising strategy to address data scarcity and the underrepresentation of clinically important phenotypes in medical imaging, yet generating images that faithfully reflect meaningful patient characteristics remains challenging. In this work, we investigate metadata-conditioned cardiac magnetic resonance (CMR) synthesis using a pretrained latent diffusion model, encoding structured clinical metadata and slice...

    arxiv.org/abs/2608.24342 · PDF

  9. 09

    Mind the Student: Behavioral and Contextual Cues for Automated Engagement Prediction in Online Learning

    Alperen Kantarci, Visvanathan Ramesh, Gemma Roig

    cs.CV · cs.AI · cs.HC · cs.LG

    The prediction of student engagement from the online tutoring videos is difficult because engagement is a multidimensional construct comprising distinct behavioral, emotional, and cognitive states. A reliable prediction requires bringing together different types of behavioral signals as well as expressive cues. Through our analysis of the CASED dataset, it is clear that engagement prediction gets even harder due to the high inter-person...

    arxiv.org/abs/2608.24340 · PDF

  10. 10

    Rethinking Pre-Training and Augmentation for Zero-Shot Cross-City Object Detection

    Long Hoang Pham, Quoc Pham-Nam Ho, Huy-Hung Nguyen, Duong Nguyen-Ngoc Tran, Ngoc Doan-Minh Huynh, Cu Quoc Le,...

    cs.CV · cs.AI

    Real-world deployment of traffic surveillance systems is bottlenecked by geographic domain shift, in which models trained in one city underperform when applied to an unseen target city. Conventional domain adaptation relies on hyperparameter-sensitive architectures or direct profiling of target data. Both are fundamentally precluded in privacy-conscious ecosystems that require completely blind training and evaluation loops. In this setting,...

    arxiv.org/abs/2608.24154 · PDF

  11. 11

    PlaceSeek: Human-Centered Geospatial Retrieval of Urban Outdoor Places via Semantic Grounding and Affective Alignment

    Ziqi Cui, Shangyu Lou

    cs.CV · cs.AI · cs.IR

    People search for urban outdoor places not only by category or function, but also by what activities a place can support and how it is perceived. Existing geospatial retrieval remains largely POIcentric and metadata-driven, making it difficult to satisfy openended, affective, or activity-oriented needs. We present PlaceSeek, a human-centered outdoor place retrieval framework that maps natural-language queries to geolocated street-view...

    arxiv.org/abs/2608.24133 · PDF

  12. 12

    Syn2RealTrack: Bridging the Gap Between Synthetic and Real-World Datasets for Online Multi-View Multi-Target Tracking

    Duong Nguyen-Ngoc Tran, Ngoc Doan-Minh Huynh, Cu Quoc Le, Hoang-Khang Nguyen, Long Hoang Pham, Huy-Hung Nguyen, Quoc...

    cs.CV · cs.AI

    Multi-camera 3D perception systems for warehouse scenes are trained largely on synthetic data and evaluated on physically captured environments. The resulting synthetic-to-real gap, which corrupts ground-plane localization and cross-camera identity association, is usually treated as one deficiency for a single domain-adaptation module to absorb; we argue instead that it enters the pipeline at three separable points: the camera calibration,...

    arxiv.org/abs/2608.24130 · PDF

  13. 13

    TransPhy: Visual In-Context Learning for Physically Grounded Image Editing

    Siyi Xie, Xuanke Shi, Jinsheng Quan, Haoran Tang, Zukai Chen, Lei Yang, Quan Wang

    cs.CV · cs.AI

    Visual demonstrations provide a natural interface for specifying image transformations that are difficult to describe exhaustively with text. However, existing visual in-context learning (VICL) methods primarily focus on appearance-level relation transfer and provide limited support for physically grounded transformations, whose outcomes depend on material properties, geometry, object interactions, and environmental conditions. Given a...

    arxiv.org/abs/2608.24119 · PDF

  14. 14

    MatReplace: A Reference-Free, Conditioning-Aligned Benchmark for Material Replacement in Interior Scenes

    Mingzhe Du, Thong Thanh Nguyen, Nguyen Tran Cong Duy, See-Kiong Ng, Luu Anh Tuan

    cs.CV · cs.AI

    Material replacement is a common interior-design operation: changing the material of a selected surface while preserving its geometry, surroundings, and illumination. Despite its commercial relevance, no public benchmark isolates this task, and evaluating it is challenging. Reference-based metrics penalize valid outputs in this inherently one-to-many setting, favor the style of the reference generator, and cannot fairly compare editors that...

    arxiv.org/abs/2608.24107 · PDF

  15. 15

    Joint-Embedding Prediction of Masked Point Tubes for Self-Supervised Learning on 4D Point Cloud Videos

    Jheng-Ling Lee, Shang-Tse Chen

    cs.CV · cs.LG

    Self-supervised representation learning for 4D point cloud videos is challenging because annotations are costly and reconstruction-based pretraining can overemphasize low-level geometric details. We propose a JEPA-style framework that learns from unlabeled spatiotemporal point clouds through latent point-tube prediction. Instead of reconstructing raw coordinates, the model masks spatiotemporal regions and predicts their target representations...

    arxiv.org/abs/2608.24093 · PDF

  16. 16

    VisCache: Visual KV Cache Pruning for Efficient Vision Large Language Model Inference

    Lyuke Wang, Zhuo Li, Guangxu Zhu

    cs.CV · cs.AI

    While Vision Large Language Models (VLLMs) have achieved remarkable success in multimodal reasoning, their long-context inference remains prohibitively expensive due to the massive computation and memory overhead of visual Key-Value (KV) caches. Existing KV compression methods often apply uniform pruning across visual tokens and layers, leading to substantial information loss and degraded performance.To address this challenge, we propose...

    arxiv.org/abs/2608.24063 · PDF

  17. 17

    IterCAD: Iterative Program Repair for CAD Code Generation from Orthographic Views

    Yuchuan Wu, Ke Niu, Haiyang Yu, Zhuofan Chen, Xiangyang Xue, Bin Li

    cs.CV · cs.AI

    Generating executable parametric CAD code from dimension-annotated orthographic drawings is a challenging task requiring geometric understanding, procedural reasoning, and precise numerical prediction. Existing vision-language approaches typically formulate this problem as one-shot generation, preventing the model from inspecting intermediate CAD results and correcting early mistakes, often leading to non-executable code or geometrically...

    arxiv.org/abs/2608.24020 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.