cs.CV · 2026-08-25 · No. 95

Computer Vision and Pattern Recognition, 2026-08-25.

13 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

13 entries
  1. 01

    EG-ARSA: An Expert-Grounded Open Model for Visual Road Safety Auditing in Low-Resource Settings

    Md Thamed Bin Zaman Chowdhury, Moazzem Hossain

    cs.CV · cs.AI

    Road traffic injuries remain a major challenge in low- and middle-income countries, where proactive road safety auditing is limited by incomplete crash records, shortages of qualified auditors, and the high cost of large-scale field inspections. To address this problem, we propose Expert-Grounded Distillation (EGD), a novel artificial intelligence framework that transfers institutional road safety expertise into a compact vision-language...

    arxiv.org/abs/2608.23563 · PDF

  2. 02

    Predicting Multiple Clinical Outcomes Related to Functional Recovery and Social Isolation Among Older Adults After Lower-Limb Fracture or Hip Replacement

    Santosh Ray, Pratik K. Mishra, Ali Abedi, Charlene H. Chu, Amir Ahmad, Shehroz S. Khan

    cs.CV · cs.LG

    Older adults recovering after lower-limb fracture or hip replacement may experience complex recovery trajectories. Most of the time, these clinical aspects are studied in isolation, masking their joint impact on recovery. This study used the MAISON-LLF dataset, which contains multimodal sensor and clinical assessment data from 18 older adults recovering in the community after lower-limb fracture or hip replacement. Participants were monitored...

    arxiv.org/abs/2608.23531 · PDF

  3. 03

    Towards Comprehensive Basketball Understanding

    Yirong Hu, Jiayuan Rao, Yu Zhang, Shangzhe Di, Weidi Xie

    cs.CV · cs.AI

    Understanding a basketball game requires recognizing events, localizing actions, identifying players, and relating these to structured game knowledge. Existing benchmarks primarily evaluate these abilities one at a time, leaving the interactions among these abilities under-explored. We introduce BasketballBench, a multimodal benchmark comprising 7,980 questions across ten tasks in text, image, and video. It is built from the 2025-2026 NBA...

    arxiv.org/abs/2608.23435 · PDF

  4. 04

    Photorealistic Novel View Synthesis of Human Faces using Next-Scale Transformers

    Federico Stella, Fei Jiang, Zhongshi Jiang, Zohar Barzelay, Emanuel Garbin, Amin Jourabloo, Liuhao Ge

    cs.CV · cs.LG

    Photorealistic novel view synthesis of people remains challenging at high spatial resolutions and across multiple target cameras, where preserving identity, fine appearance details, and geometric coherence is critical. We build on the next-scale autoregressive paradigm and adapt it for human-centric view synthesis by enabling higher image resolutions, multi-view outputs and stronger cross-view consistency in a single forward pass. We train on...

    arxiv.org/abs/2608.23410 · PDF

  5. 05

    DF-MoE: Generalizable Deepfake Detection via Multimodal Sparse Mixture-of-Experts

    Vlad Hondru, Florinel Alin Croitoru, Iuliana Georgescu, A. Sophia Koepke, Radu Tudor Ionescu

    cs.CV · cs.AI · cs.LG

    Audio-visual deepfake detection is an actively studied topic, where one of the main challenges is to develop detectors able to generalize across deepfake generation methods. We conjecture that overfitting can be mitigated by extracting multiple high-level cues from the available audio and visual modalities via pre-trained models. We therefore assemble a wide variety of pre-trained models to extract features that encode mouth movements, face...

    arxiv.org/abs/2608.23363 · PDF

  6. 06

    Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents

    Wenqi Liu, Shijie Ma, Yunxiao Wang, Meng Liu, Qile Su, Han Liu, Bohan Hou, Xuanyu Zheng, Changyi Liu, Tianke Zhang,...

    cs.CV · cs.AI

    Open-world video understanding often requires a model to locate sparse visual evidence and acquire external knowledge that is absent from the video and its parametric memory. While Thinking-with-Videos enables active temporal perception and Deep Research supports multi-step information seeking, the two capabilities are typically developed in isolation. We introduce VideoRover, a unified Video Deep Research framework that iteratively...

    arxiv.org/abs/2608.23329 · PDF

  7. 07

    E2S-Pruner: Progressive Two-Stage Evidence Fusion for Visual Token Pruning in Vision-Language Models

    Taoyu Qian, Qi Wang, Daqian Shi, Yuanhao Jiang, Shang Gao, Hualong Yu

    cs.CV · cs.AI

    Vision-language models typically encode an image into hundreds of visual tokens, incurring substantial inference latency and GPU memory overhead. Existing pruning methods largely rely on attention scores and directly aggregate outputs across attention heads and network layers, making it difficult to characterize evidential uncertainty and conflict. We propose E2S-Pruner, a progressive two-stage evidence-fusion framework for visual token...

    arxiv.org/abs/2608.23253 · PDF

  8. 08

    BenthicDINO: Physics-Informed Self-Distillation for View-Invariant Side-Scan Sonar Representations

    Taqi Hamoda, Hayat Rajani, Nuno Gracias

    cs.CV · cs.AI · cs.LG

    Automated perception in side-scan sonar (SSS) imagery is severely hindered by physical acoustic artifacts, resulting in representations that inextricably mix intrinsic seabed reflectivity with transient viewing geometries. Existing self-supervised learning (SSL) frameworks rely on augmentations designed for natural images, failing to account for acoustic degradation and explicitly enforce view-invariance. To address this gap, we introduce a...

    arxiv.org/abs/2608.23215 · PDF

  9. 09

    Cultural Moment Benchmark: Evaluating Video Cultural Reasoning and Grounding in Southeast Asia

    Burak Satar, Zhixin Ma, Cheng Yu-Tong, Huy Hoang Tran, Phuong Anh Nguyen, Chong-Wah Ngo

    cs.CV · cs.AI · cs.CL · cs.IR · cs.MM

    Cultural understanding in video means more than recognizing what is visible; it requires grasping the symbolic and temporal significance of cultural concepts. We decompose this into three abilities: naming what a concept symbolizes, visually recognizing it on video, and locating its sub-events in time. Existing video-cultural benchmarks tend to test what is seen, collapsing these three abilities into a single score that hides the bottleneck....

    arxiv.org/abs/2608.23065 · PDF

  10. 10

    Coarse Indexing, Fine Evidence: Decoupling Temporal Granularity in Long-Video RAG

    Zhe Jin, Zhimin Lin, Bin Zheng, Junhua Fang, Huihua Yang

    cs.CV · cs.AI

    Graph-based retrieval-augmented generation (RAG) provides a scalable paradigm for long-video understanding, but existing systems typically inherit a fixed temporal granularity from video segmentation when constructing their retrieval index. We argue that this design unnecessarily couples indexing granularity with evidence granularity: coarse representations can often suffice for locating relevant temporal regions, while fine-grained evidence...

    arxiv.org/abs/2608.23011 · PDF

  11. 11

    WildHandBench: A Benchmark for Handwritten Text Understanding that Challenges MLLMs and Humans

    Jun Zhang, Qiao Zhao, Cheng Cui, Jianying Qu, Zhongkai Sun, Jianwen Yang, Changda Zhou, ZhuoXin Liu, Shubin Han

    cs.CV · cs.AI

    While the top model on OmniDocBench now reaches 96.34% overall on printed-document parsing, the ability of current models to handle challenging handwritten documents remains largely uncharacterized. Existing benchmarks focus on isolated text or formulas, overlook handwritten tables and real-world degradation, and report aggregate accuracy without explaining why models fail. We present WildHandBench, a benchmark containing 500 handwritten...

    arxiv.org/abs/2608.22959 · PDF

  12. 12

    GuidedFlow: An Attention-Guided Framework for Anomaly Detection in Additive Manufacturing

    Sosmita Paul, Krishna Roy

    cs.CV · cs.LG

    Additive Manufacturing (AM) plays a vital role in the ongoing industrial revolution. However, quality control remains crucial and challenging due to printing defects or potential cyber-physical intrusions. Image or video-based anomaly detection is a key effort towards addressing these challenges. Various approaches have been explored in this domain, including reconstruction-based, embedding-based, and flow-based methods. Though normalizing...

    arxiv.org/abs/2608.22789 · PDF

  13. 13

    Object-Uni: A Unified Model for Object-Centric Spatial Understanding and Controllable Generation

    Mining Tan, Yinuo Wang, Ziqi Zhou, Weize Quan, Sifei Li, Jingdong Chen, DanDan Zheng, Libin Wang, Weiming Dong

    cs.CV · cs.AI

    Unified models for visual understanding and generation have made rapid progress, yet they still lack the ability to understand and manipulate the spatial states of object instances. Existing models can describe objects in natural language, but they struggle to precisely represent continuous object poses and generate geometrically consistent images under target viewpoints. To mitigate this, we propose \emph{Object-Uni}, a unified model for...

    arxiv.org/abs/2608.22757 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.