cs.CV · 2026-08-15 · No. 85

Computer Vision and Pattern Recognition, 2026-08-15.

18 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

18 entries
  1. 01

    AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design

    Yaxin Luo, Haobin Jiang, Jialv Zou, Xu Huang, Wenhao Yan, Haodong Li, Zhengrong Yue, Jing Li, Xiaofu Chen, Xiaohan...

    cs.CV · cs.AI · cs.CL

    Transforming multimodal sources into condensed and structured media outputs can be fundamentally conceptualized as a long-horizon agentic process centered on a model-harness system. While an ideal harness system should align with human design priors and accumulate reusable experience through empirical exploration to drive recursive self-improvement, existing paradigms remain static and fall short of this capability. In this paper, we present...

    arxiv.org/abs/2608.13560 · PDF

  2. 02

    TabSOM: A tabular-to-image encoding method based on self-organizing maps

    David Chushig-Muzo, María Ángeles Rodríguez de Cara, Eva Milara, Francisco J. Lara-Abelenda, Luis Zhinin-Vera, Diego...

    cs.CV · cs.LG

    Tabular-to-image methods have emerged as novel approaches to leverage the high predictive performance of convolutional neural networks and vision transformers. They convert tabular data into image representations, mapping each feature at a fixed pixel location derived from a dimensionality-reduction method (e.g., t-SNE, UMAP, PCA). However, they encode only the marginal value of each feature and discard information about feature...

    arxiv.org/abs/2608.13513 · PDF

  3. 03

    TraVEL: Trajectory-Guided Video Embedding Learning for Driving-Video Retrieval

    Yi-Chung Chen, Philip Jacobson, Tom Lampo, Yiren Lu, Jin Yao, David I. Inouye, Jing Gao, Danhua Guo, Burhan Yaman

    cs.CV · cs.LG

    Efficiently retrieving relevant clips from large-scale driving logs is essential for data curation, model development, and safety analysis. Structured and rule-based retrieval systems can explicitly target driving events, but typically require expert-defined rules, auxiliary data, and multi-stage perception pipelines. Multimodal embedding models offer a simpler and more efficient alternative by representing each video with a single searchable...

    arxiv.org/abs/2608.13495 · PDF

  4. 04

    MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification

    Daniel Perkins, John Squires, Janou Milligan, Chandra Raskoti, Linda Ungerboeck

    cs.CV · cs.AI · cs.CL · cs.LG

    Modern image classification models excel when trained on single task-specific datasets but often struggle to generalize across domains and difficulty levels. We propose ARMDIL, an Adaptive Router for Multi-Domain Image classification with LLMs. ARMDIL is an ensemble that uses a multimodal large language model (MLLM) agent to dynamically route each image to the most suitable vision backbone. Our diverse ensemble employs convolutional neural...

    arxiv.org/abs/2608.13463 · PDF

  5. 05

    UniTexture: Cross-Task Universal Adversarial Textures for Vision-Language-Action Models

    Yukun Dai, Mingzhe Dai, Tianshi Wang, Fengling Li, Jingjing Li, Lei Zhu

    cs.CV · cs.AI

    Vision-Language-Action (VLA) models have emerged as generalist robotic policies capable of following diverse language instructions and performing a wide range of manipulation tasks. However, their direct control over embodied agents also exposes them to adversarial interference that may cause unsafe physical behaviors. Existing attacks on robotic policies are typically optimized for a single task or instruction, leaving the cross-task...

    arxiv.org/abs/2608.13453 · PDF

  6. 06

    Sign Language Video Synthesis via Loss-Guided Multi-Expert GANs

    Dingzhan Nong, Zhihao Ren, Ziqi Li, Tim Lo

    cs.CV · cs.AI

    This preliminary technical report presents a framework for sign language video synthesis using a loss-guided multi-expert Generative Adversarial Network (GAN) to enhance communication for individuals with hearing impairments. Three specialized discriminators -- global, hand, and head -- each guide a corresponding expert branch in the generator toward a distinct visual region, enabling implicit feature specialization without explicit diversity...

    arxiv.org/abs/2608.13368 · PDF

  7. 07

    GeoCache: Training-Free Acceleration of Multi-View Texture Diffusion via Geometric Delta Transport

    Haotang Li, Zhenyu Qi, Shaohan Henry Wang, Kebin Peng, Yutong Zhao, Zi Wang, Bo Liu, Huanrui Yang, Sen He

    cs.CV · cs.AI

    Geometry-conditioned multi-view diffusion enables high-quality 3D texture generation, but its repeated per-view denoiser evaluations introduce substantial computational cost. Existing training-free accelerators primarily exploit temporal redundancy by reusing computation across denoising steps. In multi-view texturing, however, skipping a step also removes the cross-view interaction that continually aligns different observations of the same...

    arxiv.org/abs/2608.13255 · PDF

  8. 08

    CoverPrune: Coverage-Driven Token Pruning for 3D VLMs via Optimal Transport

    Peng Ling, Yingda Yin, Lingting Zhu, Weikai Chen, Shengju Qian, Zeyu Hu, Xin Wang, Wenming Yang

    cs.CV · cs.AI

    While 3D Vision-Language Models (3D VLMs) have demonstrated remarkable spatial reasoning capabilities, they suffer from massive visual token counts that create severe computational bottlenecks during inference. Existing token pruning methods primarily rely on diversity-based selection, discarding similar tokens to maximize dispersion. However, in 3D environments, this approach frequently drops representative prototype tokens in favor of...

    arxiv.org/abs/2608.13226 · PDF

  9. 09

    NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video

    Yuheng Huang, Jianlang Chen, Jiayang Song, Hua Qi, Aza Kai, Vincent Markert, Edison Marrese-Taylor, Jianjun Zhao, Lei Ma

    cs.CV · cs.AI · cs.MM

    Long-form video understanding encompasses tasks that go beyond retrieving isolated events, including tracking an evolving narrative and interpreting social meaning that may remain implicit. However, existing benchmarks rarely evaluate these capabilities jointly, particularly in high-context, non-English media. To address this gap, we introduce NARU, a benchmark designed to evaluate Narrative evolution and Reasoning on cultural Understanding...

    arxiv.org/abs/2608.13210 · PDF

  10. 10

    TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic Restraint

    Fnu Pramono, John Cai, Sourabh Kulkarni

    cs.CV · cs.AI · cs.CL · cs.LG

    When visual evidence is occluded or chaotic, models should abstain. In this paper, we show that Vision-Language Models (VLMs) can internally distinguish when abstention is required, but fail to express it anyway. We introduce TRAPSBench, a procedurally generated video benchmark of 1,404 matched physics pairs in which a single targeted change renders the outcome undeterminable from the visual evidence. Furthermore, we introduce Penalized...

    arxiv.org/abs/2608.13167 · PDF

  11. 11

    MergeOver: Post-Training Token Merging for Recursive Vision Transformers

    Junseo Kim, Uraz Odyurt, Amirreza Yousefzadeh

    cs.CV · cs.LG

    Vision Transformers (ViTs) demonstrate exceptional performance in computer vision but suffer from large parameter counts and quadratic computational complexity, severely limiting their deployment on resource-constrained edge hardware. While recursive weight-sharing reduces parameter counts and token merging mitigates computational and memory bottlenecks, integrating these two paradigms without costly retraining is non-trivial, leaving this...

    arxiv.org/abs/2608.13141 · PDF

  12. 12

    EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal Memory

    Weitao Chen, Hu Jiaxin, Xie Tianyidan, Yang Li, Yuyi Qian, Banghao Xu, Ziheng Tang, Shenyi Wang, Mingyue Yu, Duo Li,...

    cs.CV · cs.AI

    Recent advances in Multimodal Large Language Models (MLLMs) have led to substantial progress in video understanding, accompanied by a growing number of long video benchmarks. However, existing benchmarks rely predominantly on web-sourced videos that lack inter-clip spatiotemporal continuity, making it difficult to assess whether models can maintain consistent memory across days or weeks of real-world experience. We introduce EgoMonth, the...

    arxiv.org/abs/2608.13113 · PDF

  13. 13

    UniTraffic-Agent: Unified Traffic Video Reasoning for AI City Challenge 2026 Track 3 with Two Out-of-Domain Evaluations

    Peng Li, Qianqian Xu, Shilong Bao, Yangbangyan Jiang, Qingming Huang

    cs.CV · cs.AI

    Traffic video understanding has become an important problem in intelligent transportation, as road videos provide direct evidence for accidents, violations, and interactions between vehicles and vulnerable road users. A useful system should explain how a traffic event develops, why it happens, and when the relevant interaction occurs, yet this remains difficult for multimodal large language models (MLLMs) because traffic videos contain sparse...

    arxiv.org/abs/2608.13031 · PDF

  14. 14

    NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents

    Peng Cai, Zhaofan Zou, Shifa Liu, Yikun Wang, Jiawei Tang, Kaicheng Yang, Meng Tong, Zhongjiang He, Hao Sun

    cs.CV · cs.AI

    Document parsing aims to transform unstructured documents into structured and machine-readable representations. Recent advances in Vision-Language Models (VLMs) have significantly advanced document parsing. However, existing approaches still face two major challenges. First, decoupled VLM-based methods heavily rely on accurate layout analysis, where geometric distortions in camera-captured documents can introduce cascading errors. Second,...

    arxiv.org/abs/2608.12898 · PDF

  15. 15

    SPARED: Reasoning-Based AI-Generated Image Detection via Adversarially Edited Data

    Yicheng Bao, Xiahui Guo, Xuhong Wang, Xin Tan

    cs.CV · cs.AI

    Detecting AI-generated images is only half the task: a deployed detector must also justify its verdict, yet existing detectors inherit three failure modes from their training data: real and fake images collected from different sources invite provenance shortcuts, supervised explanation corpora teach templated rationales, and a static forgery corpus leaves the decision boundary standing still while generators keep moving. We introduce...

    arxiv.org/abs/2608.12876 · PDF

  16. 16

    Heterogeneous Vision-Language Ensemble with Disagreement-Aware Reranking for Text-Based Person Anomaly Retrieval

    Huu-An Vu, Cam Tu Tran Thi, Thanh Toan Le Ngo, Hoang Vo, Do Trung Hieu, Hieu Dinh Trung Pham, Khang Minh Le, Huy...

    cs.CV · cs.AI

    Text-based person anomaly retrieval aims to retrieve pedestrians exhibiting anomalous behaviors from a large image gallery using natural language descriptions. Compared with conventional text-based person retrieval, this task requires fine-grained reasoning over pedestrian appearance, behaviors, object interactions, and scene context, making robust cross-modal matching significantly more challenging. This paper presents the GENAI4E team's...

    arxiv.org/abs/2608.12843 · PDF

  17. 17

    Erase but Preserve: Controllable Removal of Copyrighted Animation Characters via Optimized Semantic Anchors

    Qiao Li, Xiaomeng Fu, Wangjia Yu, Runze He, Baisen Wang, Jiao Dai, Jizhong Han

    cs.CV · cs.AI

    The exceptional generation capabilities of text-to-image diffusion models have raised copyright concerns, particularly the unauthorized reproduction of animation characters. Existing concept erasure methods fall short for animation character erasure: model modification methods struggle to identify suitable anchors for diverse, highly distinctive characters; prompt-based steering methods lack fine-grained control for precise intervention....

    arxiv.org/abs/2608.12806 · PDF

  18. 18

    CW-BASS v2: Saturation-Aware Pseudo-Label Selection for Semi-Supervised Segmentation under Foundation-Model Teachers

    Ebenezer Tarubinga

    cs.CV · cs.LG · eess.IV

    Semi-supervised semantic segmentation has long turned on one question, which pseudo-labels to trust, and a generation of selection rules, dynamic thresholds, per-class curricula, soft confidence weights, answered it for the noisy, under-confident ResNet teachers of their day. Self-supervised foundation encoders change the regime: with a DINOv2 teacher, confidence saturates, so the filtering that helped a weak teacher can hurt a strong one. We...

    arxiv.org/abs/2608.12773 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.