cs.CV · 2026-09-28 · No. 127

Computer Vision and Pattern Recognition, 2026-09-28.

19 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

19 entries
  1. 01

    OC-GS: Gaussian Splatting for Irregular Turntable Capture

    Jae Joong Lee, Bedrich Benes

    cs.CV · cs.AI

    Uneven rotation and dropped frames make equal-angle assumptions unreliable for turntable reconstruction. We present OC-GS, an object-centric Gaussian splatting that refines each image's angle while maintaining a shared camera, rotation axis, and pivot. This orbit-consistent refinement jointly optimizes image-derived geometry and angles to reconstruct objects from sparse, irregular captures. On rendered objects with 12, 8, and 6 irregularly...

    arxiv.org/abs/2609.31572 · PDF

  2. 02

    ClearGS: Reliability-Aware Gaussian Splatting from Handheld Videos

    Xuanzhi Liu, Xinyi Wu, Hang Pan, Wensi Huang, Zhenyao Wu, Ruize Han, Song Wang

    cs.CV · cs.AI

    We present ClearGS for 3D Gaussian Splatting (3DGS) from handheld videos with uneven viewpoint coverage and mixed frame quality. Rather than selecting frames with binary decisions, ClearGS uses Reliability-aware View Allocation (RVA) to assign graded raw-supervision weights based on appearance reliability, degradation risk, and geometric utility, while weakly reactivating useful suppressed frames to maintain trajectory coverage. Since...

    arxiv.org/abs/2609.31509 · PDF

  3. 03

    From Reward Signal to Visual Utility: A Controlled Audit of Medical VLM Post-Training

    Wang Jingxin

    cs.CV · cs.AI

    Medical vision-language model (VLM) post-training is commonly evaluated through answer accuracy. We examine how changes in accuracy and training objectives relate to image-conditioned decisions in a controlled Qwen2.5-VL-3B study on PMC-VQA. We compare supervised fine-tuning (SFT) with low-rank adaptation (LoRA) restricted to the language model, expanded multimodal adaptation scopes, standard answer-only Group Relative Policy Optimization...

    arxiv.org/abs/2609.31450 · PDF

  4. 04

    Implicit Neural Representation for Hyperspectral Video Compression

    Alfredo Scalera, Paul Murray, Jaime Zabalza

    cs.CV · cs.AI · cs.LG

    With the advent of snapshot cameras, hyperspectral video is becoming more readily available. In recent years, new applications have emerged which have led to increasingly larger datasets. However, hyperspectral video compression remains in the early stages. In this study, we explore the use of implicit neural representation as a candidate solution. We propose a novel extension of an existing RGB video compression model, achieving Bjøntegaard...

    arxiv.org/abs/2609.31435 · PDF

  5. 05

    Open Vocabulary Domain Unlearning

    Sumanth Udupa, Mehrtash Harandi, Yadan Luo, Mahsa Baktashmotlagh

    cs.CV · cs.LG

    Vision-Language Models (VLMs) exhibit remarkable zero-shot generalization, yet they often encode unwanted or hazardous stylistic domains such as idealized textbook diagrams in medical AI or cartoon vehicles in autonomous driving. Approximate Domain Unlearning (ADU) aims to selectively erase a model's recognition of a target visual domain while preserving accuracy on the remaining domains. However, existing ADU methods operate under a flawed...

    arxiv.org/abs/2609.31356 · PDF

  6. 06

    DyMD: Preserving Interaction Dynamics through Distribution Matching Distillation in Few-Step Video World Models

    Haojun Xu, Jie Huang, Xin Lu, Mingchen Zhong, Zihao Fan, Linjiang Huang, Si Liu

    cs.CV · cs.AI

    Large video diffusion models offer expressive priors for embodied prediction and learning, yet their many-step sampling remains costly for interactive downstream use. Distribution Matching Distillation (DMD) enables few-step video generation, but can suppress robot--object motion while preserving visual quality. Examining DMD's teacher and fake-score signals, we find that weak re-noising keeps the teacher posterior concentrated near...

    arxiv.org/abs/2609.31349 · PDF

  7. 07

    CG-HAF: An Interpretable Global-Local Lesion-Burden Fusion Framework for Ordinal Acne Severity Grading in Agentic Skincare Support

    Muhammad Muhtasim Shahriar, Md. Naimur Asif Borno, Saad Aloteibi, Mohammad Ali Moni

    cs.CV · cs.AI

    Ordinal acne severity grading requires distinguishing visually similar neighboring grades while jointly weighing holistic facial appearance and localized lesion burden - evidence that most existing approaches collapse into a single opaque representation. We introduce CG-HAF, a global-local fusion framework that instead keeps this evidence explicit: averaged holistic severity probabilities from independently trained classifiers are combined...

    arxiv.org/abs/2609.31326 · PDF

  8. 08

    UniAR: A Unified Framework for Autism Recognition Enhanced by Multi-View Prompt Learning

    Lei Xin, Zeheng Wang, Jiayin Zhu, Shihong Huang, Fanhu Zeng, Changjiang Jiang, Dengbo He, Yutao Yue, Zhenglun Kong

    cs.CV · cs.AI

    Autism Spectrum Disorder (ASD) is a complex neurodevelopmental disorder for which early and accurate diagnosis is critical to improving long-term developmental outcomes. However, existing ASD recognition methods are often constrained by the scarcity of diagnostic text data, forcing them to rely mainly on visual analysis and limiting their ability to model clinically meaningful semantic reasoning. To address this challenge, we propose UniAR, a...

    arxiv.org/abs/2609.31298 · PDF

  9. 09

    Geometric Inconsistency Localization in Multi-View Image Sets

    Xander Staelens, Albéric Loos, Bert Ramlot, Hannes Mareen, Peter Lambert, Glenn Van Wallendael

    cs.CV · cs.AI · cs.CR · cs.MM

    Novel view synthesis (NVS) models can produce realistic new views of the same scene from different viewpoints. However, these generated views are not always geometrically consistent with one another. Multi-view (MV) consistency has shown promise as a tool for evaluating these NVS models. Its potential for multimedia forensics, however, remains largely unexplored, particularly for localizing geometric inconsistencies across wide-baseline image...

    arxiv.org/abs/2609.31247 · PDF

  10. 10

    ReG-SAM: Reference Graph-Driven SAM for 2D Foundational Vessel Segmentation

    Donghang Lyu, Zichen Zhang, Oleh Dzyubachyk, Marius Staring

    cs.CV · cs.AI

    Vessel segmentation in medical images is essential for many clinical tasks, ranging from diagnosis to treatment planning. However, it remains challenging due to complex vascular morphology and diverse imaging conditions. Existing deep learning methods rarely aim at building a generalizable vessel segmentor across anatomies and modalities. While the Seg- ment Anything Model (SAM) has shown promise for med- ical image segmentation, its original...

    arxiv.org/abs/2609.31160 · PDF

  11. 11

    FedHisto-PAST: Parameter-Efficient Stain-Aware Federated Learning for Cross-Site Lung Histopathology Classification

    Muhammad Muhtasim Shahriar, M. M. Golam Hafiz, Saad Aloteibi, Mohammad Ali Moni

    cs.CV · cs.AI

    Cross-site lung histopathology classification must account for stain variation, non-IID client data, missing classes, and the cost of adapting large pathology encoders. This study evaluates FedHisto-PAST v2 for three-way classification of adenocarcinoma (ACA), Normal, and squamous cell carcinoma (SCC). FedHisto-PAST v2 combines a frozen HIBOU-B foundation model with parameter-efficient adaptation, stain-conditioned paired-view prediction and...

    arxiv.org/abs/2609.31150 · PDF

  12. 12

    Pocket-STVG: lightweight architecture for Spatio-Temporal Video Grounding

    Alberto Presta, Michal Byra, Grzegorz Stefański, Karol Szurkowski, Eryk Kołodziejczyk, Krzysztof Arendt

    cs.CV · cs.AI · cs.MM

    Spatio-Temporal Video Grounding (STVG) aims to localize the spatio-temporal tube in a video corresponding to a natural language query. While recent methods achieve strong performance in fully supervised, weakly supervised, and zero-shot settings, they typically rely on computationally expensive architectures, complex training pipelines, or multimodal large language models. We present Pocket-STVG (P-STVG), a lightweight cascade architecture...

    arxiv.org/abs/2609.31135 · PDF

  13. 13

    DepthEvidence: Unifying Metric Depth Prediction and Geometric Reasoning in Multimodal Language Models

    Jiangning Wei, Yuan Yao, Miaomiao Cui, Mingsheng Li, Humen Zhong, Shuai Bai, Zhibo Yang

    cs.CV · cs.AI

    Spatial reasoning with metric constraints requires linking objects to geometric measurements and preserving their numerical content during language reasoning. We present DepthEvidence, a 4B model that uses its own dense metric predictions as object-grounded evidence for language generation. A camera-conditioned decoder predicts full-resolution metric depth using multi-scale visual features and high-resolution RGB refinement. A...

    arxiv.org/abs/2609.31103 · PDF

  14. 14

    FLIP: Final Layer Inference-Time Probing for Vision-Language Models

    Drandreb Earl O. Juanico, Rowel O. Atienza

    cs.CV · cs.AI

    We present FLIP, a final-layer inference-time probe for testing whether a logit-facing intervention site in an open-weight vision-language model (VLM) supports structured, task-linked computation rather than generic perturbation. Behavioral change under internal intervention is otherwise mechanistically ambiguous: it may reflect improved use of visual evidence, generic output instability, or outright degradation. FLIP applies elementwise...

    arxiv.org/abs/2609.30993 · PDF

  15. 15

    FARE: Forensic Acceptance Region Estimation for Catching Bait-and-Switch Image Generators

    Kai Yao, Marc Juarez

    cs.CV · cs.AI · cs.CR

    Modern AI image generators are increasingly deployed as opaque APIs, where customers can query the deployed service, but cannot inspect model weights or architecture. This creates a practical challenge: a provider may pass governance certification with one generator and later silently switch to a cheaper and lower-quality one for deployment, compromising public trust or even safety in high-stakes domains. We study integrity auditing at...

    arxiv.org/abs/2609.30982 · PDF

  16. 16

    MVVBench: Benchmarking 4D Reasoning in Vision-Language Models

    Hyungjin Chung, Byeongjun Park, Joonseok Lee, Hojun Kim, Jaeho Choi, Byung-Hoon Kim

    cs.CV · cs.AI

    Multi-view video understanding requires integrating spatial and temporal evidence across multiple, often non-overlapping camera streams: tracking entities as they transition between viewpoints, aligning events across time, and reasoning about latent 4D continuity rather than any single visible frame. We introduce MVVBench, a benchmark for multi-view video reasoning built from real world multi camera datasets. Questions are curated to be...

    arxiv.org/abs/2609.30952 · PDF

  17. 17

    OneWorld: Learning Consistent Physics Across Actions in World Models

    Ke He, Yichen Ding, Bin Yang

    cs.CV · cs.AI

    Action-conditioned video world models aim to predict scene evolution under different actions, a capability that is essential for reliable planning, decision-making, and interaction in dynamic environments. However, futures generated independently from the same initial scene may each appear plausible while implying incompatible physical properties, such as friction or mass. This inconsistency can lead to contradictory predictions across...

    arxiv.org/abs/2609.30946 · PDF

  18. 18

    Spackle: Completing Large View Single Image NVS with Adaptive Gaussians

    Xuanzhi Liu, Yuhe Zhou, Xinyi Wu, Zhenyao Wu, Jinghao Chen, Ruize Han, Song Wang

    cs.CV · cs.AI

    Single-image novel view synthesis (NVS) enables photorealistic rendering of un- observed viewpoints from a single input. Practical NVS systems require two key capabilities: robust reconstruction of occluded regions and high inference effi- ciency. While hybrid decoupled frameworks combining feedforward 3D Gaussian Splatting (3DGS) and diffusion models show promise for large-view-deviation NVS, they suffer from capacity competition: a fixed...

    arxiv.org/abs/2609.30941 · PDF

  19. 19

    UltraG-Bench: A Multi-task Benchmark for assessing Large Vision-Language Models on Pixel-level Evidence Grounding in Ultrasound

    Quanhao Zhu, Bo Xu, Rui Lin, Chenyuan Wang, Yu Shao, Boling Zhu, Jiuyan Sun, Liang Zhao, Hongfei Lin, Feng Xia

    cs.CV · cs.AI

    Ultrasound is one of the most widely used medical imaging modalities, and recent large vision-language models(VLMs) have shown increasing capabilities in ultrasound image understanding. However, these models fail to provide pixel-level visual evidence aligned with their semantic predictions, and their fine-grained grounding capability in ultrasound remains largely unclear. We introduce UltraG-Bench, a large-scale multi-task benchmark for...

    arxiv.org/abs/2609.30928 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.