cs.CV · 2026-09-11 · No. 112

Computer Vision and Pattern Recognition, 2026-09-11.

19 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

19 entries
  1. 01

    3D Point Splatting for mmWave Radar Novel View Synthesis

    Adnan Armouti, Yixuan Gao, Rajalakshmi Nandakumar

    cs.CV · cs.GR · cs.LG · eess.SP

    Solving novel view synthesis (NVS) for millimeter-wave (mmWave) radar requires a renderer that is physically faithful, complex-valued, and multi-viewpoint-tractable. No prior method achieves these three properties simultaneously. Differentiable Monte Carlo (MC) ray tracers implement the radar forward model directly with explicit material modeling and complex outputs, but do not scale to the multi-view optimization NVS demands. Optical-NVS...

    arxiv.org/abs/2609.11894 · PDF

  2. 02

    Logit Refiner: Improving Visual Autoregressive Models via Intra-Scale Dependency Modeling

    Meimingwei Li, Stefan Andreas Baumann, Felix Krause, Björn Ommer

    cs.CV · cs.AI · cs.LG

    Visual Autoregressive Models (VAR) generate images through next-scale prediction, producing all tokens within each scale in parallel. We show that this parallel decoding constitutes a mean-field-style approximation that discards spatial dependencies among same-scale tokens, causing locally incoherent samples regardless of backbone capacity -- a limitation of the decoding rule. Addressing this limitation, we introduce the Logit Refiner, a...

    arxiv.org/abs/2609.11804 · PDF

  3. 03

    Language-Augmented Semantic Priors for B-Spline Surface Fitting

    Yunzhong Lou, Yusheng Luo, Jiahao Li, Yu Song, Xiangdong Zhou

    cs.CV · cs.AI

    The use of B-splines and Non-Uniform Rational B-Splines surfaces constitutes the mathematical foundation of contemporary computer-aided design (CAD) systems. Despite long-term progress, geometric kernels in traditional CAD still rely heavily on predetermined heuristic initialization for surface fitting and parameterization. Meanwhile, the procedural semantics and design intent encoded in modeling histories are largely ignored during geometry...

    arxiv.org/abs/2609.11708 · PDF

  4. 04

    Multimodal Taxonomic Conditioning for Generative Plankton Imagery

    Daniela Ivanova, Ozgu Goksu, Nicolas Pugeault

    cs.CV · cs.LG

    Automated plankton imaging produces severely long-tailed datasets, where the rare taxa of greatest ecological interest have too few images to train or evaluate classifiers reliably. We generate synthetic plankton imagery conditioned on taxonomy: a CLIP encoder is adapted on a large plankton corpus with a ranked contrastive objective extended to deep, ragged taxonomies, then frozen to condition a parameter-efficient diffusion transformer. We...

    arxiv.org/abs/2609.11673 · PDF

  5. 05

    Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation

    Jintao Zhang, Kai Jiang, Jintao Chen, Xu Wang, Deyuan Liu, Jungang Li, Dechuang Chen, Ming Lin, Jingjiang Zhou,...

    cs.CV · cs.LG

    We present Vidu S2, which comprises Vidu S2-Avatar, a real-time interactive digital-character model, and Vidu S2-Editing, a real-time video editing model. Moreover, we explore the feasibility of real-time spatial video generation for both Vidu S2-Avatar and Vidu S2-Editing. Compared with Vidu S1, Vidu S2-Avatar supports real-time 720p video generation, generation with dynamic references that can be updated at any moment, and stronger...

    arxiv.org/abs/2609.11638 · PDF

  6. 06

    Learn the Solid, Not the File: Canonical Inputs for Neural Networks on CAD Boundary Representations

    Heinrich Jiang, Hager Yasser Mohamed, Alexander Hitt, Valeriia Lomakina, Henning Jiang, Jennifer Jang

    cs.CV · cs.AI · cs.CG

    Boundary representation (B-rep) is the standard format used by modern CAD systems for parametric 3D models. It turns out, the exact same solid can be represented by different B-reps: for example, two engineers using different operations, a geometry kernel rebuilding the file, and an export setting repartitioning faces will lead to different B-reps even though the underlying solid remains the same. We show that existing B-rep encoders are not...

    arxiv.org/abs/2609.11573 · PDF

  7. 07

    A Comparative Evaluation of Pre-trained Convolutional Neural Networks for Melanoma Detection

    Wagner Moreno Schmitz, Marco Antonio de Castro Barbosa, Thiago Magalhães Amaral, Dalcimar Casanova, Jefferson Tales Oliva

    cs.CV · cs.AI

    Early diagnosis of melanoma is critical for improving patient survival rates. However, accurately distinguishing melanoma from other skin lesions remains a significant clinical challenge due to the high visual similarity among lesion types and variability in image acquisition conditions. Artificial intelligence, particularly machine learning, has emerged as a promising tool to support dermatological diagnosis by automating feature extraction...

    arxiv.org/abs/2609.11550 · PDF

  8. 08

    Learning Interaction between Image and Layout Priors for Joint Image-Layout Generation in Design Templates

    Shirong Yang, Bo Yang, Ying Cao

    cs.CV · cs.AI · cs.GR

    In this paper, we address the problem of graphic design template creation, which generates a background image and a layout of foreground elements over the background to form a harmonious composition from an input text. Prior work on graphic design generation mostly adopts a sequential paradigm, where design elements are generated sequentially. We argue that such a sequential scheme falls short of faithfully capturing the dependency between...

    arxiv.org/abs/2609.11519 · PDF

  9. 09

    GRIPNet: Gaussian Radial Intensity Prior Guided Architecture for Pulmonary Nodule Detection in CT

    Haojie Yang, Ran Su

    cs.CV · cs.AI

    Lung cancer causes more deaths than any other malignancy, and low-dose CT screening is the main pathway to early diagnosis. That pathway hinges on the smallest lesions, yet nodules below six millimeters remain hard to detect, because most methods treat a nodule as a generic object and ignore the imaging physics behind its appearance. We show that this appearance is highly regular. Intensity peaks at the geometric center of a nodule and decays...

    arxiv.org/abs/2609.11312 · PDF

  10. 10

    Your Model Already Knows Don't Teach It, Learn to Ask It: Soft Prompting for Few-Shot Adaptation of Vision-Language Models

    Gautam Rajendrakumar Gare, Siyi Li, Hewei Wang, Cesar Daniel Hernandez, Wei Zhao, Wolfgang M. Pauli, John Galeotti,...

    cs.CV · cs.AI · cs.LG · eess.IV · stat.ML

    We address few-shot object detection with vision-language models (VLMs) in out-of-domain settings such as aerial, industrial, and medical imagery, using only ten annotated images for supervision. Existing adaptation methods are discrete prompt optimization and LoRA fine-tuning. We revisit a third option: soft prompting, where a small number of continuous prompt tokens are optimized while the pretrained backbone remains frozen. We identify two...

    arxiv.org/abs/2609.11310 · PDF

  11. 11

    Improving Faint Object Detection for Space Situational Awareness with Variational Autoencoders

    Angela Cratere, Luca Ghilardi, Vishnu Reddy, Francesco Dell'Olio, Charalampos S. Kouzinopoulos, Roberto Furfaro

    cs.CV · cs.AI · cs.LG

    We present a deep-learning pipeline for enhancing the detection of faint moving objects in optical space situational awareness (SSA) imagery through automated star removal and background reconstruction. Detecting low signal-to-noise ratio (SNR) objects remains extremely challenging in optical observations, particularly in the cislunar (X-GEO) environment, where structured sky backgrounds, dense stellar fields, and scattered moonlight...

    arxiv.org/abs/2609.11269 · PDF

  12. 12

    From Evaluation to Enhancement: Benchmarking and Improving Think-with-Video Reasoning for Video Generative Models

    Meng Luo, Yicheng Liu, Jiahao Wang, Yuanxing Zhang, Xin Tao, Pengfei Wan, Kun Gai, Hao Fei

    cs.CV · cs.AI

    Video generation has advanced to produce visually compelling and temporally coherent results. Yet, whether these models can genuinely think with video--executing symbolic rules, respecting physical laws, and pursuing intentional goals--remains an open question. Existing benchmarks only partially address this, often conflating visual quality with cognitive correctness. We introduce VWG-Bench (Video World Generalist Benchmark), a comprehensive...

    arxiv.org/abs/2609.11242 · PDF

  13. 13

    HALDETECT at ImageEval 2026 Shared Tasks: Answer-First Contrastive Grounding with QLoRA

    Syed Mohaiminul Hoque, Md Sakhawat Hossain

    cs.CV · cs.AI

    Large multimodal models tend to hallucinate visual detail fluently, which limits their deployment for fine-grained interpretation. We present HALDETECT, our system for the English hallucination-detection track (Task 1b) of ImageEval 2026, in which a system must identify, from an image and three culturally plausible statements, the single visually grounded one. We frame the item as one contrastive decision, emit the answer before its...

    arxiv.org/abs/2609.11236 · PDF

  14. 14

    Beyond Visual Quality: Evaluating Physical Consistency under Ego-Motion with EgoGenEval

    Yilin Long, Chenming Zhu, Zitang Gou, Jingli Lin, Tai Wang

    cs.CV · cs.AI

    Recent visual generators produce high-fidelity images yet often violate physical consistency under ego-motion, limiting their use for spatial reasoning and embodied planning. Existing benchmarks largely focus on isolated images or single-step quality, leaving this challenge underexplored. We introduce EgoGenEval, a geometry-grounded, pose-free benchmark designed to evaluate the physical consistency of visual generators under ego-motion, and...

    arxiv.org/abs/2609.11172 · PDF

  15. 15

    Beyond Benchmarks: Using VLMs to Reveal Systematic Classification Failures Under Real World Conditions

    Dieuwertje Alblas, Alma M. Liezenga, Jan Erik van Woerden, Fedor Taggenbrock, Dalia Aljawaheri, Klamer Schutte

    cs.CV · cs.AI · cs.ET

    Verification and validation (V&V) of classification models is crucial to enable a wide range of sensor processing applications. Currently, the V&V process relies on time-consuming manual inspection of erroneous samples to find meaningful patterns. This work explores the use of Vision Language Models (VLMs) to speed up this laborious process. VLMs are trained to embed images into a semantically meaningful vector representation, from which...

    arxiv.org/abs/2609.11126 · PDF

  16. 16

    TailProp: content-adaptive light- and heavy-tailed propagation for vision

    Jiahao Kong, Zihan Li

    cs.CV · cs.LG

    Science-inspired vision models show that explicit propagation dynamics can provide structured and interpretable alternatives to conventional token mixing. Existing formulations, however, typically construct and adapt visual propagation within a particular dynamical family, while visual representations can require substantially different spatial interactions across samples, channels, and network stages. We explore cross-regime adaptive...

    arxiv.org/abs/2609.11081 · PDF

  17. 17

    Meta-Learning for Classifier Selection in Image Datasets: A Feature-Driven Framework for Accuracy Prediction

    Zahra Nabizadeh_Shahre_Babak, Farzaneh Koohestani, Nader Karimi, Shahram Shirani, Shadrokh Samavi

    cs.CV · cs.LG

    No Free Lunch theorem implies that any performance gains achieved by a classifier on a particular image distribution are necessarily offset by a loss of performance over the set of all possible problems; thus, no single model is universally optimal. Selecting the most suitable classifier for image datasets is a critical yet challenging task due to the intrinsic complexity and diversity of images. This paper proposes a meta-learning framework...

    arxiv.org/abs/2609.11041 · PDF

  18. 18

    Toward Interpretable Multimodal Fusion: Heat Conduction Modeling for Hyperspectral and LiDAR Joint Classification

    Kan Wei, Jiahui Cui, Jing Yao, Xinyu Zhao, Lei Wang, Pedram Ghamisi

    cs.CV · cs.AI

    The fusion of hyperspectral (HS) and Light Detection and Ranging (LiDAR) data plays a crucial role in enhancing land-cover classification by jointly exploiting spectral, spatial, and structural cues. However, existing multimodal fusion methods still struggle to model long-range dependencies and complex anisotropic interactions while maintaining computational efficiency. This paper introduces M2Heat, a physics-inspired framework that...

    arxiv.org/abs/2609.11040 · PDF

  19. 19

    New Evidence, Same Choice: Testing Physical Experiment Selection in Vision Language Models

    Sourajit Saha, Shubhashis Roy Dipta, Nobin Sarwar, Shaswati Saha, Yuxuan Jiang, Siyuan Li, Qiheng Wang

    cs.CV · cs.AI · cs.CL · cs.LG

    A model first sees an image from one physical measurement experiment, such as how far a block coasted, and must answer a question about a new trial, such as whether the block will pass a target after a fixed push. The initial experiment may provide enough information to answer, or the model may need another measurement, such as the object's mass, friction, restitution, or spring stiffness. We study whether vision language models can decide...

    arxiv.org/abs/2609.11022 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.