cs.CV · 2026-09-20 · No. 119

Computer Vision and Pattern Recognition, 2026-09-20.

14 new papers in cs.CV. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

14 entries
  1. 01

    FAMOS: Feed-Forward 3D Articulation Modeling from Sparse Observations

    Kevin Qu, Tao Sun, Massimiliano Viola, Liyuan Zhu, Zhizhuo Zhou, Sayan Deb Sarkar, Konrad Schindler, Iro Armeni

    cs.CV · cs.AI · cs.RO

    Modeling articulated objects from sparse monocular views is challenging because each observation reveals only partial geometry and motion evidence. Most feed-forward methods infer articulation from a single observation and therefore rely heavily on learned category-level shape priors. We present FAMOS, a feed-forward model that predicts movable-part segmentation and joint parameters from a sparse, unordered set of partial point clouds. Our...

    arxiv.org/abs/2609.20817 · PDF

  2. 02

    Paint-Anything: Unified Any-Color Control for Image Generation and Editing

    Ji Xie, Dewei Zhou, Xinyu Huang, Zhennan Chen, Xun Wang

    cs.CV · cs.AI · cs.LG

    Professional design requires any-color control: the ability to specify an object's target color with any 24-bit hex value for image generation and editing. Prior work has explored color generation, editing, and colorization, but often relies on dedicated color representations or specialized inference procedures. Advances in large language models offer a simpler starting point: even compact models can associate hex values with color semantics....

    arxiv.org/abs/2609.20816 · PDF

  3. 03

    ERCPMP-Gx: Endoscopic Image and Video Dataset for Morphological, Histopathological, and Genomic Characterization of Colorectal Polyposis

    Zahra Ghaffari, Massih Bahar, Mojgan Forootan, Ali Darvishi, Hamidreza Bolhasani

    cs.CV · cs.AI

    Hereditary polyposis syndromes can be precursor lesions to colorectal cancer and are associated with a broad spectrum of extracolonic tumors. Early identification and accurate classification of these syndromes are essential for timely diagnosis, individualized patient management, and targeted surveillance strategies for affected families. However, public endoscopic datasets are largely organized around the individual sporadic polyp, and none...

    arxiv.org/abs/2609.20815 · PDF

  4. 04

    Cross-Architecture Foundation-Model Distillation for Edge Flood Segmentation

    Fabian Schmalstieg, Karsten Mueller, Wojciech Samek

    cs.CV · cs.LG

    Geospatial foundation models can provide strong flood-segmentation performance, but their size limits deployment on memory-constrained edge hardware. We distill a 300-million-parameter Prithvi-EO-2.0 teacher, fine-tuned on the 252 manually labeled Sen1Floods11 training scenes, into a 0.7-million-parameter EfficientViT-B0 student. The teacher supervises additional unlabeled Sentinel-2 imagery, allowing the student training set to grow without...

    arxiv.org/abs/2609.20441 · PDF

  5. 05

    When Do Language-Grounded Explanations Help? A Graph-Bottleneck for Farm Monitoring Interpretable Sheep Facial Pain

    Alam Noor, Miguel Guti'errez Gait'an

    cs.CV · cs.AI

    Automated pain recognition from facial expression could make continuous welfare assessment practical in sheep, but adoption depends on trust: a stockperson cannot act on a score that arrives without justification. We ground a model in the Sheep Pain Facial Expression Scale (SPFES) by letting each detected facial region attend over text embeddings of the clinical descriptors and then test whether the resulting explanations mean anything. They...

    arxiv.org/abs/2609.20427 · PDF

  6. 06

    TouchSight: Bare-Handed Tactile Prediction from Egocentric Video via Generative Visual Augmentation

    Danyan Zhou, Jinxuan Lu, Jiawei Lin, Tianxing Chen, Chuqiao Lyu, Wenbo Ding

    cs.CV · cs.AI

    Tactile signals provide direct contact and force measurements that are essential for understanding physical interactions and enabling dexterous robotic manipulation. However, tactile sensing requires direct measurement at contact interfaces, making large-scale data collection reliant on intrusive, costly, and restrictive instrumentation. We present TouchSight, a monocular egocentric vision framework for dense full-hand contact force...

    arxiv.org/abs/2609.20414 · PDF

  7. 07

    Fast Cross-Strength Multi-Contrast Brain MRI Translation using Latent Bridge Matching

    Siddharth Srivastava, Till Bretschneider

    cs.CV · cs.LG

    Magnetic Resonance Imaging (MRI) acquired at different field strengths exhibits pronounced variation in noise, resolution, homogeneity, and contrast, which limits comparability across acquisition settings and complicates downstream analysis. We address this with a unified conditional model for controllable field-to-field synthesis, built on the framework of conditional latent bridge matching. Our single model achieves highly competitive...

    arxiv.org/abs/2609.20341 · PDF

  8. 08

    Task-Oriented Semantic Feature Transmission for Multi-Task Satellite Remote Sensing over Low-SNR Channels

    Shuoyuan Sun, Hongyu Wang, Mugen Peng, Wenjia Xu

    cs.CV · cs.LG

    Conventional satellite remote sensing transmission follows a reconstruct-then-infer paradigm that optimizes pixel-level fidelity, creating an objective mismatch with downstream tasks such as classification and detection, especially at low SNR. This paper investigates a task-oriented framework that bypasses image reconstruction and directly transmits semantic features extracted by a multitask-pretrained backbone. A lightweight channel...

    arxiv.org/abs/2609.20150 · PDF

  9. 09

    Bridging Modalities on the Cortex: Surface-based MRI to PET Translation with a Diffusion Bridge

    Yitong Li, Alexandra Samoylova, Fabian Bongratz, Timo Grimmer, Dennis M. Hedderich, Igor Yakushev, Christian Wachinger

    cs.CV · cs.AI

    Cortical hypometabolism measured by Fluorodeoxyglucose Positron Emission Tomography (FDG-PET) is a highly sensitive biomarker for dementia diagnosis. However, high costs, radiation exposure, and limited accessibility constrain its clinical utility. While cross-modal synthesis from Magnetic Resonance Imaging (MRI) offers a promising alternative, existing volumetric generation methods do not explicitly account for the highly folded cortical...

    arxiv.org/abs/2609.20147 · PDF

  10. 10

    Cross-Modal Attention Acts as a Frequency Filter: Why Verbose Prompts Improve Robustness in Vision-Language Models

    Farooq Ahmad Wani, Maria Sofia Bucarelli, Mujtaba Hussain Mirza, Oleksandr Pryymak, Aryo Pradipta Gema, Iacopo Masi,...

    cs.CV · cs.AI

    Vision-language models (VLMs) are fragile under image corruption. We find that the wording of the question affects VLMs in two opposite ways. Verbose questions make VLMs substantially more robust---e.g., rephrasing "Is there a cat?" into "Please look carefully and answer: is there a cat?". Conversely, VLMs become more fragile under corruption when the question is semantically complex or finer-grained, e.g., "what colour is the cup left of the...

    arxiv.org/abs/2609.20139 · PDF

  11. 11

    PointEvent: Rethinking Event-based Tiny Object Detection via Serialized Motion Evidence Accumulation

    Zongze Wu, Baofeng Jia, Weiqi Yan, Jingyuan Zhang, Yu Zang, Xiaoyu Chen, Jing Han

    cs.CV · cs.AI

    Event cameras offer high temporal resolution and motion sensitivity for tiny UAV detection, yet distant targets generate sparse and fragmented events that are easily overwhelmed by clutter and ego-motion. Existing methods mainly rely on dense event representations or local sparse spatiotemporal modeling, resulting in redundant computation or fragmented modeling of motion continuity across distant asynchronous events. To address this...

    arxiv.org/abs/2609.20066 · PDF

  12. 12

    Astronex-World 1.0: Real-Time Interactive World Model Foundation

    Xin Zhou, Cong Miao

    cs.CV · cs.AI · cs.RO

    We present Astronex-World 1.0, an open controllable video world-model foundation. Given a text prompt (text-to-video) or an initial observation (image-to-video), the model predicts future visual states under frame-aligned camera trajectories, continuous actions, and an embodiment identifier, and accepts text events inserted at a specified position of a rollout. The family provides a bidirectional model for full-context generation and a causal...

    arxiv.org/abs/2609.20034 · PDF

  13. 13

    AVTrace: Diagnosing Audio-Visual Temporal Reasoning in Omni Models

    Longyin Zhang, Parth Sakhare Mahendra, Chengwei Wei, Ning Zhang, Lim Ming Chong, Sirui He, Ai Ti Aw

    cs.CV · cs.AI

    Omni models can describe video content, but can they locate events in time, preserve event order, and judge audio-visual synchronization? We introduce AVTrace (Audio-Visual Temporal Reasoning Assessment and Capability Evaluation), a silver-standard diagnostic suite spanning onset and span grounding, synchronization, next-step prediction, cross-modal localization, chain parsing, and event-conditioned comprehension. It contains 34,114 training...

    arxiv.org/abs/2609.19991 · PDF

  14. 14

    PACE: Precise AI Cinematic Expression: A Typed Specification for Script-Grounded Previsualization and Geometric Conformance

    Bing Duan, Qiang Guo, Linpu Li, Zhijian Mao, Min Zhu, Zhirui Ren, Yiwei Yan, Xi Chu, Xiaoding Li

    cs.CV · cs.AI

    Between a screenplay and a film sits a planning problem that is spatial first: who stands where, and what a camera sees from where it stands. An image diffusion model asked for a shot in free text settles that plan by its own defaults. We present PACE (Precise AI Cinematic Expression), a typed representation for the plan: the screenplay evidence, the characters, props and locations it needs, where each subject stands, and what the camera...

    arxiv.org/abs/2609.19853 · PDF

This edition is part of The Daily Abstract — cs.CV archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.