cs.RO · 2026-09-16 · No. 115

Robotics, 2026-09-16.

10 new papers in cs.RO. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

10 entries
  1. 01

    FluxVLA Engine: A One-Stop VLA Engineering Platform for Embodied Intelligence

    Yinhao Li, Weixin Mao, Zihan Lan, Jikun Rong, Qirui Hu, Yiming Zhang, Weipeng Deng, Bowen Shen, Minzhao Zhu, Yiming...

    cs.RO · cs.AI

    Vision-language-action (VLA) models, world-action models (WAMs), and offline reinforcement learning methods are rapidly expanding the design space of embodied policies, yet turning these algorithms into reliable robot systems remains constrained by fragmented data formats, training stacks, evaluation protocols, inference runtimes, and embodiment-specific interfaces. We present $\mathrm{FluxVLA}$ Engine, an open, configuration-driven platform...

    arxiv.org/abs/2609.17210 · PDF

  2. 02

    Kernel-Based Metrics Learning for Uncertain Opponent Vehicle Trajectory Prediction in Autonomous Racing

    Hojin Lee, Youngim Nam, Sanghun Lee, Cheolhyeon Kwon

    cs.RO · cs.AI · cs.LG

    Autonomous racing confronts significant challenges in safely overtaking Opponent Vehicles (OVs) that exhibit uncertain trajectories, stemming from unknown driving policies. To address these challenges, this study proposes heterogeneous kernel metrics for Deep Kernel Learning (DKL), designed to robustly capture the diverse driving policies of OVs, and carry out precise trajectory predictions along with the associated uncertainties. A key...

    arxiv.org/abs/2609.17147 · PDF

  3. 03

    Continual Learning for Traversability Prediction with Uncertainty-Aware Adaptation

    Hojin Lee, Yunho Lee, Daniel A Duecker, Cheolhyeon Kwon

    cs.RO · cs.AI · cs.LG

    Traversability prediction is a critical component of autonomous navigation in unstructured environments, where complex and uncertain robot-terrain interactions pose significant challenges such as traction loss and dynamic instability. Despite recent progress in learning-based traversability prediction, these methods often fail to adapt to novel terrains. Even when adaptation is achieved, retaining experience from previously trained...

    arxiv.org/abs/2609.17141 · PDF

  4. 04

    Intrinsic Robot Rewarding: Reusing VLA Representations for Autonomous Evaluation and Policy Improvement

    Tobias Schaffer, Mohab Elkhayat, Daniela Nicklas, Mustafa Almohamad, Elham Al-Fuqara

    cs.RO · cs.LG

    Vision-language-action (VLA) systems already bring together two valuable resources for robot learning: rich visual representations and demonstrations of successful task execution. Intrinsic Robot Rewarding (IRR) proposes to use these resources for a second, complementary purpose: evaluating the robot's own outcomes and providing feedback for policy improvement. Successful demonstration endpoints define task-specific references, and the...

    arxiv.org/abs/2609.17115 · PDF

  5. 05

    TEMPO: Learning Temporal Context for Dynamic Robot Manipulation

    Zhenyang Feng, Jimin Heo, Erik B. Sudderth, Unnat Jain

    cs.RO · cs.CV · cs.LG

    Vision-language-action (VLA) models have achieved impressive performance in quasi-static manipulation, but struggle in dynamic manipulation tasks because they operate on a single observation at inference time. We identify two representational failures that underlie this limitation. The first is motion ambiguity, where a single observation does not include scene dynamics and therefore cannot anticipate the future state of moving objects. The...

    arxiv.org/abs/2609.16864 · PDF

  6. 06

    The Latent That Never Was: A Forensic Re-run of the CVAE Ablation in Action Chunking Transformer

    Bo Kang

    cs.RO · cs.LG

    Action Chunking Transformers (ACT) are widely used to learn robot manipulation from demonstrations. Their conditional variational autoencoder includes an encoder meant to capture differences between demonstrations during training. The original ACT paper reported that encoder removal dropped the mean success rate from 35% to 2% on two simulated tasks with human demonstrations. We re-ran this ablation in the original code and checked whether...

    arxiv.org/abs/2609.16745 · PDF

  7. 07

    Seeing What Matters: Visual Cue Guided Video Planning for Generalizable Robot Navigation

    Hojin Lee, Sizhe Lester Li, Maximilian Hilger, Susie Lu, Achim J. Lilienthal, Vincent Sitzmann, Daniel A. Duecker

    cs.RO · cs.AI · cs.CV · cs.LG

    Generative video models can serve as a promising backbone for robot navigation by predicting future observations as video plans. Recent approaches often condition video planning on short-horizon guidance and recover geometric waypoints through scene reconstruction, leaving longer-horizon planning and precise video-to-action translation less explored. We present CueNav, a video model-based navigation framework combining visual cue guided video...

    arxiv.org/abs/2609.16737 · PDF

  8. 08

    World Models for Embodied Intelligence: From Plausible to Controllable to Actionable

    Nanjie Yao, Hao Wang, Chong Cheng, Zhikang Chen, Wenzhe Li, Jiafei Lyu, Li Shen, Peilin Zhao, Zongqing Lu, Gao...

    cs.RO · cs.AI

    World models connect perception and decision-making in embodied intelligence by maintaining hidden state, anticipating consequences, comparing interventions, and adapting when execution departs from expectations. Although progress is often measured by visual fidelity, their value lies in improving behavior. Before reaching for a cup, a person anticipates its weight and resistance to grasping, shaping the hand before contact. Such anticipation...

    arxiv.org/abs/2609.16697 · PDF

  9. 09

    Weave: Learning Whole-Body Dexterous Loco-Manipulation from Human-Object Interactions

    Liu Cao, Xingze Wu, Jingzhi Cui, Botian Xu, Mingzhi Pei, Ruoqu Chen, Mengdi Xu

    cs.RO · cs.AI · cs.LG

    Learning humanoid-object interaction requires coordinating whole-body balance, locomotion, and dexterous hand contact to control both robot and object motion. Human demonstrations provide examples of coordinated interaction, but transferring these behaviors to humanoid robots requires learning how to establish and maintain effective contacts under different embodiments and dynamics. We present Weave, a unified framework for learning...

    arxiv.org/abs/2609.16683 · PDF

  10. 10

    ProxiDex: Learning Dynamics-Guided Proximity Policy for Dexterous Manipulation

    Yushan Bai, Boyu Zheng, Zhiyang Mao, Hongzheng Sun, Yuchuang Tong, En Li, Zhengtao Zhang

    cs.RO · cs.AI

    Multi-finger dexterous manipulation relies on stable hand-object interactions, yet these interactions are partially observable in practice. Visual observations are often occluded by the hand, tactile sensors introduce hardware-specific modalities and calibration burdens, and existing policies rarely model how these cues evolve under actions, making them brittle under contact uncertainty. To address these, we present ProxiDex, a...

    arxiv.org/abs/2609.16586 · PDF

This edition is part of The Daily Abstract — cs.RO archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.