cs.RO · 2026-06-09 · No. 18

Robotics, 2026-06-09.

12 new papers in cs.RO. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

12 entries
  1. 01

    AHA-WAM:Asynchronous Horizon-Adaptive World-Action Modeling with Observation-Guided Context Routing

    Jisong Cai, Long Ling, Shiwei Chu, Zhongshan Liu, Jiayue Kang, Zhixuan Liang, Wenjie Xu, Yinan Mao, Weinan Zhang,...

    cs.RO · cs.AI · cs.CV

    World-action models have emerged as a promising paradigm for robot manipulation, jointly modeling visual scene dynamics and actions to inject physical priors into policy learning. However, existing world-action models couple world prediction and action execution at the same temporal resolution, forcing the world branch to model near-term frame variations that are redundant and weakly informative. We posit that strictly binding world...

    arxiv.org/abs/2606.09811 · PDF

  2. 02

    Difference-Aware Retrieval Policies for Imitation Learning

    Quinn Pfeifer, Ethan Pronovost, Paarth Shah, Khimya Khetarpal, Siddhartha Srinivasa, Abhishek Gupta

    cs.RO · cs.AI · cs.LG

    Parametric imitation learning via behavior cloning can suffer from poor generalization to out-of-distribution states due to compounding errors during deployment. We show that reusing the training data during inference via a semi-parametric retrieval-based imitation learning approach can alleviate this challenge. We present Difference-Aware Retrieval Policies for Imitation Learning (DARP), a semi-parametric retrieval-based imitation learning...

    arxiv.org/abs/2606.09758 · PDF

  3. 03

    Your Model Already Knows: Attention-Guided Safety Filter for Vision-Language-Action Models

    Seongbin Park, Fan Zhang, Baharan Mirzasoleiman, Shahriar Talebi, Nader Sehatbakhsh

    cs.RO · cs.LG

    Vision-Language-Action (VLA) models have demonstrated impressive end-to-end performance across a variety of robotic manipulation tasks. However, these policies offer no guarantees against collisions with task-irrelevant objects in the scene. Existing safety filters sidestep this problem by querying a vision-language model (VLM) to identify obstacles and their locations. This, however, is too slow to run in the control loop and can only be...

    arxiv.org/abs/2606.09749 · PDF

  4. 04

    ReCoVLA: VLM-Guided Reward Compilation for Failure Recovery in Vision-Language-Action Policies

    Haodi Hu, Chung-Ta Huang, Jing Liu, Ye Wang, Kei Suzuki, Matthew Brand, Toshiaki Koike-Akino

    cs.RO · cs.AI · cs.LG

    Vision-language-action (VLA) policies provide strong priors for language-conditioned manipulation, but remain brittle in off-nominal states requiring targeted recovery. We propose ReCoVLA -- a failure-conditioned residual recovery framework that keeps a pretrained VLA policy frozen, uses an external vision-language model (VLM) to infer the failure mode and recovery stage, and compiles a structured reward from task-relevant components. Rather...

    arxiv.org/abs/2606.09630 · PDF

  5. 05

    Shape Formation for the Cooperative Transportation of Arbitrary Objects Using Multi-Agent Reinforcement Learning

    Mohamed Sayed, Wolfram Burgard, Tanja Katharina Kaiser

    cs.RO · cs.AI

    Cooperative object transportation is essential in numerous domains, including industrial to domestic services. A popular transportation strategy is to carry objects on top of multi-robot systems. The corresponding task is typically solved by decomposing it into three interconnected subproblems: formation control, cooperative navigation, and collision avoidance. A particular challenge posed by real-world objects is their potentially arbitrary...

    arxiv.org/abs/2606.09610 · PDF

  6. 06

    CT-VAM: A Cerebello-Thalamic-Inspired Vision-Action Model for Efficient Visuomotor Control

    Jiacheng Li, Yize Guo, Jiabin Guo, Qingchen Liu, Jiahu Qin

    cs.RO · cs.AI

    Vision-language-action models have shown strong promise for robot manipulation, yet raw language is primarily needed to specify task intent rather than to be repeatedly processed during high-frequency low-level execution. Motivated by this separation, we propose a cerebello-thalamic-inspired vision-action model (CT-VAM) for efficient task-conditioned visuomotor control. CT-VAM acts as a compact local execution policy that predicts action...

    arxiv.org/abs/2606.09572 · PDF

  7. 07

    Targeting World Models to Compromise Robot Learning Pipelines

    Ethan Rathbun, Ahmed Agha, Saaduddin Mahmud, Christopher Amato, Alina Oprea, Eugene Bagdasarian

    cs.RO · cs.AI · cs.CR

    World models have recently seen a rapid growth in both their popularity and capability as more data efficient tools for generating robot training data or simulating real world environments, with many works proposing their integration into the robot learning pipeline. While highly practical, in this work we demonstrate that world models introduce a uniquely stealthy and effective data poisoning entry point into the robot learning supply chain...

    arxiv.org/abs/2606.09499 · PDF

  8. 08

    Dense Force Estimation with an Event-based Optical Tactile Sensor

    Agis Politis, René Zurbrügg, Valentina Cavinato

    cs.RO · cs.CV · cs.LG

    Humans rely on spatially dense, geometry and force-aware tactile feedback at high temporal resolution for dexterous manipulation. While vision-based tactile sensors enable dense force estimation, they are limited by camera frame rates, motion blur, and data bandwidth. Event-based optical tactile sensors offer an attractive alternative with microsecond temporal resolution and low motion blur, but existing methods are restricted to predicting...

    arxiv.org/abs/2606.09451 · PDF

  9. 09

    Harness Engineering for Physical AI: Robot Middleware Is the Harness Layer

    Sanghoon Lee, Jiyeong Chae, Kyung-Joon Park

    cs.RO · cs.AI · cs.SE

    Robot middleware faces a new role in the era of Physical AI. Learned policies, planners, and vision-language-action (VLA) models now enter deployed robots as causal participants on the control path, but the layer that integrates them with timing, scheduling, and network has not been named. Recent language-agent work names this layer the harness, the external system that mediates tools, manages state, bounds resources, and records execution....

    arxiv.org/abs/2606.09416 · PDF

  10. 10

    Self-Paced Curriculum Reinforcement Learning for Autonomous Superbike Racing in Simulation

    Luca Ghisi, Jacopo Essenziale, Carlo D'Eramo, Matteo Luperto

    cs.RO · cs.AI

    Autonomous Racing has seen remarkable progress through deep Reinforcement Learning (RL), primarily for four-wheeled vehicles. However, motorbikes introduce substantially greater complexity due to the need to manage balance and lean angle, in addition to more reactive steering and throttle control, and a smaller weight. In this work, we present a framework for training an autonomous agent to race a superbike in VRider SBK, a physics-accurate...

    arxiv.org/abs/2606.09236 · PDF

  11. 11

    From USD Scenes to Knowledge Graphs: Zero-Shot Ontology Grounding with LLMs

    Jiangtao Shuai, Zongxiong Chen, Manfred Hauswirth, Sonja Schimmler

    cs.RO · cs.AI · cs.CL · cs.CV · cs.GR

    Constructing knowledge graphs from 3D simulation scenes is essential for robot task reasoning, but the key bottleneck, grounding scene objects to formal ontology classes, still relies on manually curated dictionaries that are brittle and do not generalize across assets. We investigate whether large language models (LLMs) can automate this grounding step for Universal Scene Description (USD) scenes as a zero-shot, training-free alternative. On...

    arxiv.org/abs/2606.09134 · PDF

  12. 12

    RAM: Reachability Across Morphologies

    Tim Walter, Xinyu Chen, Jonathan Külz, Matthias Althoff

    cs.RO · cs.LG

    Many stages of the robotic lifecycle, from morphology synthesis to operation, rely fundamentally on the reachable workspace. However, current methods for approximating workspaces are slow, imprecise, or tied to a single morphology. We introduce Reachability Across Morphologies (RAM): a morphology-conditioned, implicit neural representation that acts as a fast, differentiable surrogate for pose reachability, generalising to unseen morphologies...

    arxiv.org/abs/2606.09108 · PDF

This edition is part of The Daily Abstract — cs.RO archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.