cs.RO · 2026-06-24 · No. 33

Robotics, 2026-06-24.

6 new papers in cs.RO. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

6 entries
  1. 01

    InSight: Self-Guided Skill Acquisition via Steerable VLAs

    Maggie Wang, Lars Osterberg, Stephen Tian, Ola Shorinwa, Jiajun Wu, Mac Schwager

    cs.RO · cs.AI · cs.LG

    Vision-language-action (VLA) models can learn manipulation skills from demonstrations, but their capabilities are bounded by the skills in the training data. We present InSight, a framework that unlocks autonomous skill acquisition by rendering VLAs steerable at the primitive-action level (e.g., "move gripper to the bowl", "lift upward", "pour the bottle"). InSight consists of two primary stages: (1) an automated segmentation pipeline that...

    arxiv.org/abs/2606.24884 · PDF

  2. 02

    TACTFUL: Tactile-Driven Exploration For Object Localization and Identification in Confined Environments

    Shivani Kamtikar, Chung Hee Kim, Camilla Tabasso, Tye Brady, Joshua Migdal, Taskin Padir

    cs.RO · cs.AI

    Humans effortlessly locate and identify objects by touch alone, even without vision. In contrast, robotic systems rely heavily on vision and struggle with autonomous tactile exploration and object identification. We present TACTFUL, a vision-free tactile exploration framework that enables a multi-fingered robot to autonomously explore confined workspaces, discover objects through contact, and identify them via tactile reconstruction. Trained...

    arxiv.org/abs/2606.24712 · PDF

  3. 03

    G$^3$VLA: Geometric inductive bias for Vision-Language-Action Models

    Yue Peng, Yongzhe Zhao, Artur Habuda, Khuyen Pham, Yanheng Zhu, Tran Nguyen Le, Fares Abu-Dakka, Li Guo

    cs.RO · cs.AI

    Vision-language-action (VLA) models have made rapid progress in generalist robot manipulation by harnessing semantic knowledge from pretrained vision-language backbones, but their visual tokens remain grounded in 2D image coordinates rather than the calibrated geometry of the robot's cameras -- a mismatch especially pronounced in multi-camera setups, where views are coupled by known intrinsics and extrinsics yet processed as independent...

    arxiv.org/abs/2606.24472 · PDF

  4. 04

    NoContactNoWorries: Estimating Contact through Vision and Proprioception for In-Hand Dexterous Manipulation

    Soham Patil, Avirup Das, Sourabh Bhosale, Spandan Roy

    cs.RO · cs.AI

    Perceiving physical contact is fundamental to dexterous manipulation. While robots often rely on dedicated hardware tactile sensors, humans exhibit a remarkable ability to infer contact by integrating visual information with an innate sense of their body's pose and movement. Inspired by this embodied perceptual skill, we investigate whether a robot can learn to infer contact from vision, an approach that also offers a scalable alternative to...

    arxiv.org/abs/2606.24450 · PDF

  5. 05

    RE4: Transformation-aware Imitation of Object Interactions Using Manipulation Modes

    Arsh Chawla, Rahul Shome

    cs.RO · cs.LG

    Object interaction tasks have been a focus of advances in imitation learning. End-to-end methods, dominated by diffusion and flow-based variants have shown leaps in performance while sacrificing interpretability. Object-centric and pose-informed variants have had a role in learning from demonstration in manipulation tasks. In this paper, we revisit a few modern imitation learning benchmarks for object interactions, with the aim of composing a...

    arxiv.org/abs/2606.24403 · PDF

  6. 06

    DynaWM: Dynamics-Aware Distillation with World Model and Momentum Targets for Smooth Locomotion over Continuous Stairs

    Haidong Hou, Zhangguo Yu, Hengbo Qi, Jianlin Zhang

    cs.RO · cs.AI

    Recent advances in control have enabled bipedal-wheeled robots to traverse slopes and single-step obstacles, yet long staircase traversal remains challenging as current teacher-student frameworks suffer from weakened dynamics-aware representations and incomplete terrain geometry encoding. To bridge this gap, we propose DynaWM, a dynamics-aware representation learning framework. To enhance terrain encoding capability and enable transparent...

    arxiv.org/abs/2606.24089 · PDF

This edition is part of The Daily Abstract — cs.RO archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.