cs.RO · 2026-07-01 · No. 40

Robotics, 2026-07-01.

12 new papers in cs.RO. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

12 entries
  1. 01

    Freeform Preference Learning for Robotic Manipulation

    Marcel Torne, Anubha Mahajan, Abhijnya Bhat, Chelsea Finn

    cs.RO · cs.AI · cs.LG

    Reward design remains a central bottleneck for autonomous robot policy improvement, especially in long-horizon manipulation tasks where sparse success labels provide too little signal and binary preferences collapse many competing notions of quality into one ambiguous signal. We introduce Freeform Preference Learning (FPL), a method for learning robot policies from freeform human preferences. Rather than asking annotators which of two...

    arxiv.org/abs/2606.32027 · PDF

  2. 02

    LeCropFollow: Latent Space Planning for Navigation in Unstructured Crop Fields

    Felipe Tommaselli, Francisco Affonso, Arthur Pompeu, Gianluca Capezzuto, Arun Narenthiran Sivakumar, Girish...

    cs.RO · cs.AI

    Unstructured navigational features, such as irregular planting or discontinuities, remain the primary failure mode for under-canopy agricultural robots. Existing geometric approaches often fail in these scenarios because they compress high-dimensional visual data into deterministic spatial references, effectively discarding the uncertainty and semantic context required to navigate ambiguous terrain. To address this, we present LeCropFollow, a...

    arxiv.org/abs/2606.31941 · PDF

  3. 03

    MVP-Nav: Multi-layer Value Map Planner Navigator

    Wenyuan Xie, Shaokai Wu, Yijin Zhou, Yanbiao Ji, Guodong Zhang, Bayram Bayramli, Qiuchang Li, Xunchu Zhou, Yue Ding,...

    cs.RO · cs.AI · cs.CV

    Zero-shot Object Goal Navigation (ZSON) with RGB-only perception poses a fundamental challenge for embodied agents, as the absence of explicit depth information introduces severe physical uncertainty and semantic-physical misalignment. Existing approaches either rely on high-level semantic reasoning without geometric grounding or learn end-to-end policies that lack explicit physical constraints, often resulting in semantically plausible but...

    arxiv.org/abs/2606.31919 · PDF

  4. 04

    Z-1: Efficient Reinforcement Learning for Vision-Language-Action Models

    Lang Cao, Renhong Chen, Luyi Li, Peng Wang, Mofan Peng, Yitong Li

    cs.RO · cs.AI

    Vision-Language-Action (VLA) models offer a promising framework for robotic manipulation by connecting language instructions, visual observations, and continuous control. However, most existing policies remain limited by behavior cloning or supervised fine-tuning (SFT) from fixed demonstrations, which provides limited opportunity to improve from the policy's own failures. In this paper, we present Z-1, a reinforcement learning (RL)...

    arxiv.org/abs/2606.31846 · PDF

  5. 05

    Bridging Local Observation and Global Simulation in Closed-Loop Traffic Modeling

    Ziyan Wang, Tan Xiang, Peng Chen, Xintao Yan

    cs.RO · cs.AI

    A local-to-global context mismatch arises when autoregressive traffic simulators trained on ego-centric driving logs are deployed in globally observable closed-loop environments. In such logs, the ego vehicle has rich local observations, while surrounding agents are only partially observed due to perception limits and occlusions. As a result, simulators may learn incomplete context--action mappings that remain hidden in log-based training but...

    arxiv.org/abs/2606.31844 · PDF

  6. 06

    RCT: A Robot-Collected Touch-Vision-Language Dataset for Tactile Generalization

    Jingbo He, Michael Färber, Roberto Calandra

    cs.RO · cs.AI · cs.CL · cs.CV

    For robots manipulating open-world objects, tactile representations must generalize to unseen materials. We introduce RCT (Robotic Contact Tactile), a robot-collected touch-vision-language dataset with 29,279 tactile frames from full robot presses on 122 industrial reference materials in 7 categories, recorded with three DIGIT sensors at multiple contact positions. RCT preserves each press as a contact sequence, enabling held-out evaluation...

    arxiv.org/abs/2606.31694 · PDF

  7. 07

    Robustness of Robotic Manipulation: Foundations and Frontiers

    Yifei Dong, Zhanyi Sun, Lujie Yang, Manuel Baum, Kei Ikemura, Shuran Song, Florian T. Pokorny, Xianyi Cheng

    cs.RO · cs.AI

    Humans and animals exhibit remarkable robustness in physical manipulation, yet robots remain far behind. Progress toward human-level manipulation robustness is hindered by the absence of a unified and systematic understanding: different subfields frame robustness in distinct ways, often leaving the concept ambiguous and limiting deeper analysis as well as communication across research areas. This paper presents a systematic study of...

    arxiv.org/abs/2606.31494 · PDF

  8. 08

    UniTac: A Unified Multimodal Model for Cross-Sensor Tactile Understanding and Generation

    Jiahang Tu, Fengyu Yang, Chenyang Ma, Xihang Yu, Ziyao Zeng, Shaokai Wu, Hanbin Zhao, Zhi Tao, Chao Zhang, Hui Qian,...

    cs.RO · cs.AI

    Unified multimodal models (UMMs) have shown great promise in integrating understanding and generation across diverse modalities. However, existing research rarely extends this paradigm to the tactile domain, where both object-level semantics and sensor-level configurations jointly determine the meaning of touch. To address this gap, we propose UniTac, the first UMM designed for tactile understanding and generation. UniTac models the tactile...

    arxiv.org/abs/2606.31451 · PDF

  9. 09

    Stage-Transition Dense Reward Modeling for Reinforcement Learning

    Yang Yang, Bingjie Chen, Zihan Wang, Yizhe Li, Guoping Pan, Yi Cheng, Houde Liu

    cs.RO · cs.AI

    Reinforcement learning for long-horizon robotic manipulation is often limited by sparse and delayed rewards, while manually designing dense shaping signals is costly and brittle to changes in environments and object configurations. This work proposes Stage-Transition Dense Reward (STDR), a visual reward-learning framework that converts unstructured expert videos into logically grounded dense rewards for training RL agents from scratch. STDR...

    arxiv.org/abs/2606.31377 · PDF

  10. 10

    3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance

    Dongyoon Hwang, Byungkun Lee, Dongjin Kim, Hyojin Jang, Hoiyeong Jin, Jueun Mun, Minho Park, Hojoon Lee, Hyunseung...

    cs.RO · cs.AI

    Hierarchical Vision-Language-Action (VLA) models decouple high-level planning from low-level control to improve generalization in robot manipulation. Recent work in this paradigm uses 2D end-effector trajectories predicted by a Vision-Language Model (VLM) as explicit guidance for a downstream policy. However, state-of-the-art low-level policies operate in 3D metric space on point clouds, and feeding them 2D guidance that lacks depth forces...

    arxiv.org/abs/2606.31329 · PDF

  11. 11

    Information-Aided DVL Calibration

    Zeev Yampolsky, Itzik Klein

    cs.RO · cs.AI

    The Doppler velocity log (DVL) velocity measurements are critical to the accuracy of autonomous underwater vehicle (AUV) navigation solutions and, consequently, to mission success. To ensure accurate measurements, the DVL is commonly calibrated before mission start while the AUV sails on the water surface, receiving global navigation satellite system (GNSS) signals that provide accurate reference measurements. Conventionally, Kalman...

    arxiv.org/abs/2606.31216 · PDF

  12. 12

    Machine Learning-based Feedback Linearization Control of Quadrotor Subject to Unmodeled Dynamics

    Amos Alwala, Gabriel da Silva Lima, Wallace Moreira Bessa

    cs.RO · cs.LG · eess.SY

    The control of agile quadrotors in dynamic and uncertain environments remains an open area of investigation to this day, particularly when the complete system dynamics are partially known or highly nonlinear. This work introduces a novel machine learning-based feedback-linearization control framework that employs a Gaussian Radial Basis Function (RBF) neural network (NN) to model and compensate for unmodeled dynamics in real time. The...

    arxiv.org/abs/2606.31199 · PDF

This edition is part of The Daily Abstract — cs.RO archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.