cs.RO · 2026-07-02 · No. 41

Robotics, 2026-07-02.

9 new papers in cs.RO. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

9 entries
  1. 01

    FurnitureVLA: Learning Long-Horizon Bimanual Furniture Assembly with Vision-Language-Action Model

    Chenyang Ma, Yue Yang, Radu Corcodel, Siddarth Jain, Andrew Wu, Chiori Hori, Diego Romeres

    cs.RO · cs.AI

    Current work on robot furniture assembly mostly focuses on toy-scale settings or single-arm manipulation. We introduce FurnitureVLA, the first systematic study of real-scale bimanual furniture assembly using Vision-Language-Action models (VLAs). We formalize the task, develop a scalable simulation pipeline for expert data generation and evaluation, and build a VR teleoperation system for single-operator bimanual control to collect...

    arxiv.org/abs/2607.01212 · PDF

  2. 02

    FAR: Failure-Aware Retry for Test-Time Recovery and Continual Policy Improvement

    Haoran Hao, Shahram Najam Syed, Jeffrey Ichnowski, Jeff Schneider

    cs.RO · cs.AI · cs.LG

    Robot policies inevitably encounter failures when deployed in real environments. Naive retries often repeat the same mistakes, while many existing recovery methods rely on human intervention. In this paper, we propose Failure-Aware Retry (FAR), a framework that enables robots to learn from previous failures at test time, adapt their behavior accordingly, and eventually complete the task autonomously. FAR combines Failure-Contrastive...

    arxiv.org/abs/2607.01111 · PDF

  3. 03

    ROSA: A Robotics Foundation Model Serving System for Robot Factories

    Wenqi Jiang, Jason Clemons, Rowland O'Flaherty, Hugo Hadfield, Alperen Degirmenci, Shuran Song, Yashraj Narang,...

    cs.RO · cs.DC

    Robotics foundation models (RFMs) are making general-purpose robots increasingly practical for factory deployments. While RFM serving systems are central to this vision, existing systems are largely shaped by a single-robot, single-model assumption: inference is treated as an edge-computing problem handled by an on-robot or dedicated nearby GPU, and the serving objective is to minimize the latency of a single action model. In this paper, we...

    arxiv.org/abs/2607.01088 · PDF

  4. 04

    DART-VLN: Test-Time Memory Decay and Anti-Loop Regularization for Discrete Vision-Language Navigation

    Shaoheng Zhang, Zhichen Li, Jie Mei

    cs.RO · cs.AI

    Memory-based discrete vision-language navigation (VLN) agents must act under partial observability, yet even strong frozen backbones remain vulnerable at test time. Two common failure modes are stale historical evidence at memory readout and inefficient local backtracking during action selection. We present DART-VLN, a training-free test-time control framework for discrete VLN. DART-VLN combines Test-Time Memory Decay, a read-side memory...

    arxiv.org/abs/2607.01043 · PDF

  5. 05

    From World Models to World Action Models: A Concise Tutorial for Robotics

    Xiaoxiong Zhang, Xiong Zeng, Wei Zhang

    cs.RO · cs.AI · eess.SY

    World models are increasingly used in embodied intelligence and generative simulation, yet their scope remains ambiguous across communities. This tutorial presents a design-space view of world models as action-conditioned predictive models that estimate the future evolution of task-relevant observations or states. We categorize existing methods into observation-space and state-space world models, comparing their trade-offs in visual fidelity,...

    arxiv.org/abs/2607.00836 · PDF

  6. 06

    Domain Arithmetic: One-Shot VLA Adaptation under Environmental Shifts

    Taewook Kang, Taeheon Kim, Donghyun Shin, Jonghyun Choi

    cs.RO · cs.CV · cs.LG

    Vision-Language-Action (VLA) models often fail to perform the same learned tasks under environmental shifts, such as changes in camera pose and shifts to a different but similar robot (e.g., from Panda to UR5e). Adapting these models to the shifted environment (i.e., target domain) often requires training on multiple demonstrations for each task, which are costly to collect. To reduce the burden of data curation and training, we propose an...

    arxiv.org/abs/2607.00666 · PDF

  7. 07

    From Technical Metrics to User Perception: A User Study of a Multimodal Human-Robot Interaction System for Object Detection and Grasping

    Jian Song, Tian Zi, Shen Guanting

    cs.RO · cs.AI

    Improvements in the technical performance of human--robot interaction (HRI) systems do not automatically translate into differences that human users can detect during live interaction. This paper investigates whether a 15 percentage point gain in end-to-end task success (from 75% in a multimodal baseline system to 90% in an improved configuration identified through a prior ablation study) is sufficient to produce consistent and measurable...

    arxiv.org/abs/2607.00530 · PDF

  8. 08

    Search-Based Spatiotemporal and Multi-Robot Motion Planning on Graphs of Space-Time Convex Sets

    Jingtao Tang, Zining Mao, Lufan Yang, Hang Ma

    cs.RO · cs.AI

    Spatiotemporal motion planning, especially in multi-robot settings, requires robots to reason about collision-free regions that change over time, which is challenging in continuous spaces when feasible regions are transient and geometrically constrained. We present an algorithmic framework based on graphs of space-time convex sets (ST-GCSs), where collision-free regions are represented as convex sets in space-time and trajectories correspond...

    arxiv.org/abs/2607.00444 · PDF

  9. 09

    Learning Gait-Aware Quadruped Locomotion with Temporal Logic Specifications

    Merve Atasever, Cagan Bakirci, Alfredo Reina Corona, Keyan Azbijari, Jyotirmoy V. Deshmukh

    cs.RO · cs.AI

    Reinforcement learning (RL) for quadruped locomotion commonly depends on fixed, hand-crafted, and Markovian reward functions that limit both interpretability of learned policies and lack explicit control over gait behaviors. We introduce a framework where distinct gaits are specified using parameterized constraints expressed in Signal Temporal Logic (STL). These include safety bounds, gait synchronization constraints, command tracking, and...

    arxiv.org/abs/2607.00442 · PDF

This edition is part of The Daily Abstract — cs.RO archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.