cs.RO · 2026-07-04 · No. 43

Robotics, 2026-07-04.

10 new papers in cs.RO. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

10 entries
  1. 01

    Controllable Sim Agents with Behavior Latents

    Juanwu Lu, Junyu Zhu, Ziran Wang

    cs.RO · cs.LG

    Realistic traffic simulation requires agents that imitate logged behavior and can also be steered along interpretable axes. Such controllability enables engineers to isolate variables, reproduce specific edge cases, and test autonomous systems without real-world risk. We introduce Controllable Neural Variational Agents (CNeVA), a controllable simulated-agent framework that learns to infer a per-agent Gaussian behavior latent from per-channel...

    arxiv.org/abs/2607.02496 · PDF

  2. 02

    Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs

    Junhao Shi, Siyin Wang, Xiaopeng Yu, Li Ji, Jingjing Gong, Xipeng Qiu

    cs.RO · cs.AI

    Vision-Language-Action (VLA) models are fundamentally bottlenecked by the scarcity of expert demonstrations -- triplets of observations, instructions, and actions that are costly to collect at scale. We argue that this bottleneck stems from conflating two distinct learning objectives: acquiring physical competence (how to move) and acquiring semantic alignment (what to do). Crucially, only the latter requires language supervision. Building on...

    arxiv.org/abs/2607.02466 · PDF

  3. 03

    WorldSample: Closed-loop Real-robot RL with World Modelling

    Yuquan Xue, Le Xu, Zeyi Liu, Zhenyu Wu, Zhengyi Gu, Xinyang Song, Bofang Jia, Ziwei Wang

    cs.RO · cs.AI

    Reinforcement learning (RL) can overcome the demonstration-coverage limitation of imitation learning (IL) by allowing robots to improve through trial-and-error interaction beyond the states observed in demonstrations. However, deploying RL on real robots remains constrained by high interaction costs, since each physical rollout is costly and reflects only one realized action-outcome path. To address this challenge, we propose WorldSample, a...

    arxiv.org/abs/2607.02431 · PDF

  4. 04

    LIME: Learning Intent-aware Camera Motion from Egocentric Video

    Boyang Sun, Jiajie Li, Yung-Hsu Yang, Chenyangguang Zhang, Tim Engelbracht, Sunghwan Hong, Cesar Cadena, Marc...

    cs.RO · cs.CV · cs.LG

    Autonomous robots often need to move their camera before they can act: to inspect an object, reveal an occluded region, or obtain a view that responds to a user's intent. While vision-language navigation translates instructions to base motion and vision-language-action policies map instructions to manipulation actions, language-conditioned camera motion remains comparatively underexplored as a first-class action. We formulate...

    arxiv.org/abs/2607.02417 · PDF

  5. 05

    ACID: Action Consistency via Inverse Dynamics for Planning with World Models

    Gawon Seo, Dongwon Kim, Suha Kwak

    cs.RO · cs.AI · cs.CV

    Decision-time planning with action-conditioned world models has become a popular paradigm for embodied control. However, the standard planning cost judges a candidate solely by how close its predicted terminal state lies to the goal, leaving the realizability of the intermediate transitions unchecked -- a predicted trajectory can look convincing while the environment rollout drifts away from it. In this paper, we propose ACID, a decision-time...

    arxiv.org/abs/2607.02403 · PDF

  6. 06

    CoFL-S: Spatially Queryable Sector Flow Fields for Local Language-Conditioned Navigation

    Haokun Liu, Zhaoqi Ma, Yicheng Chen, Wentao Zhang, Masaki Kitagawa, Zicen Xiong, Jinjie Li, Moju Zhao

    cs.RO · cs.AI

    Vision-Language Navigation has increasingly emphasized high-level instruction reasoning, memory, global map construction, and instruction decomposition, while the low-level action representation remains comparatively underexplored. We propose CoFL-S, a low-level vision-language-action framework that predicts a language-conditioned flow field over the robot's local visible sector and generates continuous trajectories by rolling out the...

    arxiv.org/abs/2607.02222 · PDF

  7. 07

    Guided Action Flow: Q-Guided Inference for Flow-Matching Vision-Language-Action Policies

    Liuhaichen Yang, Zhuang Jiang, Chenchao Sheng, Zezhi Tang

    cs.RO · cs.AI

    Flow-matching vision-language-action policies generate robot action chunks through an iterative transport process, creating an opportunity for test-time guidance without retraining the base policy. We study this opportunity in Guided Action Flow, an inference-time framework that keeps a pretrained SmolVLA policy frozen and uses a learned action-chunk critic to guide its reverse-time flow sampler. The critic is trained from real success and...

    arxiv.org/abs/2607.02092 · PDF

  8. 08

    Cross-Platform Control for Autonomous Surface Vehicles via Adaptive Reinforcement Learning

    Ruiheng Jiang, Thomas Bi, Raffaello D'Andrea, Aswin Ramachandran

    cs.RO · cs.LG

    Autonomous surface vehicles vary widely in hydrodynamic and actuation characteristics, yet most controllers are designed for single-platform deployment. We present an adaptive reinforcement learning approach for trajectory tracking that enables zero-shot cross-platform deployment using a single policy. Since the deployment platform's dynamics are unknown to the policy, we address cross-platform generalization with the standard...

    arxiv.org/abs/2607.02037 · PDF

  9. 09

    PhysMani: Physics-principled 3D World Model for Dynamic Object Manipulation

    Peng Yun, Shouwang Huang, Hao Li, Jinxi Li, Jianan Wang, Bo Yang

    cs.RO · cs.AI · cs.CL · cs.CV · cs.LG

    Manipulating fast and dynamically moving targets in unstructured 3D environments remains challenging for embodied AI. Existing visual-language-action models and world models struggle with accurate 3D geometry and physically meaningful forecasting. We propose PhysMani, a framework that couples a physics-principled 3D Gaussian world model with a future-aware action policy model. The world model learns a divergence-free Gaussian velocity field...

    arxiv.org/abs/2607.01938 · PDF

  10. 10

    Lightweight Safe Reinforcement Learning for End-to-End UAV Navigation

    Shenghui Zhang, YuXuan Gao, Songwei Zhao, Jifeng Hu, Zijing Zhang, Hechang Chen

    cs.RO · cs.AI

    With the rapid development of autonomous aerial systems, Unmanned Aerial Vehicles (UAVs) are increasingly deployed in applications such as inspection, environmental monitoring, and rescue, creating growing demand for reliable autonomous navigation. However, autonomous UAV navigation in dense environments remains challenging under sparse perception and dynamic constraints. Most reinforcement learning (RL) methods lack explicit safety...

    arxiv.org/abs/2607.01794 · PDF

This edition is part of The Daily Abstract — cs.RO archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.