cs.RO · 2026-07-19 · No. 58

Robotics, 2026-07-19.

9 new papers in cs.RO. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

9 entries
  1. 01

    RoboTTT: Context Scaling for Robot Policies

    Yunfan Jiang, Yevgen Chebotar, Ruijie Zheng, Fengyuan Hu, Yunhao Ge, Jimmy Wu, Tianyuan Dai, Scott Reed, Li Fei-Fei,...

    cs.RO · cs.AI · cs.LG

    Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-Training Robot Policies (RoboTTT), a robot model and training recipe that scale visuomotor context to 8K timesteps, three orders of magnitude beyond state-of-the-art policies, without growing inference latency. At this context length, we unlock new robot capabilities: one-shot in-context imitation from human video...

    arxiv.org/abs/2607.15275 · PDF

  2. 02

    Scaling Behavior Foundation Model for Humanoid Robots

    Weishuai Zeng, Kangning Yin, Xiaojie Niu, Shunlin Lu, Weixiang Zhong, Jiahe Chen, Feiyu Jia, Xiao Chen, Zirui Wang,...

    cs.RO · cs.AI

    Humanoid control requires natural whole-body coordination, precise real-time responses to control signals, and robust generalization across diverse environmental contexts, making it a cornerstone for generalist embodied agents. Behavior Foundation Models (BFMs) have recently emerged as a promising solution to address these challenges by leveraging large-scale behavioral data to achieve superior expressiveness, versatility and generalization....

    arxiv.org/abs/2607.15163 · PDF

  3. 03

    DriftWorld: Fast World Modeling through Drifting

    Susie Lu, Haonan Chen, Weirui Ye, Yilun Du

    cs.RO · cs.CV · cs.LG

    Predictive world models enable robots to plan by imagining the outcomes of their actions, but their value for control hinges on generating many rollouts quickly. This creates a bottleneck for diffusion-based world models: multistep sampling makes each rollout expensive, limiting large-scale action search at inference time. We introduce DriftWorld, an action-conditioned world model based on drifting generative models. Rather than denoising...

    arxiv.org/abs/2607.15065 · PDF

  4. 04

    Steering Robustness into World Action Models via Mechanistic Interpretability and Optimal Control

    Jihoon Hong, Julian Skifstad, Qiyue Dai, Alice Chan, Glen Chou

    cs.RO · cs.AI · cs.LG · eess.SY · math.OC

    World Action Models (WAMs) enable semantically- and physically-informed control but are brittle under distribution shift. In this work, we use mechanistic interpretability to study how robustness-relevant perturbations are represented in WAM activation space. Comparing activations across successful and unsuccessful rollouts, we find some WAM architectures exhibit low-dimensional linear separability for robustness-critical features, while...

    arxiv.org/abs/2607.14943 · PDF

  5. 05

    Interventional Causal Circuits for Safe Robot Action Testing and Failure Recovery

    Naren Vasantakumaar, Tom Schierenbeck, Michael Beetz

    cs.RO · cs.AI

    Safe physical AI for robot actions are required not only likely to succeed but tested to be safe before execution. In practice, however, formal testing of motion parameters is computationally expensive, and the cost scales poorly with the dimensionality of the action space. When a proposed action is rejected by a tester, the naive response is to resample blindly until a passing candidate is found. This is wasteful, uninformative, and offers...

    arxiv.org/abs/2607.14826 · PDF

  6. 06

    An Intelligent-Cloud Edge Multimodal Interaction System for Robots

    Zihan Guo, Xiaoqi Li

    cs.RO · cs.AI

    Robust human-robot interaction in complex environments requires accurate gesture perception, semantic scene understanding, and reliable task planning under limited onboard computing resources. This paper presents a cloud-edge multimodal interaction framework that integrates an enhanced YOLO-based gesture detector with coordinated large language model (LLM) and vision-language model (VLM) agents. The proposed detector, incorporates the...

    arxiv.org/abs/2607.14675 · PDF

  7. 07

    SafeRelBench: A Spatial-Relation-Aware Benchmark for Process-Level Safety in VLM-Driven Embodied Agents

    Huaigang Yang, Ya Li, Min Ren, Bo Dai, Zhenliang Zhang, Zhaofeng He

    cs.RO · cs.AI

    Vision-language models (VLMs) are increasingly used as the reasoning backbone of embodied agents, enabling robots to interpret visual scenes, follow language instructions, and plan multi-step actions. In household environments, however, safety depends not only on recognizing objects, but also on how actions change the physical scene over time. Existing embodied safety evaluations largely focus on static risk recognition, unsafe instruction...

    arxiv.org/abs/2607.14543 · PDF

  8. 08

    ConFlow: Constraints-Guided Learning with Flow Matching for Motion Generation

    Nutan Chen, Jianxiang Feng, Marvin Alles, Botond Cseke

    cs.RO · cs.AI · cs.LG

    In recent years Flow Matching has become a prominent method for generative modeling robot motion generation. In its generic form Flow Matching is an ODE-based neural sampler that is trained by regressing empirical flow fields associated with motion samples as data. However, in robot motion generation we often have additional constraints that might not be present in the collected data. The majority of current approaches train the flow on the...

    arxiv.org/abs/2607.14424 · PDF

  9. 09

    An offline approach to fNIRS-guided reinforcement learning for robot behavior

    Julia Santaniello, Madelaine Brower, Benson Jiang, Donatello Sassaroli, Robert Jacob, Jivko Sinapov

    cs.RO · cs.AI

    Human-in-the-loop Reinforcement Learning has become a popular approach to training, finetuning, and aligning robot behavior with user preferences. Our paper explores the feasibility of using brain signals via functional near-infrared spectroscopy (fNIRS) to modulate robot learning in simulation. We compare agents trained on passive (observational) versus active (demonstrative) interaction tasks, and test multiple methods for enhancing the RL...

    arxiv.org/abs/2607.14393 · PDF

This edition is part of The Daily Abstract — cs.RO archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.