cs.RO · 2026-10-01 · No. 130

Robotics, 2026-10-01.

12 new papers in cs.RO. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

12 entries
  1. 01

    DynaHarness: A Dynamic Physical Harness for Self-Evolving Robot Agents

    Haoyuan Deng, Jiebin Liu, Tengxiao Zhang, Langning Yan, Hongye Cao, Ziwei Wang

    cs.RO · cs.AI · cs.LG

    Pretrained robot policies provide useful action priors, but long-horizon manipulation still requires coordination between semantic reasoning and physical execution. Semantic reasoning operates at a coarser timescale than physical interaction, while episode-level failures provide limited guidance on which system component should be revised. We propose DynaHarness, a dynamic physical harness that couples semantic reasoning with physical...

    arxiv.org/abs/2609.40306 · PDF

  2. 02

    STARS: From Spatiotemporal Dynamics to Social Representations in Human-Robot Interaction

    Nathan Tsoi, Michael J. Munje, Tejas Oberoi, Rishab Maheshwari, Pengen Zheng, Tanush Chauhan, Peter Stone, Joydeep Biswas

    cs.RO · cs.LG

    Robot navigation in dynamic, human-centered environments requires socially-compliant decisions grounded in robust scene understanding. Recent Vision-Language Models (VLMs) exhibit promising capabilities such as object recognition, common-sense reasoning, and contextual understanding, capabilities that align with the nuanced requirements of social robot navigation. However, it remains unclear whether VLMs can accurately understand complex...

    arxiv.org/abs/2609.40245 · PDF

  3. 03

    PrefPI: Preference-Guided Steering into Out-of-Distribution Behaviors

    Seungeun Rho, Wontaek Kim, Danfei Xu, Sehoon Ha

    cs.RO · cs.AI · cs.LG

    We present PrefPI (Preference-Guided Policy Iteration), an iterative framework for steering pretrained generative robot policies using only relative preferences over self-generated trajectories. Unlike prior preference-learning methods that primarily sharpen modes already represented by the policy, we study steering beyond the initial effective support, where desired behaviors are rarely or never observed under the initial policy. Our key...

    arxiv.org/abs/2609.40165 · PDF

  4. 04

    Tactile Curiosity Drives Robot Interaction

    Klemens Iten, Alexander Proshkin, Bhavya Sukhija, Stelian Coros, Andreas Krause, Pieter Abbeel, Carmelo Sferrazza

    cs.RO · cs.AI

    Mastering robot manipulation skills via reinforcement learning (RL) remains largely sample-inefficient. The most common RL algorithms rely on random action sampling to discover new strategies, resulting in agents that allocate most of their training budget to motions in free space, away from the contacts from which manipulation skills emerge. Existing intrinsic motivation methods based on model disagreement or epistemic uncertainty improve on...

    arxiv.org/abs/2609.40134 · PDF

  5. 05

    BatSLAM 2.0: Sequence-Verified Sonar Place Recognition in a Robust Pose Graph

    Jan Steckel

    cs.RO · cs.AI · cs.LG · eess.SY

    Echolocating bats can navigate dark and cluttered spaces using echolocation. Over a decade ago, BatSLAM showed that a robot with a biomimetic binaural sonar can build a topological map of the environment, by recognizing places from the received acoustic signals. Sonar place recognition, however, is ambiguous by nature: corridors produce nearly identical echo trains, and wrong loop closure can collapse the topological map. In this paper, we...

    arxiv.org/abs/2609.40085 · PDF

  6. 06

    When Instructions Retrieve Trajectories: Diagnosing and Mitigating Generalization Failures in VLA Models

    Hung-Jen Chen, Yu-Hsun Hou, Yan-Hong Chen, Yan-Fu Chen, Binghua Cai, Min Sun, Chun-Yi Lee

    cs.RO · cs.LG

    Vision-language-action (VLA) models can exceed 90% success on in-distribution tasks and withstand nuisance changes that preserve the required action, yet fail under counterfactual changes that demand a different action. Aggregate robustness scores can therefore conceal a more specific failure, in which a policy responds to both language and vision yet does not combine them to select the action the task requires. We call this failure...

    arxiv.org/abs/2609.39971 · PDF

  7. 07

    TACTIC: Temporal and Context-Aware LLM Tactical Planning for Roadside LiDAR Attacks

    Yiming Gao, Shaocheng Luo

    cs.RO · cs.AI · eess.SY

    Physical LiDAR attacks are often evaluated using fixed primitives and manually selected parameters, despite their strong dependence on surrounding traffic. We present TACTIC, a scene-aware framework that uses a multimodal large language model (MLLM) to coordinate state-adaptive roadside LiDAR attacks. Under a gray-box threat model, TACTIC relies only on an attacker-operated roadside perception stack, without accessing the victim LiDAR's...

    arxiv.org/abs/2609.39969 · PDF

  8. 08

    Learning from Runtime Feedback through Failure-Bank Self-Evolution for Vision-Language-Action Models

    Mingyue Cui, Zheyuan Liu, Yihan Zhu, Zheyuan Zhang, Meng Jiang

    cs.RO · cs.AI

    Vision-language-action (VLA) models generalize broadly across robotic manipulation tasks, but complex environments require balancing task success with unintended contact. Runtime shields can correct individual actions, but they leave the underlying policy unchanged, so repeated disagreements may create a persistent policy-shield mismatch that blocks task progress. To address this challenge, we introduce FailBank, a four-stage self-evolving...

    arxiv.org/abs/2609.39820 · PDF

  9. 09

    DiffWAM: A Fast and Efficient Navigation World Action Model

    Mo Zhu, Yuze Wu, Xijie Huang, Xiao Cui, Fei Gao, Xin Zhou

    cs.RO · cs.AI

    Pretrained video foundation models encode rich semantic and spatiotemporal priors for embodied navigation, yet converting these priors into UAV motion typically requires expensive future-video synthesis and geometric reconstruction. We investigate whether the motion implicit in future visual prediction can instead be recovered directly from the predictive representations of a frozen video model. To this end, we present DiffWAM, a...

    arxiv.org/abs/2609.39763 · PDF

  10. 10

    RoboCoach: World Models as Active Coaches for Compositional Robot Skills

    Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

    cs.RO · cs.AI

    Long-horizon robot manipulation reuses skills across many task compositions, but improving these compositions with additional end-to-end demonstrations is costly. A practical self-improving system must decide both what to teach next and where to apply that supervision. We present ROBOCOACH, a world-model-guided coaching framework that uses imagined failures to guide demonstration requests and expert updates. Its Route-Imagine-Diagnose-Improve...

    arxiv.org/abs/2609.39685 · PDF

  11. 11

    Text-to-3D Policy: Fine-Grained Language-Behavior Alignment for Unseen Specification Generalization

    Xinhao Yang, Wenhao Wu, Ning Lv, Yanshen Ding, Zhenhong Sun, Daoyi Dong, Chunlin Chen, Zhi Wang

    cs.RO · cs.AI

    3D visuomotor policies provide a strong foundation for spatially precise manipulation, yet current text-to-3D policies struggle to follow unseen fine-grained behavioral specifications beyond those covered by demonstrations. We study this challenge as unseen specification generalization, where language specifies behaviorally significant variations, such as target position, displacement, or articulated state, that are absent from policy...

    arxiv.org/abs/2609.39599 · PDF

  12. 12

    ECHO-G: Embodied Co-speech Humanoid mOtion Generation

    Yizhao Li, Pusen Gao, Ming Wang, Shaojie Shen, Shuo Yang, Hao Xu

    cs.RO · cs.AI

    Generating full-body co-speech motion for humanoid robots requires coordinating speech prosody, linguistic content, and embodiment-specific motion. To this end, we present ECHO-G, a framework jointly conditioned on speech audio and timed transcripts. Its Speech-Grounded Diffusion Transformer (SGDiT) combines frame-aligned acoustic features with token-level linguistic context, preserving their distinct granularities. Trained with rectified...

    arxiv.org/abs/2609.39575 · PDF

This edition is part of The Daily Abstract — cs.RO archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.