cs.RO · 2026-10-06 · No. 135

Robotics, 2026-10-06.

11 new papers in cs.RO. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

11 entries
  1. 01

    Recursive Video In-Context Learning for Agentic Robot

    Wenrui Bao, Xinxin Liu, Bingxin Xu, Yuzhang Shang

    cs.RO · cs.AI · cs.CL · cs.MA

    LLM agents that orchestrate frozen vision-language-action (VLA) policies improve across episodes through text memory, which records what the agent did but not how the task is done. A demonstration video shows it, but fits poorly into an agent's context. The full video slows every turn, fixed keyframes lose the contact detail that decides whether a grasp holds, and what the agent needs shifts from the task's structure while planning to the...

    arxiv.org/abs/2610.06843 · PDF

  2. 02

    AffordCraft: Scalable Construction of Task-Ready Simulation Assets from Single Images

    Haoyun Yang, Xueyang Zhou, Ziyi Xie, Yongchao Chen

    cs.RO · cs.AI · cs.CV

    Robot learning in simulation depends on the objects the simulator offers. Many tasks need objects with separate parts, joints that allow the required motion, and physical properties that remain valid under contact. Existing methods recover this structure anew for every image: generative models predict parts and joints that mostly fail to settle or move in simulation, and general-purpose agents need a long session of model calls for each...

    arxiv.org/abs/2610.06643 · PDF

  3. 03

    SimForcing: Distilling Simulation Motion Priors into Real-Domain Robot World Models

    Xiaodong Wang, Tianle Li, Chuanxin Song, Junliang Xie, Zhanmi Zhong, Suiying Wu, Peixi Peng

    cs.RO · cs.AI · cs.CV

    Action-conditioned robot world models must respond precisely to robot trajectories while preserving realistic visual dynamics, yet learning both from heterogeneous robot videos remains challenging. Simulation offers structured motion supervision, but appearance differences hinder direct transfer, and inaccurate simulation predictions can misguide real-video generation. We present SimForcing, a simulation-guided framework that uses simulation...

    arxiv.org/abs/2610.06598 · PDF

  4. 04

    ArtifactArena: Evaluating Models by What They Build in the Physical World

    Kushagra Tiwary*, David Mayo*, Nikhil Behari, Xiangzhou Sun, Abdulrahman Alabdulkareem, Isaac Galatzer-Levy, Boris...

    cs.RO · cs.AI

    To evaluate the frontier, we must measure models not by what they say, but by what they can engineer and build in grounded physical environments. We introduce \textsc{ArtifactArena}, an open-ended platform where models face a physically grounded hardware-software co-design challenge: engineering fully functional robots to compete in a simulated arena. We evaluate a frontier model's zero-shot, verifier guided refinement, and open-ended...

    arxiv.org/abs/2610.06511 · PDF

  5. 05

    Odyssey: A Closed-Loop Benchmark for Long-Horizon Real-World Driving with Explicit Navigation Routes

    Jungho Kim, Hongjae Shin, Seunghoon Yu, Heecheol Yoo, Myeongjun Kim, Jiyong Oh, Donghyuk Kwak, Seunghyeop Nam,...

    cs.RO · cs.AI

    Closed-loop evaluation of end-to-end driving requires continuous rollouts that reveal how earlier decisions affect subsequent driving. However, existing benchmarks evaluate only short segments and fail to capture later consequences. Ambiguous directional commands also obscure the intended navigation objective. We introduce Odyssey, a closed-loop benchmark for long-horizon driving comprising 100 scenarios, each reconstructed from a 100-second...

    arxiv.org/abs/2610.06469 · PDF

  6. 06

    Dual Variational Autoencoders for Efficient Sim-to-Real Transfer in Low-Cost Robotic Navigation

    Álvaro Díez, Fidel Aznar

    cs.RO · cs.CV · cs.LG

    Vision-based autonomous navigation for low-cost robots remains a fundamental challenge, primarily due to the significant gap between simulated training environments and real-world operational conditions. Direct policy transfer from simulation is often ineffective, while training exclusively on real data is impractical. We propose a hybrid transfer learning framework that effectively bridges the sim-to-real gap by combining domain...

    arxiv.org/abs/2610.06327 · PDF

  7. 07

    GAMBIT: Learning to Plan Continuous Multi-Robot Trajectories

    Rishabh Jain, Akmaral Moldagalieva, Lorenzo Magnino, Michael Amir, Keisuke Okumura, Ajay Shankar, Wolfgang Hönig,...

    cs.RO · cs.AI · cs.LG · cs.MA

    GAMBIT is an opening chess move in which a player sacrifices a piece, typically a pawn, to gain a positional advantage later in the game. Analogously, in multi-robot coordination, individual robots may need to forgo locally reward-maximising behaviours to improve overall team performance. Such self-sacrificial behaviours are difficult to capture with manually designed heuristics, particularly in dense, interaction-rich environments. Focusing...

    arxiv.org/abs/2610.06290 · PDF

  8. 08

    Future Anchored Verification and Online Recovery for World Action Models

    Zhibin Qin, Zhenxiong Tan, Xinchao Wang

    cs.RO · cs.AI

    World action models (WAMs) have emerged as a promising paradigm for robotic manipulation. They act by first predicting how a task should be performed and then decoding the actions from that future. However, the remaining actions are invalid once execution drifts from the prediction. Simply replanning from the already out of distribution state rarely restores what the task still requires; existing execution monitors decide when to stop, but...

    arxiv.org/abs/2610.06280 · PDF

  9. 09

    VLA-ZO: Fast Zeroth-Order Adaptation for Vision-Language-Action Models

    Jaemin Kim, Jiahn Kim, Taesik Gong

    cs.RO · cs.AI

    Adapting vision-language-action (VLA) models to deployment-time distribution shifts is important for reliable robotic operation, but conventional first-order adaptation can exceed the memory budget of inference-oriented deployment platforms. Zeroth-order (ZO) optimization offers a forward-only alternative with inference-level memory, but accurate gradient estimation requires many perturbation queries, making naive ZO prohibitively slow for...

    arxiv.org/abs/2610.06271 · PDF

  10. 10

    Encoded but Not in Control: Revealing the Grounding Gap in Vision-Language Robot Policies

    Shaohan Jiang, Jiahang Cao, Qiduo He, Fengting Deng, Kun Wu, Jingkai Sun, Jiaxu Wang, Qiang Zhang, Qihao Zheng,...

    cs.RO · cs.LG

    Instruction following is central to language-conditioned robot policies: language should determine what to do when the same scene permits multiple valid actions. Yet successful execution alone cannot establish whether a policy follows the instruction or infers the task from the scene. We study this ambiguity through scene-preserving instruction interventions, using valid target substitutions, arbitrary nouns, and unrelated sentences while...

    arxiv.org/abs/2610.06235 · PDF

  11. 11

    AUTOPILOT An Advanced Perception, Localization and Path Planning Techniques for Autonomous Vehicles Using YOLOv7 and MiDaS

    Harshkumar Devmurari, Gautham Kuckian, Prajjwal Vishwakarma

    cs.RO · cs.LG

    Self driving vehicles have emerged as a reliable technology that has the capability to transform transportation and mobility. The development of self driving cars requires significant advances in a number of areas, including perception, localization, decision making, and control. This research paper is based on the project implementation of the combination of object detection using YOLO (You Only Look Once), depth sensing using MiDaS for the...

    arxiv.org/abs/2610.06232 · PDF

This edition is part of The Daily Abstract — cs.RO archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.