cs.RO · 2026-08-19 · No. 89

Robotics, 2026-08-19.

6 new papers in cs.RO. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

6 entries
  1. 01

    Dijkstra as an Oracle for Online Stochastic Shortest Path Navigation with Provable Guarantees

    Mansur M. Arief, Ali Akarma, Ahmad Alfan Alfian Irfan

    cs.RO · cs.AI · math.OC

    Mobile robots that operate in side by side with humans and critical facilities must reach their goals at low cost, despite often unknown true traversal costs of the map apriori and imperfect actuation. Planners that solve the underlying stochastic shortest path problem exactly, such as value iteration, require computation that grows with the diameter of the map, whereas Dijkstra's algorithm is fast but is usually considered inexact once...

    arxiv.org/abs/2608.17703 · PDF

  2. 02

    Iterative Grasp Pose Refinement: A Deep Reinforcement Learning Approach for 2D Vision

    Amir Arsalan Nematollahi, Shayan Ahmadi, Mehdi Tale Masouleh, Ahmad Kalhor

    cs.RO · cs.AI · cs.LG · eess.IV

    Developing robots capable of understanding and manipulating objects requires compact, interpretable, and generalizable representations. This work proposes a reinforcement learning-based framework for robotic grasp refinement, integrating keypoint-based object representations with a Deep Q-Network (DQN). Using 2D overhead images captured in a simulated environment, a geometric-based algorithm generates initial grasp candidates, which are...

    arxiv.org/abs/2608.17628 · PDF

  3. 03

    tinyDSM: A Framework for Skill Modeling and Development for Resource-Constrained Millirobots

    Markus D. Kobelrausch, Michael Miedler, Axel Jantsch

    cs.RO · cs.AI

    In this study, we investigate developmental mechanisms that enable small, resource-constrained systems such as cm-sized millirobots to autonomously explore, learn, and adapt their capabilities throughout their lifespan. Reinforcement learning algorithms guide the agent's skill acquisition and adaptation through the interplay of our proposed tinyDSM, which integrates intrinsic motivation and fitness-based assessment. We strive for minimal,...

    arxiv.org/abs/2608.17596 · PDF

  4. 04

    Prism-GRPO: Faster VLA Policy Optimization via Splitting Same-outcome Groups

    Zeyun Deng, Yuzhe Lu, Yawei Wang, Linbo Liu, Qing Ping, Han Ding, Guande Wu, Panpan Xu, Jun Huan

    cs.RO · cs.LG

    GRPO is increasingly used for reinforcement learning of vision-language-action (VLA) policies because, unlike PPO, it does not require training a critic. This simplification comes with a sampling cost: group-relative advantages require multiple rollouts from each scene. Under binary success rewards, groups whose rollouts all succeed or all fail have zero advantage and are discarded by dynamic sampling. These groups are especially common early...

    arxiv.org/abs/2608.17423 · PDF

  5. 05

    ORPA: Online Residual Policy Adaptation for Robot Manipulation Control with Human Feedback

    Muhammad A. Muttaqien, Tomohiro Motoda, Ryo Hanai, Yukiyasu Domae

    cs.RO · cs.AI

    Robotic manipulation policies trained via imitation learning, such as Action Chunking with Transformers (ACT), can achieve strong performance under ideal conditions but often remain sensitive to small execution errors and distribution shifts. Correcting these failures typically requires dataset aggregation and full-policy retraining, which is computationally expensive and unsuitable for real-time deployment. In this work, we propose Online...

    arxiv.org/abs/2608.17323 · PDF

  6. 06

    Teach and Grow: An Agent-Centered Architecture for General Robot Learning

    Chang Nie, Zhe Liu, Hesheng Wang

    cs.RO · cs.AI · cs.CV · cs.LG

    End-to-end vision-language-action (VLA) and world-action models offer an elegant route to general-purpose robotics, but their reliability is bounded by validated physical coverage. When an unfamiliar object, sensor, embodiment, or contact falls outside that coverage and no validated fallback exists, correcting the failure requires new robot data, a policy update, and regression testing. This recurring burden is the retraining tax. Unlike...

    arxiv.org/abs/2608.17209 · PDF

This edition is part of The Daily Abstract — cs.RO archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.