cs.RO · 2026-09-06 · No. 107

Robotics, 2026-09-06.

6 new papers in cs.RO. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

6 entries
  1. 01

    Adaptive Vision-Language Grasping via Composable Foundation Priors and Generalizable Grasp Synthesis

    Sixu Yan, Shikang Wang, Binhua Huang, Xuanlai Tang, Guohua Fan, Fan Huang, Haoxuan Li, Yongkang Li, Yuhan Li,...

    cs.RO · cs.AI · cs.CV

    This paper proposes AdaRoboVLG, a task-adaptive Vision-Language-Grasp (VLG) framework that supports generalizable grasp synthesis across different robotic hands. Unlike existing VLG methods that tightly couple foundation models with end-to-end grasp policies, AdaRoboVLG learns an efficient generalizable base policy that generates and evaluates physically feasible grasp candidates through explicit kinematic mapping and force-closure-based...

    arxiv.org/abs/2609.04096 · PDF

  2. 02

    FWBC-VLA: Force-Aware Whole-Body Compensation for Contact-Rich Loco-Manipulation

    Yutian Zhang, Siyuan Ma, Liwen Yang, Yang Li, Ce Hao, Haozhen Chi, Dong We, Qiaojun Yu, Dibo Hou

    cs.RO · cs.AI

    Contact-rich loco-manipulation requires a bridge between semantic action generation and physical interaction control. Existing Vision-language-action (VLA) models generate task-level actions from visual and linguistic observations, but cannot interpret the physical interactions induced by those actions. While the whole-body control (WBC) policy can stabilize the robot, it cannot distinguish task-relevant interaction forces from forces induced...

    arxiv.org/abs/2609.03889 · PDF

  3. 03

    FailBench: How Reliable are VLMs at Judging Robot Task Success?

    Zaruhi Navasardyan, Tatul Danielyan, Hrant Davtyan

    cs.RO · cs.AI

    Vision-Language Models (VLMs) are increasingly used to evaluate robot manipulation outcomes, but existing benchmarks offer limited evidence of cross-domain generalization. We introduce FailBench, a benchmark for robot failure detection comprising 2,197 manipulation attempts across 14 public sources (12 real-world, 2 simulated). In FailBench, 75% of failures occur naturally, and six real-world sources come from non-failure-detection datasets....

    arxiv.org/abs/2609.03611 · PDF

  4. 04

    Toward Physically Grounded JEPA World Models for Goal-Conditioned Robotic Planning

    Muyuan Liu, Yue Huang, Zheng Liang, Xiang Gao

    cs.RO · cs.AI · cs.LG

    Action-conditioned JEPA world models enable planning toward visually specified goals without reconstructing future pixels, yet latent prediction alone does not explicitly encourage the learned representations to retain information relevant to robotic control. We introduce an end-to-end JEPA world model that augments latent prediction with inverse dynamics (IDM) and state alignment (SA). While inverse dynamics discourages latent collapse and...

    arxiv.org/abs/2609.03565 · PDF

  5. 05

    BRIDGE: An Open-Source Humanoid Platform via Morphology-Control Co-Design for Physical AI

    Jianren Wang, Letian Qian, Zikai Wang, Weiwei Wu, Junjie Zong, Abhinav Gupta, Deepak Pathak

    cs.RO · cs.AI

    Developing humanoid robots capable of leveraging human behavioral data is essential for general-purpose embodiment, yet conventional development remains bottlenecked by a decoupled paradigm that isolates hardware design from whole-body control. This approach leads to suboptimal systems that compromise human-like fluidity and agility. To bridge this gap, we introduce a data-driven morphology-control co-design framework that optimizes humanoid...

    arxiv.org/abs/2609.03497 · PDF

  6. 06

    Air-Ground Collaborative Vision-and-Language Navigation via Shared Bird's-Eye Maps

    Shuning Zhang, Liang Li, Yunheng Wang, Tao Wang, Yihang Kang, Renjing Xu

    cs.RO · cs.AI

    Air-ground collaborative Vision-and-Language Navigation (VLN) pairs an unmanned aerial vehicle (UAV) with a global bird's-eye view and an unmanned ground vehicle (UGV) with a local first-person view, yet the setting remains largely unexplored: existing training-free methods solve single-agent tasks but offer no collaboration mechanism, and a recent CARLA-Air evaluation found no stable cooperative behavior across five state-of-the-art VLA...

    arxiv.org/abs/2609.03483 · PDF

This edition is part of The Daily Abstract — cs.RO archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.