cs.RO · 2026-09-09 · No. 110

Robotics, 2026-09-09.

10 new papers in cs.RO. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

10 entries
  1. 01

    TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model

    Anqi Li, Yuxin Chen, Zhaobo Li, Zhuo Cao, Junli Ren, Masayoshi Tomizuka, Dhruv Shah

    cs.RO · cs.AI

    We study the problem of navigating cluttered indoor environments with a humanoid robot. Unlike conventional methods that model navigation as a 2D path planning problem, humanoid traversal in cluttered environments requires continuous geometry-aware whole-body adaptation, including coordinated arm placement, torso adjustment, and gait modulation for collision-free movement through complex 3D spaces. We introduce TANGO, the first whole-body...

    arxiv.org/abs/2609.09158 · PDF

  2. 02

    DeCAL: Towards Physically-Grounded Dexterous Vision-Language-Action Models via Contact-Aware Latent Co-Imagination

    Yankai Fu, Ning Chen, Junkai Zhao, Heng Zhang, Guocai Yao, Pengwei Wang, Zhongyuan Wang, Shanghang Zhang

    cs.RO · cs.AI

    Dexterous manipulation involves contact-rich and fine-grained interactions with the physical world, posing significant challenges for existing vision-language-action (VLA) models due to severe visual occlusions and complex contact dynamics. While recent works have incorporated tactile sensing into robotic manipulation, most approaches still rely on homogeneous multimodal fusion, lacking adaptive tactile integration and explicit modeling of...

    arxiv.org/abs/2609.09119 · PDF

  3. 03

    Ostrich: Taking Large Strides Through Stiff Contact in Differentiable Dynamics

    Aleš Kučera, Karel Zimmermann

    cs.RO · cs.GR · cs.LG

    Three properties determine whether a differentiable simulator can drive gradient-based optimization through contact: simulation accuracy, gradient reliability, and per-iteration cost. Tape-based engines such as MJX and Newton Semi-Implicit require timesteps small enough to keep contacts numerically tractable, and their backpropagation memory grows linearly with the number of timesteps T. Surrogate models bound memory by approximating contact...

    arxiv.org/abs/2609.08800 · PDF

  4. 04

    BIFTA: Brain-Inspired Few-Shot Tactile Adaptation for Unknown Sensors

    Boheng Liu, Ziyu Li, Xia Wu

    cs.RO · cs.AI

    Advances in tactile sensing have made contact-rich perception possible, accelerating progress in robotic manipulation, material understanding, and embodied interaction. However, because optical design, elastomer mechanics, and imaging geometry differ substantially across tactile sensors, models trained on known sensor types can suffer an abrupt performance collapse on unknown sensors. To address this problem, we propose the Brain-Inspired...

    arxiv.org/abs/2609.08673 · PDF

  5. 05

    Learning to build covering structures with continuous adjustments

    Gabriel Vallat, Maryam Kamgarpour, Stefana Parascho

    cs.RO · cs.LG

    Robotic construction offers the potential to use materials more efficiently and create complex geometries, but current methods rely on rigid, high-precision plans that cannot accommodate the tolerances, inaccuracies, and unexpected changes inherent in physical fabrication. In this work, we introduce a reinforcement learning approach that forgoes predefined plans entirely, instead generating construction sequences adaptively as the structure...

    arxiv.org/abs/2609.08669 · PDF

  6. 06

    CASD: Chunk-Aligned Semantic Distillation for Multi-StageRobot Manipulation

    Tinghe Ding, Jiahao Li, He Wang

    cs.RO · cs.AI

    An action chunk can span several stages of a manipulation task, yet a label for its first step describes only the current stage. We introduce Chunk-Aligned Semantic Distillation (CASD), which derives semantic targets for entire action chunks. An offline vision--language model segments demonstrations into described stages. Their occupancy within each action chunk determines a weighted semantic target, including transitions between stages. A...

    arxiv.org/abs/2609.08638 · PDF

  7. 07

    RoboCousin: Build Your Own Simulation Playground for Robust Bimanual Robotic Manipulation

    Jingxuan Zhu, Jingyi Li, LiangLiang Chen, Zhiyuan Jing, Jidong Zhang, Hongming Li

    cs.RO · cs.AI

    Bimanual manipulation policies require large and diverse training datasets, yet collecting demonstrations on physical robots is expensive and difficult to scale. Simulation can generate data efficiently, but existing pipelines typically operate within closed asset libraries and predefined scenes: adding a newly observed object or environment still requires substantial effort to reconstruct geometry, specify physical and semantic properties,...

    arxiv.org/abs/2609.08339 · PDF

  8. 08

    A Multi-Modal Perception Pipeline for Object Detection and Tracking in Autonomous Racing

    Davide Malvezzi, Michele Pestarino, Vittoria Cavicchioli, Valentina La Gamba, Silvia Severi, Fabio Bagni, Luca...

    cs.RO · cs.AI

    Object detection and tracking are fundamental components of perception systems for autonomous driving. Achieving robust performance under adverse conditions such as limited visibility, sensor noise, and failures remains an open challenge, particularly in autonomous racing, where vehicles operate at very high speeds, experience strong vibrations, and interact under small safety margins. This paper presents a multi-modal late-fusion perception...

    arxiv.org/abs/2609.08338 · PDF

  9. 09

    CALIPER: Clean Scenes Cannot Rank Physical Inference in Pretrained Visual Representations

    Aman Mehta, Riya Baviskar

    cs.RO · cs.AI · cs.CV

    How far a pushed object slides depends on its mass and friction, which no single image reveals. Pretrained visual encoders are increasingly used as the perception front end of world models for manipulation, and their physical competence is assessed with perturbation benchmarks and linear probes, almost always in a clean, fixed-camera scene. We show that these assessments cannot distinguish an encoder that infers physics from one that does...

    arxiv.org/abs/2609.08250 · PDF

  10. 10

    3DWay: Generalizing Robot Manipulation via 3D Consistent Waypoints

    Ziqin Huang, Yingyue Li, Chenyangguang Zhang, Ruida Zhang, Yuxin Chen, Gu Wang, Xingyu Liu, Masayoshi Tomizuka, Xiangyang Ji

    cs.RO · cs.AI

    Intermediate representations are key to bridging the modality gap between generalizable manipulation policies and large-scale pretrained vision-language models (VLMs). Among these, trajectory-based representations compactly represent motion-relevant cues, yet most existing approaches predict trajectories in 2D image space, resulting in intrinsic 3D ambiguity. Moreover, using 2D trajectories with depth still leaves the free-space waypoints...

    arxiv.org/abs/2609.08224 · PDF

This edition is part of The Daily Abstract — cs.RO archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.