<?xml version="1.0" encoding="UTF-8"?>
<?xml-stylesheet type="text/xsl" href="/feed.xsl"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <title>Khenda Resources Radar</title>
  <subtitle>Recent humanoid-robotics research from arXiv: manipulation, VLA models, imitation learning, locomotion, and more.</subtitle>
  <id>https://www.khendarobotics.com/resources</id>
  <link href="https://www.khendarobotics.com/resources" />
  <link rel="self" href="https://www.khendarobotics.com/research.xml" />
  <updated>2026-08-11T06:47:57.940Z</updated>
  <entry>
    <title>RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance</title>
    <id>arxiv:2608.09853</id>
    <link href="https://arxiv.org/abs/2608.09853" />
    <published>2026-08-10T17:09:37.000Z</published>
    <updated>2026-08-10T17:09:37.000Z</updated>
    <author><name>Dongchi Huang</name></author>
    <author><name>Hongyin Zhang</name></author>
    <author><name>Bohan Hou</name></author>
    <author><name>Siteng Huang</name></author>
    <author><name>Zhian Su</name></author>
    <author><name>Hang Guo</name></author>
    <category term="cs.RO" />
    <summary>General-purpose reward models are increasingly the bottleneck for scaling robot learning, yet the recipe for learning value-related capabilities from large-scale heterogeneous corpora remains underexplored. Existing approaches tie supervision to task-internal anchors such as preferences or normalized progress, none of which transfer cleanly across embodiments and data sources. We introduce RynnValue, an open-source value foundation model for robotic manipulation that replaces these anchors with temporal distance, the directed cost-to-go from an observation to the language-specified goal. Beca…</summary>
  </entry>
  <entry>
    <title>RoboSeg: Online Part-Level Semantic Reconstruction for Robotic Manipulation via a Single Eye-in-Hand Camera</title>
    <id>arxiv:2608.09778</id>
    <link href="https://arxiv.org/abs/2608.09778" />
    <published>2026-08-10T16:05:38.000Z</published>
    <updated>2026-08-10T16:05:38.000Z</updated>
    <author><name>Zhaochen Lan</name></author>
    <author><name>Mengxiang Lin</name></author>
    <category term="cs.RO" />
    <summary>Robotic manipulation requires perception systemsthat identify actionable parts such as handles, rims, triggers,and tool tips, not merely object categories or point clouds. This paper presents RoboSeg, a part-level semantic reconstructionsystem that links vision-language model (VLM) functional-partdiscovery, asynchronous online RGB-D semantic reconstruc-tion, and task-oriented grasp generation without requiring CAD models or pre-scanned meshes. RoboSeg queries a VLM onthe initial RGB observation to obtain compact functional part prompts, then scans with two asynchronous streams: a high-frequen…</summary>
  </entry>
  <entry>
    <title>SLIM-0.5B: Learning Action-Grounded Predictive Latents for Robot Manipulation</title>
    <id>arxiv:2608.09771</id>
    <link href="https://arxiv.org/abs/2608.09771" />
    <published>2026-08-10T15:58:39.000Z</published>
    <updated>2026-08-10T15:58:39.000Z</updated>
    <author><name>Jingkai Wang</name></author>
    <author><name>Zihan Tang</name></author>
    <author><name>Gu Zhang</name></author>
    <author><name>Mingyu Cao</name></author>
    <author><name>Jiapeng Chen</name></author>
    <author><name>Jingjiao Zhao</name></author>
    <category term="cs.RO" />
    <summary>Vision-language-action policies rely on large multimodal backbones to jointly perform perception, language conditioning, and action generation at every control step. Much of this capacity supports open-domain semantics, whereas continuous robot manipulation primarily requires compact representations of observations, actions, and the transitions induced by actions. Pixel-level world models provide another route, but predicting visual details irrelevant to control can be unnecessarily expensive. We propose SLIM (Self-supervised Latent Interaction Model), a compact 0.5B-parameter latent interact…</summary>
  </entry>
  <entry>
    <title>Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition</title>
    <id>arxiv:2608.09762</id>
    <link href="https://arxiv.org/abs/2608.09762" />
    <published>2026-08-10T15:54:25.000Z</published>
    <updated>2026-08-10T15:54:25.000Z</updated>
    <author><name>Changhao Li</name></author>
    <author><name>Yifang Zhang</name></author>
    <author><name>Heng Zhang</name></author>
    <author><name>Davide Torielli</name></author>
    <author><name>Damiano Gasperini</name></author>
    <author><name>Arturo Laurenzi</name></author>
    <category term="cs.RO" />
    <summary>Real-world online reinforcement learning (RL) provides a promising approach for training robotic manipulation policies directly in the physical world, avoiding the sim-to-real gap and enabling continuous policy refinement through human-in-the-loop interaction. Recent methods have demonstrated sample-efficient learning through human intervention but remain limited to small randomization ranges and encounter challenges with the non-stationarity induced by concurrently training multiple agents. To address these limitations, we introduce a unified framework that combines centralized training with…</summary>
  </entry>
  <entry>
    <title>TAMS: Task-Aware Multi-View Adaptive Streaming for Wireless Telerobotic Manipulation</title>
    <id>arxiv:2608.09731</id>
    <link href="https://arxiv.org/abs/2608.09731" />
    <published>2026-08-10T15:30:51.000Z</published>
    <updated>2026-08-10T15:30:51.000Z</updated>
    <author><name>Zexin Deng</name></author>
    <author><name>Zhenhui Yuan</name></author>
    <author><name>Lu Tian</name></author>
    <author><name>Subhash Lakshminarayana</name></author>
    <author><name>Longhao Zou</name></author>
    <category term="cs.RO" />
    <summary>Wireless telerobotic manipulation relies on timely multi-view video feedback, but the available uplink bandwidth is often limited and dynamic. This paper presents Task-Aware Multi-View Adaptive Streaming (TAMS), a system that allocates video bitrate according to the current manipulation phase. TAMS infers task phase from lightweight robot-side signals and prioritizes the camera view most relevant to the operator while preserving baseline visibility for secondary views. Experiments on a six-degree-of-freedom (6-DoF) teleoperation testbed under three constrained network conditions show that TAM…</summary>
  </entry>
  <entry>
    <title>World Tokens: Enhancing Embodied Policies with Training-Time World Modeling</title>
    <id>arxiv:2608.09730</id>
    <link href="https://arxiv.org/abs/2608.09730" />
    <published>2026-08-10T15:30:38.000Z</published>
    <updated>2026-08-10T15:30:38.000Z</updated>
    <author><name>Qu Tang</name></author>
    <author><name>Benhui Zhuang</name></author>
    <author><name>Bo Yuan</name></author>
    <author><name>Xue Yu</name></author>
    <author><name>Longteng Guo</name></author>
    <author><name>Junlan Feng</name></author>
    <category term="cs.CV" />
    <summary>Vision-language-action (VLA) models are a widely adopted paradigm for embodied policies. They excel at efficient closed-loop control but do not explicitly model how physical scenes evolve as a task unfolds. Recently emerging world-action models (WAMs) leverage pretrained video world models to capture spatiotemporal evolution, yet retaining future generation or a large video backbone in the control loop substantially increases inference cost. We introduce World Tokens, an embodied policy architecture built around a World Adapter that bridges visual-language understanding, world-dynamics modeli…</summary>
  </entry>
  <entry>
    <title>Robotic Fabric Alignment System for Sewing Using Global Local Weighted ICP</title>
    <id>arxiv:2608.09528</id>
    <link href="https://arxiv.org/abs/2608.09528" />
    <published>2026-08-10T12:30:17.000Z</published>
    <updated>2026-08-10T12:30:17.000Z</updated>
    <author><name>Wenbo Dong</name></author>
    <author><name>Dipankar Bhattacharya Member</name></author>
    <author><name>Kai Tang</name></author>
    <author><name>Akinari Kobayashi</name></author>
    <author><name>Fuyuki Tokuda</name></author>
    <author><name>Akira Seino</name></author>
    <category term="cs.RO" />
    <summary>Accurate fabric alignment is a critical step that must be performed before sewing. This paper presents a novel automated fabric alignment system. The system estimates the poses of top and bottom fabric panels, lying flat and wrinkle-free in arbitrary positions, using a new Global Local Weighted Iterative Closest Point (GLW-ICP) method. The system then manipulates the top panel to achieve precise alignment at both edges and sewing lines. Unlike conventional approaches, GLW-ICP robustly aligns both global edges and local sewing lines by globally aligning fabric edge points and locally aligning…</summary>
  </entry>
  <entry>
    <title>Graph-Guided Safe Diffuser: Topological Graph Guidance for Safe Diffusion Planning</title>
    <id>arxiv:2608.09484</id>
    <link href="https://arxiv.org/abs/2608.09484" />
    <published>2026-08-10T11:49:24.000Z</published>
    <updated>2026-08-10T11:49:24.000Z</updated>
    <author><name>Nakgyu Yang</name></author>
    <author><name>KwangBin Lee</name></author>
    <author><name>SooJean Han</name></author>
    <category term="cs.RO" />
    <summary>Many diffusion-based planners enforce safety through inference-time guidance, but such interleaved trajectory deformations often degrade kinematic feasibility due to manifold rupture. We propose Graph-Guided Safe Diffuser (G2SD), a hierarchical framework that leverages a high-level topological graph planner to guide a low-level diffusion model. G2SD enforces safety at a structural level by abstracting the data manifold into a learned latent graph, on which high-level planning is performed. Continuous trajectories are generated by diffusion planners, which are conditioned on the graph node rep…</summary>
  </entry>
  <entry>
    <title>RecoverFly: A Failure-Aware Reinforcement Learning Post-Training Framework for Aerial Vision-Language Navigation</title>
    <id>arxiv:2608.09467</id>
    <link href="https://arxiv.org/abs/2608.09467" />
    <published>2026-08-10T11:37:46.000Z</published>
    <updated>2026-08-10T11:37:46.000Z</updated>
    <author><name>Boxiong Wang</name></author>
    <author><name>Hui Kang</name></author>
    <author><name>Geng Sun</name></author>
    <author><name>Jiahui Li</name></author>
    <author><name>Chao Yu</name></author>
    <author><name>Daxin Tian</name></author>
    <category term="cs.CV" />
    <summary>Unmanned aerial vehicle vision-language navigation (UAV-VLN) requires agents to translate visual observations and language instructions into reliable flight actions in complex environments. Although recent end-to-end UAV vision-language-action (UAV-VLA) policies reduce reliance on separately designed perception, planning, and control modules, their behavior-cloning objectives provide limited corrective supervision for interactive closed-loop execution. Reinforcement learning (RL) offers a promising solution, while its effectiveness is constrained by inefficient use of samples, long-tailed sce…</summary>
  </entry>
  <entry>
    <title>VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction</title>
    <id>arxiv:2608.09448</id>
    <link href="https://arxiv.org/abs/2608.09448" />
    <published>2026-08-10T11:22:54.000Z</published>
    <updated>2026-08-10T11:22:54.000Z</updated>
    <author><name>Hongjin Ji</name></author>
    <author><name>Guoyang Xia</name></author>
    <author><name>Luoyang Sun</name></author>
    <author><name>Fangxiang Feng</name></author>
    <author><name>Lei Ren</name></author>
    <category term="cs.RO" />
    <summary>Test-time training (TTT) offers a lightweight way to adapt vision--language--action (VLA) policies from unlabeled deployment streams, but it remains difficult to use reliably in closed-loop manipulation. A shared adaptation space can mix incompatible task corrections, while an online update can alter subsequent actions before its consequences are known. We introduce a reliable TTT framework for VLA policies (VANE). VANE conditions prompt adaptation on the current vision--language context and learns from the future visual consequences of executed actions. Candidate updates are isolated from th…</summary>
  </entry>
  <entry>
    <title>Skills in Weights, Memory in Code: Hybrid Learning for Memory-Dependent Robot Manipulation</title>
    <id>arxiv:2608.09410</id>
    <link href="https://arxiv.org/abs/2608.09410" />
    <published>2026-08-10T10:35:47.000Z</published>
    <updated>2026-08-10T10:35:47.000Z</updated>
    <author><name>Yunhao Zhao</name></author>
    <author><name>Zhenyang Ni</name></author>
    <author><name>Haoyang Chen</name></author>
    <author><name>Ruohan Zhang</name></author>
    <author><name>Qi Zhu</name></author>
    <category term="cs.RO" />
    <summary>Modern vision-language-action (VLA) policies have acquired broad manipulation skills, but typically generate each action chunk from the current observation or a short fixed-length history. However, real-world manipulation is often non-Markovian, requiring robots to retain and reason over task-relevant information from long-horizon interaction histories to determine the next action. To address this challenge, we propose HyMeS, a hybrid learning framework that leverages the reasoning and memory-management capabilities of coding agents to steer a Markovian VLA for memory-dependent manipulation.…</summary>
  </entry>
  <entry>
    <title>JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling</title>
    <id>arxiv:2608.09381</id>
    <link href="https://arxiv.org/abs/2608.09381" />
    <published>2026-08-10T09:57:54.000Z</published>
    <updated>2026-08-10T09:57:54.000Z</updated>
    <author><name>Yihan Lin</name></author>
    <author><name>Jiawei He</name></author>
    <author><name>Shifeng Bao</name></author>
    <author><name>Chen Zhao</name></author>
    <author><name>Yang Li</name></author>
    <author><name>Xiaobo Wang</name></author>
    <category term="cs.RO" />
    <summary>Robust robot control benefits from explicitly modeling state transitions, but video-generation world action models (WAMs) introduce substantial deployment cost. Existing latent WAMs avoid explicit future generation, but often compress predictive representations or separate predictive modeling from the representations used for action generation. We introduce JEPA-WAM, a latent WAM built in a pretrained V-JEPA space, which couples latent transition prediction with continuous action generation through a shared predictor. JEPA-WAM predicts a spatially structured joint current-future target that c…</summary>
  </entry>
  <entry>
    <title>SAFE-CHEM: Uncertainty-Aware Policy Switching for Robust Robotic Chemistry</title>
    <id>arxiv:2608.09303</id>
    <link href="https://arxiv.org/abs/2608.09303" />
    <published>2026-08-10T08:51:43.000Z</published>
    <updated>2026-08-10T08:51:43.000Z</updated>
    <author><name>Laura Jones</name></author>
    <author><name>Shazil Shahzad</name></author>
    <author><name>Ayesha Sana</name></author>
    <author><name>Gabriella Pizzuto</name></author>
    <category term="cs.RO" />
    <summary>The deployment of autonomous robotic systems in chemistry laboratories is accelerating experimental workflows and providing the foundational data for AI-driven scientific discovery. However, despite the success of data-driven methods in acquiring dexterous skills, safety remains a primary barrier to their deployment in high-risk domains, such as early-stage materials chemistry experiments. Specifically, learning-based policies frequently struggle to distinguish between safe and unsafe actions, leading to overconfident extrapolation and potentially catastrophic failures. To mitigate these safe…</summary>
  </entry>
  <entry>
    <title>Ultra-Low-Impedance Robotic Gripper for High-Bandwidth and Transparent Physical Interaction</title>
    <id>arxiv:2608.09198</id>
    <link href="https://arxiv.org/abs/2608.09198" />
    <published>2026-08-10T07:11:51.000Z</published>
    <updated>2026-08-10T07:11:51.000Z</updated>
    <author><name>Joon Lee</name></author>
    <author><name>Ari Choi</name></author>
    <author><name>Seokhwan Jeong</name></author>
    <category term="cs.RO" />
    <summary>Conventional robotic grippers often use high-ratio transmissions to generate grasping torque and external force sensors to measure physical interaction. High-ratio transmissions increase friction, reflected inertia, and mechanical impedance, while external sensors add hardware complexity. To address these trade-offs, this study proposes a novel 9-DOF, three-fingered Differential Direct-Drive (DDD) gripper that combines DD motors with a low-ratio (1:2) differential transmission. The mechanism centralizes actuator mass at the base to minimize moving-link inertia, while the differential architec…</summary>
  </entry>
  <entry>
    <title>Intuitive Directional Sense Presentation to the Torso Using McKibben-Based Surface Haptic Sensation in Immersive Space</title>
    <id>arxiv:2608.09177</id>
    <link href="https://arxiv.org/abs/2608.09177" />
    <published>2026-08-10T06:42:24.000Z</published>
    <updated>2026-08-10T06:42:24.000Z</updated>
    <author><name>Kenta Yokoe</name></author>
    <author><name>Tadayoshi Aoyama</name></author>
    <author><name>Yuki Funabora</name></author>
    <author><name>Masaru Takeuchi</name></author>
    <author><name>Yasuhisa Hasegawa</name></author>
    <category term="cs.HC" />
    <summary>In recent years, systems that utilize immersive space have been developed in various fields. Immersive spaces often contain considerable amounts of visual information; therefore, users often fail to obtain their desired information. Therefore, various methods have been developed to guide users toward haptic sensations. However, many of these methods have limitations in terms of the intuitive perception of haptic sensation and require practice for familiarization with haptic sensation. Fabric actuators are wearable haptic devices that combine fabric and McKibben artificial muscles to provide w…</summary>
  </entry>
  <entry>
    <title>Particle-Based Conformal Prediction for Contact-Aware Uncertainty Calibration in Stratified Configuration Spaces</title>
    <id>arxiv:2608.09166</id>
    <link href="https://arxiv.org/abs/2608.09166" />
    <published>2026-08-10T06:22:10.000Z</published>
    <updated>2026-08-10T06:22:10.000Z</updated>
    <author><name>Luís Marques</name></author>
    <author><name>Kristian Popov</name></author>
    <author><name>Dmitry Berenson</name></author>
    <category term="cs.RO" />
    <summary>Reliable uncertainty representation is essential for deploying autonomous systems that interact with their environment, as robots must reason about how uncertainty arising from both stochasticity and model mismatch is impacted by contacts with obstacles (e.g., when navigating through a cluttered environment or inserting a part into an assembly). We propose Calibrated Particle-sets for Trans-dimensional Uncertainty Representation (CaPTURe), a geometry-aware, conformal prediction-based algorithm that generates probabilistically valid prediction regions of the unknown future system configuration…</summary>
  </entry>
  <entry>
    <title>SpeedTuning: Speeding Up Policy Execution with Lightweight Reinforcement Learning</title>
    <id>arxiv:2608.09138</id>
    <link href="https://arxiv.org/abs/2608.09138" />
    <published>2026-08-10T05:31:41.000Z</published>
    <updated>2026-08-10T05:31:41.000Z</updated>
    <author><name>David D. Yuan</name></author>
    <author><name>Tony Z. Zhao</name></author>
    <author><name>Kaylee Burns</name></author>
    <author><name>Chelsea Finn</name></author>
    <category term="cs.RO" />
    <summary>While learned robotic policies hold promise for advancing generalizable manipulation, their practical deployment is often hindered by suboptimal execution speeds. Imitation learning policies are inherently limited by hardware constraints and the speed of the operator during data collection. In addition, there are no established methods for accelerating policies learned via imitation, and the empirical relationship between execution speed and task success remains underexplored. To address these issues, we introduce SpeedTuning, a reinforcement learning framework specifically designed to enhanc…</summary>
  </entry>
  <entry>
    <title>High Fidelity Capture, Reconstruction, and Transfer of Human Demonstrations for Robot-Assisted Bathing</title>
    <id>arxiv:2608.09127</id>
    <link href="https://arxiv.org/abs/2608.09127" />
    <published>2026-08-10T05:07:08.000Z</published>
    <updated>2026-08-10T05:07:08.000Z</updated>
    <author><name>Arjun S. Lakshmipathy</name></author>
    <author><name>Jonathan P. King</name></author>
    <author><name>Ethan Zuo</name></author>
    <author><name>Rohit Satishkumar</name></author>
    <author><name>Hongyi Chen</name></author>
    <author><name>Jeffrey Ichnowski</name></author>
    <category term="cs.RO" />
    <summary>Despite the demand for robots in high-value clinical tasks like bathing, contemporary systems still lack the safety and reliability required for complex, sustained physical interaction with humans. A key challenge hindering the development of such systems is that collecting, understanding, and effectively transferring highly dynamic, contact-rich human bathing demonstrations is difficult, even with modern motion and tactile sensing equipment. We present a straightforward, but effective framework for doing so with high fidelity by utilizing contact regions as a key processing primitive. We use…</summary>
  </entry>
  <entry>
    <title>Trajectory Divergence Horizon Decision for Reliable Dual-Arm Surgical Subtask Manipulation</title>
    <id>arxiv:2608.09125</id>
    <link href="https://arxiv.org/abs/2608.09125" />
    <published>2026-08-10T05:04:02.000Z</published>
    <updated>2026-08-10T05:04:02.000Z</updated>
    <author><name>Mingwu Su</name></author>
    <author><name>Guankun Wang</name></author>
    <author><name>Jinsong Lin</name></author>
    <author><name>Rulin Zhou</name></author>
    <author><name>Ziyi Hao</name></author>
    <author><name>Zhiwei Fang</name></author>
    <category term="cs.RO" />
    <summary>Surgical robotic systems are increasingly being adopted as clinical workload rises, motivating autonomous solutions for repetitive manipulation subtasks. Learning-based controllers improve generalization compared with rule-based and analytic approaches, but most are trained for individual tasks and remain difficult to reuse across procedures. Vision-Language-Action (VLA) models provide a unified framework that integrates visual perception, language grounding, and action generation, offering a promising path toward more composable surgical autonomy. However, existing VLA policies rely on fixed…</summary>
  </entry>
  <entry>
    <title>From Recovery to Drop-off: How Action Post-training Reduces a VLM&apos;s Late-Layer Depth Decodability</title>
    <id>arxiv:2608.08904</id>
    <link href="https://arxiv.org/abs/2608.08904" />
    <published>2026-08-09T20:31:52.000Z</published>
    <updated>2026-08-09T20:31:52.000Z</updated>
    <author><name>Alexander Hackett</name></author>
    <author><name>Arnaud Denis-Remillard</name></author>
    <author><name>Axel Cassou</name></author>
    <category term="cs.CV" />
    <summary>How much of a vision-language model&apos;s (VLM) spatial understanding remains after the action post-training process of building a vision-language-action model (VLA)? We probe depth perception, a primitive of spatiogeometric understanding, from every decoder layer of a weight-matched open-source base VLM/VLA pair: Molmo2-ER and MolmoAct2-LIBERO. First, the VLA decodes depth worse at every layer, a persistent gap we call the floor. Second, the degradation is not uniform: while the base VLM&apos;s depth decodability improves through its final layers, the VLA&apos;s collapses, an additional late-layer drop we…</summary>
  </entry>
  <entry>
    <title>Hierarchical Topology-Aware Planning and Control of Underwater Vehicle-Manipulator Systems in Confined Environments</title>
    <id>arxiv:2608.08871</id>
    <link href="https://arxiv.org/abs/2608.08871" />
    <published>2026-08-09T19:22:51.000Z</published>
    <updated>2026-08-09T19:22:51.000Z</updated>
    <author><name>Mohamed Abdelwahab</name></author>
    <author><name>Ruggero Carli</name></author>
    <author><name>Damiano Varagnolo</name></author>
    <author><name>Alberto Dalla Libera</name></author>
    <category term="cs.RO" />
    <summary>This paper addresses autonomous intervention with an underwater vehicle--manipulator system (UVMS) in confined, cluttered, and partially known environments, where poor maneuverability, narrow passages, and uncertain execution may cause the robot to enter unrecoverable regions. We propose MANTA, a three-layer hierarchical planning-and-control framework that couples passage accessibility, manipulation feasibility, and closed-loop execution. The first layer performs global connectivity reasoning in a conservative reduced base space to extract traversable corridor candidates toward the task regio…</summary>
  </entry>
  <entry>
    <title>SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models</title>
    <id>arxiv:2608.08839</id>
    <link href="https://arxiv.org/abs/2608.08839" />
    <published>2026-08-09T17:59:35.000Z</published>
    <updated>2026-08-09T17:59:35.000Z</updated>
    <author><name>Junjie He</name></author>
    <author><name>Junfeng Li</name></author>
    <author><name>Zhide Zhong</name></author>
    <author><name>Haodong Yan</name></author>
    <author><name>Ruixin Li</name></author>
    <author><name>Yangyang Zheng</name></author>
    <category term="cs.RO" />
    <summary>World-Action Models (WAMs) have emerged as a promising paradigm for robotic manipulation. However, most existing WAMs generate future videos and actions by relying mainly on visual cues rather than language instructions, since off-the-shelf text encoders embed instructions independently of visual observations. As a result, the videos predicted by these WAMs are often semantically misaligned with their corresponding language instructions, which degrades the accuracy of the predicted actions. To overcome this limitation, we propose SG-WAM, a semantic guidance method for world-action models that…</summary>
  </entry>
  <entry>
    <title>PEEL: Parallel Extraction for Long-Horizon Disassembly Planning via Scale-Invariant Sampling</title>
    <id>arxiv:2608.08773</id>
    <link href="https://arxiv.org/abs/2608.08773" />
    <published>2026-08-09T15:44:26.000Z</published>
    <updated>2026-08-09T15:44:26.000Z</updated>
    <author><name>Servet B. Bayraktar</name></author>
    <author><name>Andreas Orthey</name></author>
    <author><name>Zachary Kingston</name></author>
    <author><name>Marc Toussaint</name></author>
    <category term="cs.RO" />
    <summary>Long-horizon multi-part object disassembly requires robots to compute feasible sequences of collision-free removal motions, even in the presence of tight, narrow escape corridors. To efficiently solve such disassembly problems, we propose Parallel Extraction for Long-Horizon Disassembly (PEEL), an algorithm which efficiently computes disassembly motions for object assemblies and feeds them to a robot manipulator for execution. PEEL uses sampling-based motion planning to compute single-object motions through the use of a scale-invariant sampling scheme, where the object scale is estimated in a…</summary>
  </entry>
  <entry>
    <title>OnEvoMemory: Evolving Memory through Online Robot Rollouts for Pretrained Robot Policies</title>
    <id>arxiv:2608.08749</id>
    <link href="https://arxiv.org/abs/2608.08749" />
    <published>2026-08-09T14:53:19.000Z</published>
    <updated>2026-08-09T14:53:19.000Z</updated>
    <author><name>Zhongxi Chen</name></author>
    <author><name>Shenqi Zong</name></author>
    <category term="cs.RO" />
    <summary>Long-horizon robot manipulation requires policies to track completed subtasks and critical interaction events. However, existing memory mechanisms heavily rely on external models or predefined update rules. To address this, we propose OnEvoMemory, a value-guided memory module for pretrained robot policies. It maintains recent context, high-value experiences, and salient transitions, while learning which experiences should be retained from trajectory outcomes. Offline demonstrations initialize the memory prior, whereas successful and unsuccessful online rollouts refine memory selection, helpin…</summary>
  </entry>
  <entry>
    <title>WA-SpecDec: World-Aware Speculative Decoding for Vision-Language-Action Models</title>
    <id>arxiv:2608.08725</id>
    <link href="https://arxiv.org/abs/2608.08725" />
    <published>2026-08-09T14:17:46.000Z</published>
    <updated>2026-08-09T14:17:46.000Z</updated>
    <author><name>Zikang Wen</name></author>
    <author><name>Yuning Zhang</name></author>
    <author><name>Dong Yuan</name></author>
    <category term="cs.RO" />
    <summary>Vision-language-action (VLA) policies generate robot controls autoregressively, making closed-loop latency dominated by repeated target-model forward passes. Speculative decoding reduces this cost by verifying blocks of draft action tokens in parallel, and recent VLA methods further relax token-level acceptance because small differences in action-token space often map to similar continuous controls. However, this relaxation remains scene-agnostic. A fixed token-distance tolerance treats the same action-token deviation as equally safe across states, although deviations that are harmless in fre…</summary>
  </entry>
  <entry>
    <title>SAGE: SLO-Aware Adaptive Retrieval for Production RAG Systems</title>
    <id>arxiv:2608.08237</id>
    <link href="https://arxiv.org/abs/2608.08237" />
    <published>2026-08-08T17:05:36.000Z</published>
    <updated>2026-08-08T17:05:36.000Z</updated>
    <author><name>Muhammad Faizan Raza</name></author>
    <author><name>Shuo</name></author>
    <author><name>Yang</name></author>
    <author><name>Satish Mahadevan Srinivasan</name></author>
    <category term="cs.LG" />
    <summary>Retrieval-Augmented Generation (RAG) systems in production operate under strict service level objectives (SLOs) on tail latency and infrastructure cost. However, standard retrieval pipelines rely on fixed retrieval budgets that ignore query difficulty, over-retrieving for easy queries and under-serving hard ones, forcing operators to trade answer quality against SLO compliance. This paper proposes SAGE, a learned SLO-aware adaptive retrieval policy that dynamically selects the number of passages k per query. SAGE uses lightweight features derived from initial retrieval (e.g., score distributi…</summary>
  </entry>
  <entry>
    <title>Spatiotemporal Context-dependent Personalized Movement Compensation in Delayed Telemanipulation</title>
    <id>arxiv:2608.08200</id>
    <link href="https://arxiv.org/abs/2608.08200" />
    <published>2026-08-08T15:51:09.000Z</published>
    <updated>2026-08-08T15:51:09.000Z</updated>
    <author><name>Sai Jiang</name></author>
    <author><name>Zonghe Chua</name></author>
    <category term="cs.RO" />
    <summary>Communication delay remains a central challenge in telerobotics, where it disrupts visuomotor coordination and reduces task precision. Motion scaling is an effective countermeasure to delay-induced overshoot, yet typical deployments rely on uniform gains that neglect individual and contextual variability. We propose a human-centered method that fits personalized delay-, direction-, and distance-specific scaling parameters for each participant. We conducted experiments with twenty participants who performed delayed reaching tasks in a virtual simulator. Scaling gains were computed to minimize…</summary>
  </entry>
  <entry>
    <title>Multi-modal Interactive Control of Robotic Arm based on Offline Large Language Models</title>
    <id>arxiv:2608.08183</id>
    <link href="https://arxiv.org/abs/2608.08183" />
    <published>2026-08-08T15:21:57.000Z</published>
    <updated>2026-08-08T15:21:57.000Z</updated>
    <author><name>Hanxiao Chen</name></author>
    <category term="cs.RO" />
    <summary>Large Language Models (LLMs) have significantly revolutionized the modern society with numerous advanced interactions between humans and AI agents, whereas the usage of most large language models including ChatGPT are not friendly open-sourced and must require the users paying a lot for such AI services continuously. Therefore, deploying open-sourced large language models on local servers can be considered as an efficient approach to design and implement creative embodied AI algorithms with lower cost and more stable free usage. Inspired by this ordinary motivation, we originally propose and…</summary>
  </entry>
  <entry>
    <title>Auditing Instruction-Trajectory Mismatches in Multimodal Robot Demonstrations</title>
    <id>arxiv:2608.07895</id>
    <link href="https://arxiv.org/abs/2608.07895" />
    <published>2026-08-08T03:43:58.000Z</published>
    <updated>2026-08-08T03:43:58.000Z</updated>
    <author><name>Simon Holk</name></author>
    <author><name>Ryosuke Takanami</name></author>
    <author><name>Tatsuya Matsushima</name></author>
    <author><name>Yusuke Iwasawa</name></author>
    <author><name>Yutaka Matsuo</name></author>
    <author><name>Yueh-Hua Wu</name></author>
    <category term="cs.RO" />
    <summary>Robot demonstration datasets used to train vision-language-action policies can contain a subtle but harmful failure mode: trajectories that are behaviorally correct but paired with the wrong language instruction. We study post-hoc auditing of these Instruction-Trajectory Mismatches (ITMs). Unlike failed rollouts, ITMs often look plausible, and can corrupt the language-behavior mapping learned by the policy. We propose Multimodal Probabilistic Fusion (MMPF), a training-free auditing framework that treats each modality as an expert, estimates a task-label distribution from local neighborhood ag…</summary>
  </entry>
  <entry>
    <title>LUCID: Latent-Skill Unified Control via Imagined Dynamics for Long-Horizon Humanoid Loco-Manipulation</title>
    <id>arxiv:2608.07746</id>
    <link href="https://arxiv.org/abs/2608.07746" />
    <published>2026-08-07T20:26:34.000Z</published>
    <updated>2026-08-07T20:26:34.000Z</updated>
    <author><name>Cheng Guo</name></author>
    <author><name>Mingzhe Ni</name></author>
    <author><name>Angelo Cangelosi</name></author>
    <author><name>Arash Ajoudani</name></author>
    <category term="cs.LG" />
    <summary>Long-horizon humanoid loco-manipulation requires composing versatile whole-body skills and reliable high-level decision making. Existing methods often coordinate pretrained skills with scripted planners, finite-state machines or task-specific model-free policies, restricting their ability to handle complex task sequences. To address this limitation, we propose \textbf{LUCID}, a hierarchical model-based reinforcement learning framework that plans over reusable skills through imagined rollouts of a learned dynamics model. LUCID first trains a structured latent-conditioned low-level policy via a…</summary>
  </entry>
  <entry>
    <title>Hölder Signed Distance: A Differentiable, Signed, Parallelizable Metric for Robotics</title>
    <id>arxiv:2608.07707</id>
    <link href="https://arxiv.org/abs/2608.07707" />
    <published>2026-08-07T18:51:26.000Z</published>
    <updated>2026-08-07T18:51:26.000Z</updated>
    <author><name>Felipe Bartelt</name></author>
    <author><name>Ali Umut Kaypak</name></author>
    <author><name>Anthony Tzes</name></author>
    <author><name>Farshad Khorrami</name></author>
    <author><name>Luciano C. A. Pimenta</name></author>
    <author><name>Vinicius M. Gonçalves</name></author>
    <category term="cs.RO" />
    <summary>Computing distances between sets is essential in robotic motion planning and control, where differentiable gradients enable real-time optimization. The Euclidean Signed Distance Function (SDF), however, is not differentiable everywhere, and existing alternatives often sacrifice differentiability, sign information, or computational efficiency. In this letter, we introduce a novel differentiable signed distance between convex polyhedra. To this end, we first propose differentiable versions of the minimum and maximum operators, termed the Hölder minimum and Hölder maximum. We then replace the or…</summary>
  </entry>
  <entry>
    <title>Depth-Wise Probing and Pruning of the Planning Token in a Driving Vision-Language-Action Model</title>
    <id>arxiv:2608.07361</id>
    <link href="https://arxiv.org/abs/2608.07361" />
    <published>2026-08-07T16:02:34.000Z</published>
    <updated>2026-08-07T16:02:34.000Z</updated>
    <author><name>Harisankar Babu</name></author>
    <author><name>Benjamin Coors</name></author>
    <author><name>Christopher Lang</name></author>
    <author><name>Hendrik Berkemeyer</name></author>
    <author><name>Tamim Asfour</name></author>
    <author><name>Simon Foell</name></author>
    <category term="cs.RO" />
    <summary>Vision-language-action (VLA) models route driving decisions through a deep language model, but it is unclear how much of that depth the action itself requires. We study a representative driving VLA whose entire plan is carried by a single planning token that a generative planner decodes into a trajectory. Borrowing the planner as a trajectory-space logit lens, we decode the planning token from every one of the 32 decoder layers and measure two signals: the linear decodability of the navigation command and trajectory compatibility with the frozen native planner. Our diagnostic shows that seman…</summary>
  </entry>
  <entry>
    <title>Learning Fault-Tolerant Locomotion with Adaptive Gait Timing</title>
    <id>arxiv:2608.07328</id>
    <link href="https://arxiv.org/abs/2608.07328" />
    <published>2026-08-07T15:27:03.000Z</published>
    <updated>2026-08-07T15:27:03.000Z</updated>
    <author><name>Giovanbattista Gravina</name></author>
    <author><name>Luca Rossini</name></author>
    <author><name>Carlo Rizzardo</name></author>
    <author><name>Arturo Laurenzi</name></author>
    <author><name>Nikos Tsagarakis</name></author>
    <category term="cs.RO" />
    <summary>Hardware failures require legged robots to rapidly reorganize coordination and gait timing to maintain stability and mobility. This is particularly challenging for larger quadrupeds, where increased mass and tighter actuation limits reduce the feasibility of aggressive, high-frequency compensation strategies often observed on smaller platforms. In this work, we propose a deep reinforcement learning approach for fault-tolerant locomotion under actuator power loss. The method employs an asymmetric actor-critic architecture in which the critic has access to privileged information during training…</summary>
  </entry>
  <entry>
    <title>TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models</title>
    <id>arxiv:2608.07314</id>
    <link href="https://arxiv.org/abs/2608.07314" />
    <published>2026-08-07T15:09:51.000Z</published>
    <updated>2026-08-07T15:09:51.000Z</updated>
    <author><name>Ziheng Liu</name></author>
    <author><name>Quantao Yang</name></author>
    <category term="cs.RO" />
    <summary>Vision-language-action (VLA) models are commonly adapted to downstream manipulation tasks via supervised fine-tuning (SFT) or online reinforcement learning (RL) post-training. SFT is prone to distribution mismatch, and existing RL approaches typically apply a single, uniform update strategy to all model components, ignoring their distinct functional roles. We propose TEMPO, a semantic-action decoupled, two-timescale RL post-training framework for VLA models. TEMPO freezes the pretrained vision-language backbone to preserve general semantic representations, and restricts adaptation to two comp…</summary>
  </entry>
  <entry>
    <title>WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN</title>
    <id>arxiv:2608.07267</id>
    <link href="https://arxiv.org/abs/2608.07267" />
    <published>2026-08-07T14:29:00.000Z</published>
    <updated>2026-08-07T14:29:00.000Z</updated>
    <author><name>Yuehao Huang</name></author>
    <author><name>Yunzi Wu</name></author>
    <author><name>Xiaotao Zhang</name></author>
    <author><name>Xinhai Li</name></author>
    <author><name>Jiankun Dong</name></author>
    <author><name>Jiajun Lv</name></author>
    <category term="cs.AI" />
    <summary>Recent vision-language navigation (VLN) systems increasingly adapt pretrained vision-language models (VLMs) into vision-language-action (VLA) policies that map egocentric observations and language instructions directly to navigation actions. Although semantically capable, such action-centric training does not explicitly model how the agent&apos;s visual observations should evolve under its predicted motion. Generative world-action models (WAMs) jointly predict future observations and actions, yet existing WAMs for continuous VLN do not condition joint future-view and action generation on geometry-…</summary>
  </entry>
  <entry>
    <title>Representation Handoffs for OpenArm-Based Laboratory Mobile Manipulation</title>
    <id>arxiv:2608.07154</id>
    <link href="https://arxiv.org/abs/2608.07154" />
    <published>2026-08-07T12:19:50.000Z</published>
    <updated>2026-08-07T12:19:50.000Z</updated>
    <author><name>Yang Shen</name></author>
    <author><name>Chonghao Cheng</name></author>
    <author><name>Ziyi Zhao</name></author>
    <author><name>Jialuo Zhu</name></author>
    <author><name>Zhenyi Yi</name></author>
    <author><name>Qi Zhao</name></author>
    <category term="cs.RO" />
    <summary>Open-source robotics and foundation models have lowered the barrier to embodied AI, yet language-guided laboratory automation still requires reliable alignment from instructions and observations to safe actions. This field report presents an OpenArm-based mobile manipulation prototype for laboratory-style tasks, built by integrating dual OpenArm manipulators with a mobile base, vertical slide, RGB-D sensing, lidar-based mapping, ROS2/MoveIt execution, and profile-defined skill interfaces. The system is organized around representation handoffs: natural language requests are constrained into re…</summary>
  </entry>
  <entry>
    <title>Identifying the Key Biomechanical Features of Movement Adaptation during Exoskeleton-Assisted Locomotion</title>
    <id>arxiv:2608.07140</id>
    <link href="https://arxiv.org/abs/2608.07140" />
    <published>2026-08-07T12:00:45.000Z</published>
    <updated>2026-08-07T12:00:45.000Z</updated>
    <author><name>Peter Seungjune Lee</name></author>
    <author><name>Katja Mombaur</name></author>
    <category term="cs.RO" />
    <summary>The understanding of natural human adaptation during exoskeleton-assisted locomotion - particularly individual differences in adaptation behaviors and temporal progression - remains limited. In this work, we investigate temporal evolution of biomechanical variables to uncover participant-specific adaptation strategies across different exoskeleton-assisted locomotion scenarios. Nine healthy participants performed treadmill walking under three conditions: without an exoskeleton, with exoskeleton active ankle assistance, and with exoskeleton zero-torque. Lower limb kinematics, inter-joint coordi…</summary>
  </entry>
  <entry>
    <title>Detection and Ranging of Transient Extrinsic Contacts Based on 6D Dynamic Tactile Sensing</title>
    <id>arxiv:2608.07075</id>
    <link href="https://arxiv.org/abs/2608.07075" />
    <published>2026-08-07T10:27:37.000Z</published>
    <updated>2026-08-07T10:27:37.000Z</updated>
    <author><name>Haowen Zheng</name></author>
    <author><name>Yinghao Wu</name></author>
    <author><name>Fuyuan Liu</name></author>
    <author><name>Yichen Li</name></author>
    <author><name>Yitian Shao</name></author>
    <category term="cs.RO" />
    <summary>Delicate manipulation often involves transient and subtle collisions between a grasped object and the environment. While the human hand localizes these contacts effortlessly thanks to superior tactile sensitivity, robotic systems often lack the requisite resolution to acquire the information necessary for motion planning, resulting in clumsy manipulation or even task failure. Here, we propose transient extrinsic contact detection and ranging (TECDAR), a simple yet fast and efficient method for detecting and ranging extrinsic contact of grasped objects. Our design of gripper tips employs dynam…</summary>
  </entry>
  <entry>
    <title>AutoIntervene: Calibrated Intervention for Action-Chunking Imitation Learning Policies</title>
    <id>arxiv:2608.07065</id>
    <link href="https://arxiv.org/abs/2608.07065" />
    <published>2026-08-07T10:14:29.000Z</published>
    <updated>2026-08-07T10:14:29.000Z</updated>
    <author><name>Jinhe Tang</name></author>
    <author><name>Weiming Zhi</name></author>
    <category term="cs.RO" />
    <summary>Action-chunking visuomotor policies learn from demonstrations and improve temporal consistency by predicting short action sequences rather than single-step commands. Yet perception errors and execution drift can move the robot outside the demonstration distribution, while the policy continues to produce smooth action chunks that are inconsistent with the observed state. We present AutoIntervene, an online framework that selectively transfers control between an action-chunking policy and an operator during deployment. AutoIntervene evaluates proposed chunks against a visual-action support memo…</summary>
  </entry>
  <entry>
    <title>C2Dex: Contact-Consistent Reconstruction and Retargeting for Dexterous Manipulation from Monocular Video</title>
    <id>arxiv:2608.07045</id>
    <link href="https://arxiv.org/abs/2608.07045" />
    <published>2026-08-07T09:54:00.000Z</published>
    <updated>2026-08-07T09:54:00.000Z</updated>
    <author><name>Jie Ren</name></author>
    <author><name>Zhehao Jiang</name></author>
    <author><name>Yinhong Yang</name></author>
    <author><name>Haorui Jia</name></author>
    <author><name>Han Jiang</name></author>
    <author><name>Ben Li</name></author>
    <category term="cs.RO" />
    <summary>High-quality demonstrations for dexterous robot manipulation are costly and difficult to collect, whereas monocular human videos provide a scalable source of diverse manipulation behaviors. However, transferring such demonstrations to dexterous robots remains challenging: monocular hand-object interaction (HOI) reconstruction often produces temporally unstable contacts and physically implausible interactions, while conventional retargeting methods struggle to preserve task-relevant contacts and local interaction geometry across different hand embodiments. We present C2Dex, a video-to-dexterou…</summary>
  </entry>
  <entry>
    <title>A Reconfigurable Tracked Robot for Enhanced Obstacle Traversal Through Movable Articulation Point and Internal Mass Relocation</title>
    <id>arxiv:2608.07624</id>
    <link href="https://arxiv.org/abs/2608.07624" />
    <published>2026-08-07T09:50:51.000Z</published>
    <updated>2026-08-07T09:50:51.000Z</updated>
    <author><name>Yuki Uda</name></author>
    <author><name>Yasutaka Nakashima</name></author>
    <author><name>Motoji Yamamoto</name></author>
    <author><name>Ayato Kanada</name></author>
    <category term="cs.RO" />
    <summary>Tracked robots are widely used in unstructured environments; however, their obstacle traversal capability is fundamentally limited by a tradeoff between front-end reachability and locomotion stability. This study presents TRASER (Tracked Robot with Articulated Spine for Extended Reach), a reconfigurable tracked robot capable of relocating both its articulation point and internal mass. TRASER employs a tape-spring mechanism that localizes compliance to the bending region while maintaining high stiffness in the remaining body, thereby improving both front-end reachability and center-of-mass (Co…</summary>
  </entry>
  <entry>
    <title>Real-time Whole-Body Motion Planning for Mobile Manipulators Carrying Arbitrarily Shaped Payloads via Kinematically-Coupled SVSDF</title>
    <id>arxiv:2608.07005</id>
    <link href="https://arxiv.org/abs/2608.07005" />
    <published>2026-08-07T09:20:31.000Z</published>
    <updated>2026-08-07T09:20:31.000Z</updated>
    <author><name>Yisheng Li</name></author>
    <author><name>Longji Yin</name></author>
    <author><name>Tingrui Zhang</name></author>
    <author><name>Ruize Xue</name></author>
    <author><name>Haoda Zhu</name></author>
    <author><name>Nan Chen</name></author>
    <category term="cs.RO" />
    <summary>Mobile manipulators are increasingly tasked with transporting large, non-convex payloads through cluttered environments, yet existing planners either oversimplify the payload geometry or fail to handle the kinematic coupling between manipulator links, leading to lost feasible space or stalled optimization. This letter presents a real-time whole-body motion planning framework for mobile manipulators carrying arbitrarily shaped payloads. The front-end employs a chain-decomposed kernel-based collision check that preserves the true geometry of the robot and payload, with compact storage and fast…</summary>
  </entry>
  <entry>
    <title>A Haptic Robot Finger Designed for Guqin Instrument Playing</title>
    <id>arxiv:2608.07002</id>
    <link href="https://arxiv.org/abs/2608.07002" />
    <published>2026-08-07T09:18:58.000Z</published>
    <updated>2026-08-07T09:18:58.000Z</updated>
    <author><name>Tianwei Zhang</name></author>
    <author><name>Hanming Yan</name></author>
    <author><name>Yang Yang. Ziya Wang</name></author>
    <category term="cs.RO" />
    <summary>With the rapid advancement of humanoid robotics and embodied intelligence technologies, numerous musical instrument-playing robots have emerged in recent years, such as pianos, chime bells, and taiko drums. These robots primarily employ open-loop positional control, rendering them incapable of operating instruments requiring dexterous hands and precise tactile perception, such as a violin, guitar, and guqin. This paper describes the design and validation of a high-precision tactile-sensing finger. By mimicking the shape of the fingertip and fingernail found on a human finger, we develop a bio…</summary>
  </entry>
  <entry>
    <title>Decoupling Intention from Trajectory: A Representational Deduction Framework for World Action Models</title>
    <id>arxiv:2608.06994</id>
    <link href="https://arxiv.org/abs/2608.06994" />
    <published>2026-08-07T09:12:15.000Z</published>
    <updated>2026-08-07T09:12:15.000Z</updated>
    <author><name>Xiangkai Ma</name></author>
    <author><name>Yue Ma</name></author>
    <author><name>Junjie Wang</name></author>
    <author><name>Sheng Xu</name></author>
    <author><name>Mingyang Li</name></author>
    <author><name>Han Zhang</name></author>
    <category term="cs.RO" />
    <summary>World Action Models (WAMs) aim to construct a unified architecture capable of understanding world state evolution and guiding to generative motion planning. However, existing visual branches focus on predicting static visual observation, rather than reflecting potential transition information that captures the evolution of world states under motion interactions. This leads to representational entanglement between high-level physical condition evolution and low-level action trajectory generation within the Action Model, creating a structural bottleneck while weakening the predictive capability…</summary>
  </entry>
  <entry>
    <title>CMU-Drive and V2V-VLA: Cooperative Multi-agent Unified Driving with Reasoning Benchmark and Vehicle-to-Vehicle Vision-Language-Action Models</title>
    <id>arxiv:2608.07621</id>
    <link href="https://arxiv.org/abs/2608.07621" />
    <published>2026-08-07T09:00:39.000Z</published>
    <updated>2026-08-07T09:00:39.000Z</updated>
    <author><name>Hsu-kuang Chiu</name></author>
    <author><name>Stephen F. Smith</name></author>
    <category term="cs.AI" />
    <summary>Vision-Language-Action (VLA) models have recently achieved impressive performance for end-to-end autonomous driving, yet existing approaches are primarily designed for an individual single autonomous driving agent with limited support for cooperative perception, reasoning, and planning. We present Cooperative Multi-agent Unified Driving with Reasoning (CMU-Drive), a closed-loop end-to-end benchmark for evaluating cooperative autonomous driving with multiple connected autonomous vehicles (CAVs) operating in safety-critical driving scenarios with background traffic participants. We further prop…</summary>
  </entry>
  <entry>
    <title>Cross-View Action Consistency for Camera-Robust Vision-Language-Action Policies</title>
    <id>arxiv:2608.06965</id>
    <link href="https://arxiv.org/abs/2608.06965" />
    <published>2026-08-07T08:41:31.000Z</published>
    <updated>2026-08-07T08:41:31.000Z</updated>
    <author><name>Bingqi Huang</name></author>
    <author><name>Bingchuan Wei</name></author>
    <author><name>Xuan Wang</name></author>
    <author><name>Yingkai Cai</name></author>
    <author><name>Zhaokui Wang</name></author>
    <category term="cs.RO" />
    <summary>Vision-language-action (VLA) policies fine-tuned from a fixed scene camera can fail when the camera is moved, even when the task, objects, language, and robot state are unchanged. We study scene-camera viewpoint robustness using only a scene RGB image, language, and proprioception, without camera labels, extrinsics, depth, or point-cloud inputs. The wrist stream is masked throughout to prevent an unperturbed visual shortcut from confounding attribution to scene-camera variation. For flow-based VLAs, we propose to regularize the action-flow velocity field, the quantity directly integrated to g…</summary>
  </entry>
  <entry>
    <title>Spatiotemporal Agility: Time-Constrained Reinforcement Learning for Vision-Guided Dynamic Quadrupedal Interception</title>
    <id>arxiv:2608.06907</id>
    <link href="https://arxiv.org/abs/2608.06907" />
    <published>2026-08-07T07:42:19.000Z</published>
    <updated>2026-08-07T07:42:19.000Z</updated>
    <author><name>Yidong Zhu</name></author>
    <author><name>Zibo Dai</name></author>
    <author><name>Tongning Zhang</name></author>
    <author><name>Leixin Chang</name></author>
    <author><name>Hua Chen</name></author>
    <category term="cs.RO" />
    <summary>Legged robots require robust agility to perceive and interact with complex and dynamic environments within a constrained time. However, most existing quadruped locomotion works rely on velocity-tracking policy, which struggle to reach precise targets within strict temporal constraints. Moreover, integrating real-time perception with agile locomotion for highly dynamic targets remains challenging due to sensor latency and processing delays. To concretely study and benchmark such agility in dynamic settings, we introduce a challenging ball-catching task for legged robots. This paper proposes an…</summary>
  </entry>
  <entry>
    <title>GWM-VLA: Geometry-Aware Latent World Modeling for Vision-Language-Action Learning</title>
    <id>arxiv:2608.07619</id>
    <link href="https://arxiv.org/abs/2608.07619" />
    <published>2026-08-07T06:41:51.000Z</published>
    <updated>2026-08-07T06:41:51.000Z</updated>
    <author><name>Yanping Zhao</name></author>
    <author><name>Hang Yu</name></author>
    <author><name>Yiwei Wang</name></author>
    <author><name>Chen Ye</name></author>
    <author><name>Siyu Tian</name></author>
    <author><name>Di Zhang</name></author>
    <category term="cs.RO" />
    <summary>Vision-Language-Action (VLA) models achieve strong robotic manipulation performance but often degrade under visual and environmental shifts. Latent world modeling offers a promising approach to improving robustness, yet existing methods commonly encode camera views independently and predict holistic scene dynamics without explicitly modeling their geometric relationships. We propose GWM-VLA, a geometry-aware latent world modeling framework for VLA learning. GWM-VLA combines geometry-aware multi-view state encoding, global context-conditioned target-view prediction, and shared latent-action re…</summary>
  </entry>
  <entry>
    <title>AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models</title>
    <id>arxiv:2608.06729</id>
    <link href="https://arxiv.org/abs/2608.06729" />
    <published>2026-08-07T02:49:39.000Z</published>
    <updated>2026-08-07T02:49:39.000Z</updated>
    <author><name>Guiyu Zhao</name></author>
    <author><name>Longteng Guo</name></author>
    <author><name>Yanghong Mei</name></author>
    <author><name>Zilin Zhu</name></author>
    <author><name>Yu Zhang</name></author>
    <author><name>Bin Cao</name></author>
    <category term="cs.RO" />
    <summary>While Vision-Language-Action (VLA) models have advanced embodied AI, their fundamentally reactive paradigm severely limits performance in partially observable and long-horizon tasks. When restricted to a single wrist-mounted camera, they inevitably suffer from perception forgetting as objects exit the field of view, and temporal task-progress forgetting} during multi-step execution. To overcome these bottlenecks, we propose AtlasVLA, a novel framework that transitions from direct reactive manipulation to proactive reasoning through a persistent world-ego state. AtlasVLA features a dual-memory…</summary>
  </entry>
  <entry>
    <title>CrossTracer: Cross-Embodiment Navigation via VLA Model Reasoning and Trace Residuals Adapting</title>
    <id>arxiv:2608.06688</id>
    <link href="https://arxiv.org/abs/2608.06688" />
    <published>2026-08-07T01:24:53.000Z</published>
    <updated>2026-08-07T01:24:53.000Z</updated>
    <author><name>Yao Wang</name></author>
    <author><name>Siyuan Wang</name></author>
    <author><name>Zhirui Sun</name></author>
    <author><name>Wenzheng Chi</name></author>
    <author><name>Liang Lin</name></author>
    <author><name>Jiankun Wang</name></author>
    <category term="cs.RO" />
    <summary>Vision-language-action (VLA) models provide strong semantic priors for robot navigation, but they often ignore embodiment-specific mobility constraints. A path that is semantically plausible for one robot may be physically infeasible for another. We propose CrossTracer, a hierarchical framework for cross-embodiment navigation through adaptive trace residuals. CrossTracer represents navigation plans as normalized image-plane waypoints, forming a unified pixel-space interface between semantic reasoning and physical grounding. First, Vision-Language Trace Proposer (VL-Tracer) adapts a pretrained…</summary>
  </entry>
</feed>
