Robotics & embodied AI

Robot learning, manipulation, navigation, vision-language-action models, embodied agents in the physical world

41 articles

Robotics & embodied AI

Replay-free continual learning method prevents forgetting in vision-language-action models

The authors introduce SAMBAR, a continual learning algorithm for vision-language-action (VLA) models that requires no replay of previously seen demonstrations. SAMBAR frames continual learning as a constrained optimization problem, solved via the method of multipliers, and selectively anchors parameters important to earlier tasks. On the LIBERO benchmark and in hardware experiments, SAMBAR retained every learned task, while all replay-free baselines fully forgot the first task they were taught.

3 Oct 2026
Robotics & embodied AI

A single world model enables robots to insert unseen parts zero-shot

Hansen et al. present InsertionWM, a framework that trains a single world model on up to 90 geometrically diverse insertion tasks using wrist-mounted camera images and robot proprioception. The model achieves 56% zero-shot success on unseen objects with unknown geometry, versus 7% for a model-free baseline. Finetuning the generalist model on held-out objects improves data-efficiency and asymptotic performance.

3 Oct 2026
Robotics & embodied AI

Canonical gripper-frame representation achieves zero-shot cross-embodiment transfer for two-finger manipulation

The authors propose an interaction-centric framework that uses a parameterized universal gripper abstraction to transform RGB-D observations into a canonical gripper frame. This enables zero-shot cross-embodiment and cross-viewpoint transfer for two-finger gripper manipulation tasks, demonstrated in both simulation and real-world experiments. The approach simultaneously achieves competitive benchmark performance and extreme generalization to heterogeneous robot platforms.

3 Oct 2026
Robotics & embodied AI

Trajectory optimization that preserves endpoints improves closed-loop driving

Zhang et al. introduce Endpoint-Constrained Optimization (ECO), a training-free post-processor for end-to-end driving models that fixes the vehicle's executed history and the policy's predicted endpoint, then optimises the intermediate waypoints for physical plausibility. Across two closed-loop simulators and five datasets, ECO improved the aggregate score of every evaluated policy, with gains up to 71% in HUGSIM score for the VaVAM model.

3 Oct 2026
Robotics & embodied AI

Aiming at observed intermediate targets beats final-goal scoring in frozen world model planning

The authors show that scoring predicted outcomes by their distance to the final goal can break latent world model planning even when dynamics are exact and short-horizon search is globally optimal, whenever the route to the goal initially moves away. They introduce Anchored Planning, a training-free method that reuses the frozen model's own offline trajectories to select an intermediate target. Across Cube, PushT, Reacher, and TwoRoom, it outperforms the LeWM planner and additional final-goal search on long-range goals.

3 Oct 2026
Robotics & embodied AI

Semantic messaging helps distributed robots coordinate long manipulation tasks

Zhou and colleagues present DuoMind, a hierarchical system in which each robot independently plans, communicates with peers and executes detailed actions. Across the new RoboPoly benchmark and the existing RoboTwin benchmark, the authors report that this combination improves coordinated task performance, while ablations indicate that both semantic communication and hierarchical control contribute to the gains.

2 Oct 2026
Robotics & embodied AI

Robots improve reusable manipulation skills through guided simulation practice

Wang and colleagues present Reconstruct, Practice, Go Real, a framework that improves a robot execution system through simulated practice without changing the underlying model weights. Across 22 manipulation tasks, success on held-out initializations rose from 28.6% after one practice round to 95.0% after 15 rounds, while a frozen version completed 30 of 30 reported physical trials across three tasks.

2 Oct 2026
Robotics & embodied AI

Simulation-first framework lets coding agents learn real robot skills in 10 minutes

The authors introduce SimEX, a two-stage autoresearch framework that tightly integrates simulation with LLM-based coding agents. In the first stage, the agent conducts open-ended probe-and-optimize iterations in simulation to develop a robust robot toolbox; in the second stage, the toolbox is adapted to a real robot through only a few physical trials, each correcting the simulator and using the corrected simulator to diagnose failures and screen repairs. On three challenging real-world manipulation tasks—towel folding, barcode scanning, and plate manipulation—SimEX achieved success without any demonstrations and with only 10 minutes of real-robot interaction.

1 Oct 2026
Robotics & embodied AI

Touch-directed curiosity helps robots discover useful manipulation skills

Iten and colleagues introduce TacEx, an exploration framework that rewards robots for reducing uncertainty specifically about tactile outcomes rather than about every unpredictable transition. The authors report that this focus drives contact-rich exploration without task rewards or expert demonstrations, yielding data that supports offline pick-and-place learning and sample-efficient post-training of vision-language-action models.

1 Oct 2026
Robotics & embodied AI

Robot self-improvement stalls when perception, skill chains and tests mislead

Wang built an agentic robotics system that diagnoses failures, writes skills or installs external models, and tests changes without human-written robot code. The agent successfully discovered some missing capabilities, but its improvements did not accumulate into task success because relational perception, sequential skill testing and the evaluation harness constrained what it could learn.

30 Sept 2026
Robotics & embodied AI

Robot-relative 3D features improve pretrained manipulation policies across diverse tasks

Liu et al. introduce Spatial Grafting, which converts features from frozen 3D reconstruction models into robot-relative spatial tokens and feeds them to flow-matching action generators through cross-attention. Across four simulation benchmarks and three physical robot platforms, the authors report consistent gains, with the largest improvements appearing in cluttered and long-horizon manipulation.

30 Sept 2026
Robotics & embodied AI

Coding agents synthesize robot planners that generalize to unseen instances

The authors gave coding agents task descriptions and simulator access, then asked them to develop reusable programs for 28 simulated task-and-motion planning environments. Across 98,000 held-out evaluation episodes, the resulting frozen programs achieved mean success rates of 56% to 95%; on the 16 environments with planner baselines, all three agent configurations exceeded the planners’ 47% mean success.

27 Sept 2026
Robotics & embodied AI

Synthetic demonstrations let robot policies escape sparse-reward failures

The authors introduce SynthDemo-RL, a teacher-student pipeline that generates demonstrations from simulator-privileged state, uses them to fine-tune a vision-language-action policy, and then refines it with reinforcement learning. On LIBERO-PRO, the method achieved nonzero success on all 27 tasks where the starting policy had failed completely, while direct PPO rescued only 10 under matched reinforcement-learning compute.

22 Sept 2026
Robotics & embodied AI

Adaptive action chunking improves robot control by varying horizon based on prediction reliability.

The authors propose GeoAAC, a method for Vision-Language-Action (VLA) policies that adaptively sets the action horizon—the number of future actions predicted at once—based on the geometric properties of the denoising process. By measuring uncertainty through the variation of denoising trajectories across action prefixes, GeoAAC selects a horizon from a single generation without extra training, achieving consistent improvements over fixed horizons in both simulation and real-world tasks.

21 Sept 2026
Robotics & embodied AI

A semantic harness lets vision-language models control different robots

Chen et al. introduce an interface that lets vision-language models make fine-grained physical decisions using semantic action units, which deterministic interpreters translate into commands for particular robots. The authors report that this approach transfers across tasks, environments, and robot embodiments while outperforming representative agent-based and vision-language-action baselines.

11 Sept 2026