Research manuscript

A Skeleton-Guided World-Action Model for
Zero-Shot Cross-Embodiment Manipulation

One skeleton. Many embodiments.

Franka demonstrations train the simulation policy; JAKA mini2 data train the physical policy deployed on Feagine A03. Fig. 1
25Dshared skeleton state
10 × 10tasks × target robots
43.3%zero-shot simulation success
+36.2 ppover the best evaluated baseline

LIBERO-Cross10 · 1,000 episodes per method · 95% Wilson interval for SkelWAM: 40.3–46.4%. Table I

Overview

Change the robot.
Keep the representation.

Reusing manipulation experience across embodiments changes both the visual observations and the action geometry: a new body alters appearance, control coordinates, and the whole-body configurations that can realize the same tool pose. SkelWAM is a skeleton-guided world–action model that connects perception and control through one explicit geometric representation. Six centerline samples, tool-center-point (TCP) pose, and a parallel-jaw command define a shared 25-D state. The same geometry defines canonical third-person and wrist observations and future whole-body action targets. Trained with predictive visual supervision, paired video–action experts predict skeleton action chunks; embodiment-specific constrained decoders realize the predicted intent through joint or continuum-robot controls. Task learning and model selection use source demonstrations only, with no one-to-one joint correspondence or target-task policy updates.

One geometric language

Visual observations, numerical state, and future actions share the same geometry and timestamps.

Whole-body intent

Centerline posture guidance complements task-essential TCP pose and jaw commands.

Frozen target policy

Target-specific kinematics and calibration change. The learned policy stays fixed.

Method

A shared guiding skeleton
for vision and action.

Canonicalize the observation.
Predict a skeleton chunk.
Decode to the target body.

Future-video paths are auxiliary supervision in Stage A only. No future video is generated at deployment. Fig. 2 / Sec. III-D
The canonical state

15 + 3 + 6 + 1 = 25

Six samples follow centerline arc length—not native joint locations. The task-frame TCP preserves the interaction goal.

B = {(pi − p5) / Lr}i=0…4

Five non-TCP centerline points, expressed relative to the TCP and normalized by sampled centerline length. The full chain contains six points; p₅ is the TCP.

Two distinct normalizations. Body offsets use sampled centerline length; TCP position uses the task-frame mapping νF. Source-fitted per-coordinate normalization is applied afterward for policy learning and remains fixed on targets. Sec. III-B

Stage A · 6k updates

Learn visual dynamics
and skeleton actions

For simulation source training, adapt the final four Video Expert blocks, the video output head, and the Action Expert for 6k updates with video prediction and 25-D skeleton-action objectives.

LA = λvLvideo + Laction25D
Stage B · 24k updates

Match the deployment
conditioning path

Freeze the Video Expert and train the action pathway on current-frame features for 24k updates. Predict 24-D continuous geometry plus jaw-switch presence and first-switch timing.

LB = Lflow24D + λg(Lpresence + Ltiming)
At inference

Observe once.
Predict. Execute. Repeat.

Cache current-frame video K/V, use 20 flow-matching steps to predict 32 absolute states, execute ten states, then replan. Each chunk contains at most one jaw-state switch.

Current-frame K/V32 × 25Joint IK / PCC

Simulation checkpoint selection uses held-out Franka episodes: Stage-B step 23k after 6k Stage-A updates. Sec. IV-B

How native decoding and closed-loop feedback workSec. III-E +

Decode at the target scale

In simulation, body offsets are rescaled using the target’s sampled centerline length at decoder initialization, then translated by the task-frame TCP position. Joint IK or PCC decoding prioritizes TCP and jaw commands while using the centerline as soft posture guidance; exact centerline agreement is not required.

Retain intent, update interaction feedback

With no prior intent, the simulation bridge starts from measured geometry. Rigid-link targets retain body intent and update measured TCP position, orientation, and jaw command. The continuum target retains body and interaction-orientation intent while updating measured TCP position and jaw feedback. Decoders still use measured native configuration for kinematics and feasibility.

Read the decoding and feedback definition →

LIBERO-Cross10

One source.
Ten unseen robots.

Ten tasks. Four morphology groups.
The objects, instructions and success
predicates remain from LIBERO.

3 Cross-Spatial3 Cross-Object2 Cross-Goal2 Cross-Long

467 source demonstrations447 train / 20 validation

1,000 episodes per method10 targets × 10 tasks × 10 trials

0 target-task policy updatesFranka excluded from target averages

Simulation results

Shared geometry.
Measurable transfer.

43.3%SkelWAM mean success
7.1% best evaluated baseline

Zero-shot success by method

0255075100%

These are comparisons of deployed systems, including their adapters, under the same target protocol; pretraining and source-adaptation budgets differ. Non-SkelWAM methods keep their method-specific image processing and native action contracts with target-specific EEF-IK. They receive no skeleton images, centerline state, or body-tracking objective. Mirage reuses FastWAM weights; RoVi-Aug uses source-data augmentation and a 4k adapted checkpoint. Sec. IV-B · Table I

Full per-robot comparison8 methods · 10 targets +
LIBERO-Cross10 zero-shot success (%). Each target contributes 100 episodes.

Bold: highest; underline: second-highest distinct rate in each column, including ties. Mean equally weights all ten targets.

Download Table I as CSV

Component analysis

Both sides of
the skeleton matter.

The source-retrained native-RGB visual variant and the variant without body-action supervision achieve 0.0% and 0.1% success. Wrist input and the decoder’s body objective have task-dependent effects, rather than uniform gains across task families.

Table II · intervention definitions
Component interventionSuccessΔ (pp)
Full SkelWAM43.3%
Without skeleton vision0.0%−43.3
Without body-action supervision*0.1%−43.2
Without wrist input43.0%−0.3
Without decoder body objective40.9%−2.4

* The body-action variant keeps the 25-D architecture, removes supervision on its 15 body coordinates, and also uses an EEF-only decoder objective. The V, A and W variants are source-retrained; only the decoder-body intervention reuses the full-model checkpoint at inference.

Real-robot deployment

From a rigid arm
to a continuum arm.

JAKA mini2 demonstrations
Feagine A03 deployment
Three tabletop tasks · qualitative study

For the physical study, 277 quality-controlled JAKA mini2 teleoperation demonstrations are split into 249 training and 28 validation trajectories and compiled into 10 Hz skeleton windows. The Action Expert and final two video blocks are trained on JAKA mini2 data for 6k updates with the action objective, with source-only model selection. The JAKA-trained policy is then deployed on Feagine A03 with fixed weights and a target-specific PCC decoder; no Feagine task demonstrations are used for fine-tuning. Sec. IV-E

Task setups · Fig. 6
01

Toy Cleanup

Move the toy into the container.

02

Ring Placement

Place the ring over the post.

03

Peg Insertion

Align and insert the cylindrical peg.

Figure 7 source and target physical visual–action interface montage
Behind the observation

The same geometry, in both views.

Figure 7 shows four selected times on each platform. For Toy Cleanup, inspect external RGB, mask overlays, skeleton composites, wrist RGB, and wrist illustrations; lower bands show Peg Insertion and Ring Placement.

Evidence scope. The manuscript presents qualitative JAKA-to-Feagine deployments, not a physical success-rate table. Figure 6 uses arrows and translucent poses to illustrate task motions; Figure 7 presents the selected task images and associated external and wrist-view processing. Fig. 6 · Fig. 7 / Sec. IV-E

Calibration and physical visual processingJAKA → Feagine +

Calibrated JAKA forward kinematics and Feagine PCC models encode native feedback into the shared representation. Task-frame, tool and camera calibration register the numerical skeleton with both camera views; the TCP is the gripper contact center, accounting for different tool lengths. Masking and skeleton-point tracking combine static-ROI feature matching with RANSAC, forward–backward Lucas–Kanade flow, and color-assisted TCP relocking.

Physical implementation · Sec. IV-E →

Scope & limitations

Zero-shot does not mean
calibration-free.

Target geometry, camera and TCP calibration, native coordinate limits, and kinematic control remain necessary. Exact centerline matching is not required: the decoder prioritizes the interaction goal while using the body as soft posture guidance.

Performance is lower on Cross-Goal and Cross-Long. Reachability, gripper contact geometry, calibration, and background filling under occlusion remain important limits. Baseline training budgets and coupled interventions also limit component-level attribution.

Read the discussion

Citation

BibTeX

@misc{skelwam,
  title  = {SkelWAM: A Skeleton-Guided World-Action Model for Zero-Shot Cross-Embodiment Manipulation},
  author = {Niu, Pengjun and Xie, Yujia and Peng, Rui and Zhao, Hang and Liu, Ke},
  note   = {Research manuscript},
  url    = {http://www.liukepku.com/skelwam/index.html}
}

Citation for the current manuscript. Publication venue and year are not specified in this version. Download .bib