One geometric language
Visual observations, numerical state, and future actions share the same geometry and timestamps.
One skeleton. Many embodiments.
LIBERO-Cross10 · 1,000 episodes per method · 95% Wilson interval for SkelWAM: 40.3–46.4%. Table I
Overview
Reusing manipulation experience across embodiments changes both the visual observations and the action geometry: a new body alters appearance, control coordinates, and the whole-body configurations that can realize the same tool pose. SkelWAM is a skeleton-guided world–action model that connects perception and control through one explicit geometric representation. Six centerline samples, tool-center-point (TCP) pose, and a parallel-jaw command define a shared 25-D state. The same geometry defines canonical third-person and wrist observations and future whole-body action targets. Trained with predictive visual supervision, paired video–action experts predict skeleton action chunks; embodiment-specific constrained decoders realize the predicted intent through joint or continuum-robot controls. Task learning and model selection use source demonstrations only, with no one-to-one joint correspondence or target-task policy updates.
Visual observations, numerical state, and future actions share the same geometry and timestamps.
Centerline posture guidance complements task-essential TCP pose and jaw commands.
Target-specific kinematics and calibration change. The learned policy stays fixed.
Method
Canonicalize the observation.
Predict a skeleton chunk.
Decode to the target body.
Six samples follow centerline arc length—not native joint locations. The task-frame TCP preserves the interaction goal.
B = {(pi − p5) / Lr}i=0…4
Five non-TCP centerline points, expressed relative to the TCP and normalized by sampled centerline length. The full chain contains six points; p₅ is the TCP.
Two distinct normalizations. Body offsets use sampled centerline length; TCP position uses the task-frame mapping νF. Source-fitted per-coordinate normalization is applied afterward for policy learning and remains fixed on targets. Sec. III-B
For simulation source training, adapt the final four Video Expert blocks, the video output head, and the Action Expert for 6k updates with video prediction and 25-D skeleton-action objectives.
Freeze the Video Expert and train the action pathway on current-frame features for 24k updates. Predict 24-D continuous geometry plus jaw-switch presence and first-switch timing.
Cache current-frame video K/V, use 20 flow-matching steps to predict 32 absolute states, execute ten states, then replan. Each chunk contains at most one jaw-state switch.
Simulation checkpoint selection uses held-out Franka episodes: Stage-B step 23k after 6k Stage-A updates. Sec. IV-B
In simulation, body offsets are rescaled using the target’s sampled centerline length at decoder initialization, then translated by the task-frame TCP position. Joint IK or PCC decoding prioritizes TCP and jaw commands while using the centerline as soft posture guidance; exact centerline agreement is not required.
With no prior intent, the simulation bridge starts from measured geometry. Rigid-link targets retain body intent and update measured TCP position, orientation, and jaw command. The continuum target retains body and interaction-orientation intent while updating measured TCP position and jaw feedback. Decoders still use measured native configuration for kinematics and feasibility.
LIBERO-Cross10
Ten tasks. Four morphology groups.
The objects, instructions and success
predicates remain from LIBERO.
Cross-Object, task 3
Native demonstration
and its canonical views
467 source demonstrations447 train / 20 validation
1,000 episodes per method10 targets × 10 tasks × 10 trials
0 target-task policy updatesFranka excluded from target averages
Simulation results
These are comparisons of deployed systems, including their adapters, under the same target protocol; pretraining and source-adaptation budgets differ. Non-SkelWAM methods keep their method-specific image processing and native action contracts with target-specific EEF-IK. They receive no skeleton images, centerline state, or body-tracking objective. Mirage reuses FastWAM weights; RoVi-Aug uses source-data augmentation and a 4k adapted checkpoint. Sec. IV-B · Table I
Bold: highest; underline: second-highest distinct rate in each column, including ties. Mean equally weights all ten targets.
Download Table I as CSVComponent analysis
The source-retrained native-RGB visual variant and the variant without body-action supervision achieve 0.0% and 0.1% success. Wrist input and the decoder’s body objective have task-dependent effects, rather than uniform gains across task families.
Table II · intervention definitions| Component intervention | Success | Δ (pp) |
|---|---|---|
| Full SkelWAM | 43.3% | — |
| Without skeleton vision | 0.0% | −43.3 |
| Without body-action supervision* | 0.1% | −43.2 |
| Without wrist input | 43.0% | −0.3 |
| Without decoder body objective | 40.9% | −2.4 |
* The body-action variant keeps the 25-D architecture, removes supervision on its 15 body coordinates, and also uses an EEF-only decoder objective. The V, A and W variants are source-retrained; only the decoder-body intervention reuses the full-model checkpoint at inference.
Real-robot deployment
JAKA mini2 demonstrations
Feagine A03 deployment
Three tabletop tasks · qualitative study
For the physical study, 277 quality-controlled JAKA mini2 teleoperation demonstrations are split into 249 training and 28 validation trajectories and compiled into 10 Hz skeleton windows. The Action Expert and final two video blocks are trained on JAKA mini2 data for 6k updates with the action objective, with source-only model selection. The JAKA-trained policy is then deployed on Feagine A03 with fixed weights and a target-specific PCC decoder; no Feagine task demonstrations are used for fine-tuning. Sec. IV-E
Move the toy into the container.
Place the ring over the post.
Align and insert the cylindrical peg.

Figure 7 shows four selected times on each platform. For Toy Cleanup, inspect external RGB, mask overlays, skeleton composites, wrist RGB, and wrist illustrations; lower bands show Peg Insertion and Ring Placement.
Evidence scope. The manuscript presents qualitative JAKA-to-Feagine deployments, not a physical success-rate table. Figure 6 uses arrows and translucent poses to illustrate task motions; Figure 7 presents the selected task images and associated external and wrist-view processing. Fig. 6 · Fig. 7 / Sec. IV-E
Calibrated JAKA forward kinematics and Feagine PCC models encode native feedback into the shared representation. Task-frame, tool and camera calibration register the numerical skeleton with both camera views; the TCP is the gripper contact center, accounting for different tool lengths. Masking and skeleton-point tracking combine static-ROI feature matching with RANSAC, forward–backward Lucas–Kanade flow, and color-assisted TCP relocking.
Physical implementation · Sec. IV-E →Scope & limitations
Target geometry, camera and TCP calibration, native coordinate limits, and kinematic control remain necessary. Exact centerline matching is not required: the decoder prioritizes the interaction goal while using the body as soft posture guidance.
Performance is lower on Cross-Goal and Cross-Long. Reachability, gripper contact geometry, calibration, and background filling under occlusion remain important limits. Baseline training budgets and coupled interventions also limit component-level attribution.
Read the discussionCitation
@misc{skelwam,
title = {SkelWAM: A Skeleton-Guided World-Action Model for Zero-Shot Cross-Embodiment Manipulation},
author = {Niu, Pengjun and Xie, Yujia and Peng, Rui and Zhao, Hang and Liu, Ke},
note = {Research manuscript},
url = {http://www.liukepku.com/skelwam/index.html}
}Citation for the current manuscript. Publication venue and year are not specified in this version. Download .bib