UNITAS / WORLD ACTION MODELING

UNITASA 3D-Native World Action Model for Embodied Manipulation

Ruixiang Wang1,2*Yongyi Su3*†Wenlve Zhou3Bo Yue1,2Hengyan Liu1,2Dekun Lu2
Yuxin Tian1,2Yihan Fang1,2Zerui Wu2Xing Hu2Jietao Chen2Yong Guo2
Ziyan He2Junbin Yuan4Guiliang Liu1Xiaofen Xing5Kui Jia1,2
1 The Chinese University of Hong Kong, Shenzhen2 DexForce Technology Co., Ltd.3 Foshan University4 Sun Yat-sen University5 South China University of Technology
* Equal contribution† Corresponding author

A 3D-native world action model.

A future you can look around.

Drag to orbit. Scroll to zoom.
Play the prediction, or move through time.

Input

Initial observation RGB at source resolution
Language instruction

Choose an interaction above.

Initial query points Drag to orbit ↔

Output

Loading scene

↔ Drag to explore
0 / 20

Loading 3D trajectories…

Many interactions.
A world in motion.

Lifting. Inserting. Opening. Releasing.
Watch predicted scene responses across simulation and real-world interactions.

A shared representation.
Two complementary capabilities.

POLICY MODE

From instruction to action.

Generate task-conditioned gripper motion and executable robot commands. Direct control can skip flow generation when only actions are needed.

WORLD-MODEL MODE

From motion to scene evolution.

Condition on a gripper trajectory and predict how queried scene points move over the same physical interval. On this page, the conditioning trajectory is generated by the policy from the same observation and instruction.

UNITAS architecture with shared observation grounding and coupled motion and action branches

Metric groundingWorld-aligned 3D positional embeddings ground visual tokens in a shared physical frame.

Physical timeA trajectory tokenizer encodes each point trajectory as one token, anchored at its current 3D position.

Coupled expertsA point dynamics expert models action and scene flow, while an action expert generates robot commands.

Strong performance across benchmarks.

99.8%

LIBERO

Average success
89.1%

LIBERO-Plus

Overall success
93.04%

RoboTwin Randomized

Success rate
56.2%

VLABench

Success rate
Policy benchmark comparison Β· higher is better
MethodParams (B)LIBERO Avg.LIBERO-PlusRoboTwin CleanRoboTwin Rand.VLABench SRVLABench PSVLABench IS
VLA Policies
X-VLA1.098.171.472.8872.8429.451.270.2
VLA-JEPA2.097.277.9β€”β€”β€”β€”β€”
Ο€β‚€3.094.1553.665.9258.4029.444.155.0
Ο€β‚€.β‚…3.096.985.782.7476.7648.162.364.9
ABot-M04.098.680.580.4281.16β€”β€”β€”
Video-Based WAMs
Cosmos Policy2.098.582.2β€”β€”β€”β€”β€”
GE-Act2.296.580.3β€”β€”β€”β€”β€”
LingBot-VA5.398.569.592.9391.55β€”β€”β€”
Fast-WAM6.097.651.591.8891.78β€”β€”β€”
Motus8.097.7β€”88.6687.02β€”β€”β€”
3D-Aware Policies
SpatialVLA3.578.1β€”β€”β€”β€”β€”β€”
LaMP5.098.379.3β€”β€”β€”β€”β€”
X-WAM6.7β€”β€”89.890.7β€”β€”β€”
Dex-BEVβ€”97.8β€”76.042.0β€”β€”β€”
UNITAS
UNITAS (w/o Pretrain)1.799.586.591.0890.7450.364.371.2
UNITAS1.799.889.192.9493.0456.267.669.3