Rethinking World-Action Model for
Compositional and In-Context Robotic Manipulation

ViGAR — Visual Goal-conditioned Action Reasoning

1Peking University 2AgiBot 3CocoMatrix
*Equal Contribution    †Corresponding Author
Peking University AgiBot

Overview video

ViGAR overview: a subgoal planner predicts the next subgoal image from the observation and instruction (optionally a global goal), and a world-action policy generates the video and action chunk that reach it.

Overview of ViGAR. Long-horizon compositional manipulation is factorized into visual subgoal planning and subgoal execution, with both modules trained on task-diverse, long-horizon robot data. ViGAR performs strongly in simulation and on a real AgiBot A2, and generalizes to unseen task compositions through in-context steering with a global goal image.

RoboTwin 2.0 Leaderboard

Simulation · RoboTwin Clean2Random

average success rate (%) over the Clean and Random settings · 50 tasks × 100 rollouts each

Real robot · AgiBot A2

average task progress over 5 long-horizon tasks × 20 rollouts

About ViGAR

ViGAR architecture: the subgoal planner generates the next subgoal image from the instruction and current observation (optionally a global goal image); the policy takes the predicted subgoal and jointly predicts future frames and an action chunk.

Model architecture. (1) The subgoal planner Gθ generates the next subgoal image from the task instruction and current observation, optionally conditioned on a global goal image for in-context learning. (2) The policy πWAMφ takes the predicted subgoal and jointly predicts the future frames and the action chunk. Both share the Cosmos3-Nano reasoner–generator backbone.

Long-horizon compositional manipulation has become increasingly important for real-world robot deployment, where a single task involves multiple coordinated subtasks. Existing world-action models (WAMs) jointly predict short-horizon visual futures and actions, but typically lack explicit subtask-level reasoning. We propose Visual Goal-conditioned Action Reasoning (ViGAR), a hierarchical framework that factorizes manipulation into a visual subgoal planner and a subgoal executor. Given the current observation and global instruction, the subgoal planner predicts a visual subgoal for the next subtask. The subgoal executor then jointly generates future visual trajectories and actions conditioned on the predicted subgoal. Both components share a pretrained world-model representation, enabling task-level planning and action generation to benefit from common physical knowledge.

Moreover, our framework naturally supports in-context learning: using a global goal image as context can induce different subtask decompositions and behaviors without parameter updates. On the RoboTwin Clean2Random benchmark, ViGAR achieves 82.00% and 67.02% success rates under the Clean and Random settings, surpassing the strongest baseline by 12.81% in average success rate. Real-world robot experiments on five compositional and two in-context learning tasks further confirm the effectiveness of ViGAR.

In-context manipulation from a global goal image

With the initial scene and the language instruction held fixed, only a global goal image G is handed to the subgoal planner. Training demonstrations cover a restricted set of placement patterns; at test time an unseen goal specifies a task composition that was never demonstrated, and the planner decomposes it into subgoals that the unchanged policy executes.

In-domain goalOut-of-distribution goal

Task progress over 20 rollouts each. OOD averages combine two unseen goal configurations with different numbers of stages.

BibTeX

@article{gong2026rethinking,
  title={Rethinking World-Action Model for Compositional and In-Context Robotic Manipulation},
  author={Gong, Shukai and Zhai, Xuanran and Zhang, Yintianrun and Cui, Ruopeng and Huang, Ye and Fu, Yiyang and Lyu, Dexuan and Li, Chaojie and Song, Xinyi and Lin, Peiwen and others},
  journal={arXiv preprint arXiv:2610.02368},
  year={2026}
}