WorldGuideGoal-Directed Video World Model
for Procedural Task Execution
Abstract
Long-horizon procedural video generation requires a model to determine the next action from the state it has generated, carry out that action, and recognize task completion. WorldGuide studies this setting through a closed-loop planner–executor framework. Starting from an initial image and a task goal, a ContextPlanner predicts an atomic action, an Executor generates its video clip, and the generated result informs the next decision or termination.
The two components learn from the same step-level procedural demonstrations. Actions are represented as interpretable language-level instructions, while hierarchical visual memory provides context across clips at bounded history-token cost. We introduce WorldGuide Bench, containing 58,679 videos across 245 tasks and 27 procedural categories. We evaluate task execution and video quality on WorldGuide Bench and Video-CraftBench, with ablations for replanning, visual feedback, and memory.
Relation to open-loop and planner–executor generation
Open-loop generation follows instructions specified before rollout. Existing planner–executor approaches introduce planning or feedback in different ways. WorldGuide combines procedural action supervision with decisions based on generated visual progress, including the decision to stop.
Method overview
WorldGuide alternates between action prediction and video generation. Each generated clip updates the planner’s recent history and the Executor’s visual memory.
Interpretable atomic actions
The ContextPlanner represents each decision as a natural-language instruction or a task-completion token. Pairing an instruction with its generated clip makes the intended action readable and allows planning and visual execution to be inspected separately.
Choose the next atomic action.
The planner reads the goal and the latest three clip–action pairs, then predicts an interpretable language-level atomic action. Each instruction describes the intended next step and is grounded in visible progress.
Procedural training
The ContextPlanner is initialized from Qwen2.5-VL-7B and learns next-action and completion prediction. The HunyuanVideo-1.5 Executor learns from paired action–video demonstrations, conditioned on the planner’s action embeddings. The planner uses the latest three clip–action pairs.
Hierarchical visual memory
The Executor represents recent history at finer spatial resolution and older history at coarser resolution. The compression schedule follows Yume-1.5 and FramePack. It bounds history-token cost while preserving temporal context across generated clips.
Quantitative results
We report the manuscript’s results on both benchmarks. Task Success measures whether the final generated state visibly achieves the goal; it is separate from the validity of a predicted text plan.
| Benchmark | Test samples | WorldGuide | MiniMax-H3 | Conditioning |
|---|---|---|---|---|
| WorldGuide Bench | 980 | 33.33 | 29.90 | Goal-only for WorldGuide; reference actions for MiniMax-H3 |
| Video-CraftBench | 294 | 47.69 | 32.73 | Initial image and task goal for both models |
WorldGuide selects its actions from the task goal and generated history.
WorldGuide and starred planner–executor baselines receive the initial image and task goal. Unstarred baselines receive reference action plans.
The reported WorldGuide–MiniMax-H3 difference on this benchmark is not statistically significant (p = 0.10); the input conditions also differ.
Full benchmark results · 21 models +
WorldGuide and starred methods receive the initial image and task goal. Unstarred baselines receive reference action plans.
| Model | Aesthetic ↑ | Imaging ↑ | Dynamic ↑ | Smoothness ↑ | Consistency ↑ | Alignment ↑ | Plan accuracy ↑ | Order ↑ | Repeat/Skip ↓ | Task success ↑ |
|---|---|---|---|---|---|---|---|---|---|---|
| Cosmos-Predict2.5-2B | 38.77 | 62.11 | 82.52 | 96.32 | 17.46 | 63.75 | 38.05 | 84.08 | 3.98 | 11.94 |
| Open-Sora 2.0-11B | 33.42 | 44.91 | 60.19 | 95.06 | 17.42 | 62.43 | 31.27 | 79.56 | 6.53 | 8.54 |
| HunyuanVideo-1.5-8.3B | 42.43 | 65.57 | 88.84 | 96.00 | 18.91 | 62.28 | 43.53 | 80.00 | 3.00 | 12.00 |
| Wan2.2-14B | 41.46 | 58.87 | 16.02 | 93.67 | 17.52 | 67.33 | 4.37 | 21.65 | 2.06 | 2.58 |
| CogVideoX-5B | 38.76 | 54.99 | 64.08 | 96.58 | 18.09 | 68.30 | 34.33 | 75.74 | 3.96 | 9.90 |
| Helios-14B | 39.32 | 56.06 | 72.33 | 95.90 | 18.60 | 63.99 | 25.71 | 76.47 | 1.51 | 8.04 |
| UniVideo-13B | 45.75 | 61.92 | 48.00 | 97.09 | 18.05 | 65.63 | 10.77 | 52.00 | 4.00 | 4.00 |
| FlashMotion-34B | 41.69 | 67.26 | 90.29 | 95.46 | 17.93 | 62.92 | 24.10 | 77.39 | 2.51 | 5.53 |
| MAGI-1-24B | 39.14 | 63.08 | 2.30 | 98.31 | 17.00 | 69.33 | 17.70 | 61.63 | 2.33 | 3.49 |
| RoboMaster-5B | 32.60 | 17.68 | 4.37 | 97.09 | 6.66 | 37.13 | 4.29 | 5.50 | 0.50 | 4.00 |
| SpMem-5B | 35.79 | 47.69 | 96.73 | 98.94 | 20.72 | 60.92 | 33.74 | 78.95 | 2.94 | 10.53 |
| Astra-1.3B | 32.00 | 43.43 | 97.67 | 95.73 | 15.03 | 65.40 | 27.23 | 60.71 | 0.64 | 11.90 |
| Yume-1.5-5B | 40.61 | 65.72 | 87.86 | 96.90 | 17.39 | 60.79 | 19.72 | 74.26 | 2.39 | 2.97 |
| SkyReels-V3-14B | 42.64 | 58.78 | 14.56 | 99.66 | 17.03 | 72.54 | 11.23 | 48.52 | 1.01 | 3.90 |
| MiniMax-H3-33B | 41.61 | 56.60 | 80.58 | 97.99 | 19.84 | 70.04 | 57.99 | 86.79 | 4.06 | 29.90 |
| LTX-2.5-22B | 43.85 | 57.92 | 85.92 | 99.29 | 15.63 | 70.10 | 16.14 | 49.76 | 1.46 | 7.32 |
| HY-WorldPlay-8.3B | 39.25 | 64.56 | 85.95 | 96.30 | 15.28 | 55.47 | 15.67 | 73.46 | 3.07 | 1.26 |
| PhysAgent* | 42.76 | 65.09 | 52.68 | 98.39 | 16.72 | 70.61 | 18.20 | 56.56 | 1.98 | 8.42 |
| TempAct-1.3B* | 44.68 | 71.06 | 100.00 | 97.81 | 15.23 | 65.09 | 9.11 | 29.68 | 0.50 | 3.98 |
| Bernini-14B* | 45.12 | 63.13 | 93.20 | 98.23 | 11.91 | 53.48 | 5.30 | 22.06 | 0.00 | 1.96 |
| WorldGuide | 49.36 | 64.77 | 98.06 | 98.96 | 31.13 | 73.85 | 62.85 | 92.42 | 6.06 | 33.33 |
The first five metrics assess video quality; the remaining five assess planning and execution. Higher is better except Repeat/Skip. Bold: best; underlined: second best. * Methods with their own planning–execution loops. Task Success measures visible goal achievement.
Ablation studies
The following variants examine the contributions of replanning, visual feedback, and Executor memory. The oracle variant replaces predicted actions with reference actions.
| Variant | WorldGuide Bench | Video-CraftBench |
|---|---|---|
| Open-loop | 11.71 | — |
| Without planner visual feedback | 14.72 | — |
| Without Executor memory | 27.56 | 29.78 |
| WorldGuide | 33.33 | 47.69 |
| Oracle ContextPlanner | 38.82 | — |
The memory variants use the same trained Executor checkpoint with its memory branch enabled or disabled. A dash denotes a result not reported in the ablation table.
Planner evaluation
We evaluate action prediction from ground-truth histories and autonomous planning from generated histories separately. In the latter setting, visual feedback raises Plan Success from 86.41% to 91.75% and reduces Repeat/Skip from 20.19 to 15.73.
| Setting | Fine-tuned | History K | Plan accuracy ↑ | Order ↑ | Repeat/Skip ↓ | Plan success ↑ |
|---|---|---|---|---|---|---|
| Ground-truth histories | ||||||
| Qwen2.5-VL baseline | 3 | 1.04 | 0.00 | 8.82 | 0.00 | |
| ContextPlanner | 1 | 56.48 | 42.35 | 21.16 | 39.17 | |
| ContextPlanner | 2 | 75.83 | 71.28 | 15.56 | 44.66 | |
| ContextPlanner | 3 | 79.66 | 72.56 | 14.71 | 47.04 | |
| Autonomous rollout · generated histories | ||||||
| WorldGuide (without visual feedback) | 3 | 63.84 | 77.46 | 20.19 | 86.41 | |
| WorldGuide planner–executor | 3 | 64.03 | 79.45 | 15.73 | 91.75 | |
K is the number of recent clip–action pairs. Fine-tuned denotes supervised fine-tuning. These text-plan scores judge whether a predicted procedure would achieve the goal if executed correctly; Plan Success is distinct from video-judged Task Success and termination accuracy.
Video-judged task scores and VBench video-quality scores measure different properties.
Video comparisons
Browse all 18 examples and compare WorldGuide with any of eight baselines. The videos retain their original dimensions, durations, and playback speeds.
MiniMax-H3
COMPARISONLoading the video collection…
Videos retain their original durations and playback speeds. Paired playback starts both videos; individual controls let you inspect each result.
Open the full nine-model comparisonFried-rice comparison +
WorldGuide Bench
WorldGuide Bench provides task goals, temporally grounded atomic-action clips, and explicit completion signals. The same demonstrations support planner decisions and action-conditioned video generation.
- Source videos
- 58,679
- Tasks / categories
- 245 / 27
- Training videos
- 57,699
- Test videos
- 980
The video-disjoint split holds out four demonstrations per task. It evaluates new demonstrations of known tasks. Annotations were produced with Gemini 2.5 Flash and assessed through a separate human audit.
Browse step-level annotations ↗Why a dedicated procedural dataset?
Closed-loop execution needs supervision for three connected decisions: which action follows observed progress, how that action changes the visual state, and when the task is complete. Existing instructional datasets provide narration, action labels, or localized steps, but need additional processing for this training formulation. WorldGuide combines ordered atomic-action clips, task goals, and explicit completion signals within each demonstration, supporting both ContextPlanner decisions and Executor training.
| Dataset | Size | Tasks / domains | Step-Clip | Atomic action | Coverage target | Task goal | Completion | Type | Public |
|---|---|---|---|---|---|---|---|---|---|
| EgoPlan-IT | 50K QA | — / 1 | ✗ | ✓ | ✗ | ✓ | ✗ | Planning QA | ✓ |
| Ego4D Goal-Step | 48K segments | 86 / 1 | ✓ | ✗ | ✗ | ✓ | ✗ | Localization | ✓ |
| COIN | 11.8K videos | 180 / 12 | Fixed-vocab | Partial | ✗ | ✓ | ✗ | Localization | ✓ |
| CrossTask (primary) | 2.75K videos | 18 / 4 | ✓ | Partial | ✗ | ✓ | ✗ | Weak localization | ✓ |
| YouCook2 | 2K videos | 89 / 1 | ✓ | ✗ | ✓ | ✓ | ✗ | Segmentation | ✓ |
| HT-Step | 19.7K videos | 433 / 1 | ✓ | Partial | ✗ | ✓ | ✗ | Grounding | ✓ |
| Assembly101 | 4.3K videos | 101 / 1 | ✓ | ✓ | ✓ | ✗ | ✗ | Recognition | ✓ |
| CaptainCook4D | 384 videos | 24 / 1 | ✓ | Partial | ✓ | ✓ | ✗ | Error/localization | ✓ |
| SemComp-Data | 1.27K | 21 / 6 | ✗ | ✗ | ✗ | ✓ | ✗ | Evaluation only | Partial |
| ActVideoGen | 118K clips | — / Multi | ✓ | Partial | Partial | ✓ | ✗ | Generation training | ✗ |
| EgoForge / X-Ego | 15K clips | — / Multi | ✗ | ✓ | ✗ | ✓ | ✗ | Generation training + evaluation | ✗ |
| Ego-Exo4D Keystep | 27.6K segments | 17 / 3 | ✓ | Partial | ✗ | ✓ | ✗ | Recognition | ✓ |
| WorldGuide (ours) | 59K videos | 245 / 27 | ✓ | ✓ | ✓ | ✓ | ✓ | Generation training + evaluation | ✓ |
Sizes retain their original units; total of 58,679 WorldGuide videos. Step-Clip denotes temporally localized step annotations; atomic action denotes single-action granularity; completion denotes explicit task-completion supervision. Coverage target describes an annotation protocol targeting all visible task-relevant steps, excluding background intervals; it is not a verified coverage rate.
Annotation audit. Three annotators assessed 245 videos (one per task), covering 3,168 steps. Majority-agreement approval rates across six annotation criteria range from 81.6% to 84.5%. This audit assesses dataset annotations separately from the generated-video human evaluation above.
Citation
@misc{worldguide,
title = {WorldGuide: Goal-Directed Video World Model for Procedural Task Execution},
author = {Deria, Ankan and Kumar, Komal and Cholakkal, Hisham and Khan, Fahad Shahbaz and Khan, Salman}
}For questions about WorldGuide, contact: ankan.deria@mbzuai.ac.ae