WorldGuide

WorldGuideGoal-Directed Video World Model
for Procedural Task Execution

Ankan DeriaKomal KumarHisham CholakkalFahad Shahbaz KhanSalman Khan

01 / FOLD
Paper boatPaper folding
02 / BUILD
Block horseBlock assembly
03 / COOK
Fried riceCooking

Selected WorldGuide generations from an initial image and a task goal. Use the play controls to inspect each procedure.

Abstract

Long-horizon procedural video generation requires a model to determine the next action from the state it has generated, carry out that action, and recognize task completion. WorldGuide studies this setting through a closed-loop planner–executor framework. Starting from an initial image and a task goal, a ContextPlanner predicts an atomic action, an Executor generates its video clip, and the generated result informs the next decision or termination.

The two components learn from the same step-level procedural demonstrations. Actions are represented as interpretable language-level instructions, while hierarchical visual memory provides context across clips at bounded history-token cost. We introduce WorldGuide Bench, containing 58,679 videos across 245 tasks and 27 procedural categories. We evaluate task execution and video quality on WorldGuide Bench and Video-CraftBench, with ablations for replanning, visual feedback, and memory.

Relation to open-loop and planner–executor generation

Open-loop generation follows instructions specified before rollout. Existing planner–executor approaches introduce planning or feedback in different ways. WorldGuide combines procedural action supervision with decisions based on generated visual progress, including the decision to stop.

Figure. From open-loop synthesis to learned procedural execution.

Method overview

WorldGuide alternates between action prediction and video generation. Each generated clip updates the planner’s recent history and the Executor’s visual memory.

Architecture. Generated clips update the planner’s history and the Executor’s visual memory. The loop ends at predicted completion.

Interpretable atomic actions

The ContextPlanner represents each decision as a natural-language instruction or a task-completion token. Pairing an instruction with its generated clip makes the intended action readable and allows planning and visual execution to be inspected separately.

CONTEXTPLANNER

Choose the next atomic action.

The planner reads the goal and the latest three clip–action pairs, then predicts an interpretable language-level atomic action. Each instruction describes the intended next step and is grounded in visible progress.

Goal + recent visual and action historyNext atomic action

Procedural training

The ContextPlanner is initialized from Qwen2.5-VL-7B and learns next-action and completion prediction. The HunyuanVideo-1.5 Executor learns from paired action–video demonstrations, conditioned on the planner’s action embeddings. The planner uses the latest three clip–action pairs.

Hierarchical visual memory

The Executor represents recent history at finer spatial resolution and older history at coarser resolution. The compression schedule follows Yume-1.5 and FramePack. It bounds history-token cost while preserving temporal context across generated clips.

Figure. The longest schedule supports up to 1,366 history latent slots. Its nominal cost, including the separate reference, is at most eight full-resolution latent-frame equivalents before spatial padding. This is a token-cost bound, not eight retained frames.

Quantitative results

We report the manuscript’s results on both benchmarks. Task Success measures whether the final generated state visibly achieves the goal; it is separate from the validity of a predicted text plan.

Task Success (%) · overview of both benchmarks
BenchmarkTest samplesWorldGuideMiniMax-H3Conditioning
WorldGuide Bench98033.3329.90Goal-only for WorldGuide; reference actions for MiniMax-H3
Video-CraftBench29447.6932.73Initial image and task goal for both models
TASK SUCCESS RATE ↑
33.33%

WorldGuide selects its actions from the task goal and generated history.

+3.43%vs. MiniMax-H3

WorldGuide and starred planner–executor baselines receive the initial image and task goal. Unstarred baselines receive reference action plans.

The reported WorldGuide–MiniMax-H3 difference on this benchmark is not statistically significant (p = 0.10); the input conditions also differ.

Full benchmark results · 21 models +

WorldGuide and starred methods receive the initial image and task goal. Unstarred baselines receive reference action plans.

WorldGuide Bench Results
ModelAesthetic ↑Imaging ↑Dynamic ↑Smoothness ↑Consistency ↑Alignment ↑Plan accuracy ↑Order ↑Repeat/Skip ↓Task success ↑
Cosmos-Predict2.5-2B38.7762.1182.5296.3217.4663.7538.0584.083.9811.94
Open-Sora 2.0-11B33.4244.9160.1995.0617.4262.4331.2779.566.538.54
HunyuanVideo-1.5-8.3B42.4365.5788.8496.0018.9162.2843.5380.003.0012.00
Wan2.2-14B41.4658.8716.0293.6717.5267.334.3721.652.062.58
CogVideoX-5B38.7654.9964.0896.5818.0968.3034.3375.743.969.90
Helios-14B39.3256.0672.3395.9018.6063.9925.7176.471.518.04
UniVideo-13B45.7561.9248.0097.0918.0565.6310.7752.004.004.00
FlashMotion-34B41.6967.2690.2995.4617.9362.9224.1077.392.515.53
MAGI-1-24B39.1463.082.3098.3117.0069.3317.7061.632.333.49
RoboMaster-5B32.6017.684.3797.096.6637.134.295.500.504.00
SpMem-5B35.7947.6996.7398.9420.7260.9233.7478.952.9410.53
Astra-1.3B32.0043.4397.6795.7315.0365.4027.2360.710.6411.90
Yume-1.5-5B40.6165.7287.8696.9017.3960.7919.7274.262.392.97
SkyReels-V3-14B42.6458.7814.5699.6617.0372.5411.2348.521.013.90
MiniMax-H3-33B41.6156.6080.5897.9919.8470.0457.9986.794.0629.90
LTX-2.5-22B43.8557.9285.9299.2915.6370.1016.1449.761.467.32
HY-WorldPlay-8.3B39.2564.5685.9596.3015.2855.4715.6773.463.071.26
PhysAgent*42.7665.0952.6898.3916.7270.6118.2056.561.988.42
TempAct-1.3B*44.6871.06100.0097.8115.2365.099.1129.680.503.98
Bernini-14B*45.1263.1393.2098.2311.9153.485.3022.060.001.96
WorldGuide49.3664.7798.0698.9631.1373.8562.8592.426.0633.33

The first five metrics assess video quality; the remaining five assess planning and execution. Higher is better except Repeat/Skip. Bold: best; underlined: second best. * Methods with their own planning–execution loops. Task Success measures visible goal achievement.

Ablation studies

The following variants examine the contributions of replanning, visual feedback, and Executor memory. The oracle variant replaces predicted actions with reference actions.

Task Success (%) · manuscript Table 4
VariantWorldGuide BenchVideo-CraftBench
Open-loop11.71—
Without planner visual feedback14.72—
Without Executor memory27.5629.78
WorldGuide33.3347.69
Oracle ContextPlanner38.82—

The memory variants use the same trained Executor checkpoint with its memory branch enabled or disabled. A dash denotes a result not reported in the ablation table.

Planner evaluation

We evaluate action prediction from ground-truth histories and autonomous planning from generated histories separately. In the latter setting, visual feedback raises Plan Success from 86.41% to 91.75% and reduces Repeat/Skip from 20.19 to 15.73.

ContextPlanner on WorldGuide Bench — manuscript Table 5
SettingFine-tunedHistory KPlan accuracy ↑Order ↑Repeat/Skip ↓Plan success ↑
Ground-truth histories
Qwen2.5-VL baseline❄️31.040.008.820.00
ContextPlanner🔥156.4842.3521.1639.17
ContextPlanner🔥275.8371.2815.5644.66
ContextPlanner🔥379.6672.5614.7147.04
Autonomous rollout · generated histories
WorldGuide (without visual feedback)🔥363.8477.4620.1986.41
WorldGuide planner–executor🔥364.0379.4515.7391.75

K is the number of recent clip–action pairs. Fine-tuned denotes supervised fine-tuning. These text-plan scores judge whether a predicted procedure would achieve the goal if executed correctly; Plan Success is distinct from video-judged Task Success and termination accuracy.

Video-judged task scores and VBench video-quality scores measure different properties.

Video comparisons

Browse all 18 examples and compare WorldGuide with any of eight baselines. The videos retain their original dimensions, durations, and playback speeds.

WorldGuide

OURS

MiniMax-H3

COMPARISON

Loading the video collection…

Videos retain their original durations and playback speeds. Paired playback starts both videos; individual controls let you inspect each result.

Open the full nine-model comparison
Fried-rice comparison +
Expanded fried-rice comparison. A plausible individual action does not guarantee a coherent procedure. Complete rollouts, rather than these sampled frames alone, are used for task scoring.

WorldGuide Bench

WorldGuide Bench provides task goals, temporally grounded atomic-action clips, and explicit completion signals. The same demonstrations support planner decisions and action-conditioned video generation.

Source videos
58,679
Tasks / categories
245 / 27
Training videos
57,699
Test videos
980

The video-disjoint split holds out four demonstrations per task. It evaluates new demonstrations of known tasks. Annotations were produced with Gemini 2.5 Flash and assessed through a separate human audit.

Browse step-level annotations ↗

Figure 4. A diverse collection of everyday procedures.

Why a dedicated procedural dataset?

Closed-loop execution needs supervision for three connected decisions: which action follows observed progress, how that action changes the visual state, and when the task is complete. Existing instructional datasets provide narration, action labels, or localized steps, but need additional processing for this training formulation. WorldGuide combines ordered atomic-action clips, task goals, and explicit completion signals within each demonstration, supporting both ContextPlanner decisions and Executor training.

Procedural datasets and their supervision — manuscript Table 1
DatasetSizeTasks / domainsStep-ClipAtomic actionCoverage targetTask goalCompletionTypePublic
EgoPlan-IT50K QA— / 1✗✓✗✓✗Planning QA✓
Ego4D Goal-Step48K segments86 / 1✓✗✗✓✗Localization✓
COIN11.8K videos180 / 12Fixed-vocabPartial✗✓✗Localization✓
CrossTask (primary)2.75K videos18 / 4✓Partial✗✓✗Weak localization✓
YouCook22K videos89 / 1✓✗✓✓✗Segmentation✓
HT-Step19.7K videos433 / 1✓Partial✗✓✗Grounding✓
Assembly1014.3K videos101 / 1✓✓✓✗✗Recognition✓
CaptainCook4D384 videos24 / 1✓Partial✓✓✗Error/localization✓
SemComp-Data1.27K21 / 6✗✗✗✓✗Evaluation onlyPartial
ActVideoGen118K clips— / Multi✓PartialPartial✓✗Generation training✗
EgoForge / X-Ego15K clips— / Multi✗✓✗✓✗Generation training + evaluation✗
Ego-Exo4D Keystep27.6K segments17 / 3✓Partial✗✓✗Recognition✓
WorldGuide (ours)59K videos245 / 27✓✓✓✓✓Generation training + evaluation✓

Sizes retain their original units; total of 58,679 WorldGuide videos. Step-Clip denotes temporally localized step annotations; atomic action denotes single-action granularity; completion denotes explicit task-completion supervision. Coverage target describes an annotation protocol targeting all visible task-relevant steps, excluding background intervals; it is not a verified coverage rate.

Annotation audit. Three annotators assessed 245 videos (one per task), covering 3,168 steps. Majority-agreement approval rates across six annotation criteria range from 81.6% to 84.5%. This audit assesses dataset annotations separately from the generated-video human evaluation above.

Citation

@misc{worldguide,
  title = {WorldGuide: Goal-Directed Video World Model for Procedural Task Execution},
  author = {Deria, Ankan and Kumar, Komal and Cholakkal, Hisham and Khan, Fahad Shahbaz and Khan, Salman}
}

For questions about WorldGuide, contact: ankan.deria@mbzuai.ac.ae