This repository contains the code and checkpoints for Lifting Embodied World Models for Planning and Control. A low-level PEVA world model is lifted using a diffusion policy that predicts low-level joint actions conditioned on high-level waypoint actions. Planning is performed using CEM by optimizing the high-level waypoints through the lifted world model using a DreamSIM objective.
- 2026-07-21 — Fixed default masking mode; options are
no_masking(default) anduniform.
train/
train_ddp.py # training entry point (DDP / torchrun)
plan_cem.py # CEM planning + evaluation entry point
config/
waypoint_policy.yaml # the canonical training config
defaults.yaml # default config values
vint_train/ # model + dataset + training-loop code
models/nomad/ # NoMaD policy (encoder + diffusion U-Net, vendored)
planning/ # CEM planner, world-model wrappers, rendering
peva/ # PEVA video world model
wm_generate_data.py # build the full ep_info.pt dataset from a download manifest
wm_preprocess.py # per-track worker: raw recording -> ep_info.pt (+ images/)
preprocess_camera_data.py # build per-track camera_data.pt (skeleton overlays)
scripts/repack_ep_info_dist8.py # optional: slim/repack ep_info.pt for fast loading
data_splits/nymeria/{train,test}/traj_names.txt # train/test split track lists
nomad_train.yml / nomad_train2.yml # conda envs (training / planning)
This codebase uses two conda environments — one for training, one for planning (they pin different dependency versions). Create whichever you need:
# training environment (name: nomad_train)
conda env create -f train/nomad_train.yml
conda activate nomad_train
# planning environment (name: nomad_train2)
conda env create -f train/nomad_train2.yml
# install the package (in whichever env you activate)
pip install -e train/A CUDA-capable GPU is required for both training and planning.
Download the pretrained checkpoints and point the commands below at them:
| Checkpoint | What it is | Link |
|---|---|---|
| Waypoint policy | the trained diffusion policy weights (ema_9.pth); the architecture is read from the in-repo config/waypoint_policy.yaml |
HF Hub |
| PEVA world model (Nymeria) | the video world model used by the CEM planner (*.pth.tar) |
HF Hub |
First get Nymeria access from the official dataset page
and download its JSON manifest (the per-track list + time-limited data links); save it as
data_jsons/all_data.json. The dataset loader then expects one directory per recording containing an
ep_info.pt (poses, projection/visibility matrices, joint offsets) and an images/ directory of
egocentric frames, produced as follows.
-
Build the per-track
ep_info.pt— generate the whole dataset at once (1a) or one track at a time (1b).1a — whole dataset from the download manifest (recommended).
wm_generate_data.pyreads the Nymeria download JSON, then for each track downloads the raw recording, builds itsep_info.pt, and (optionally) deletes the raw download to reclaim disk:python wm_generate_data.py \ --json_file ../data_jsons/all_data.json \ --raw_dir <NYMERIA_RAW_ROOT> \ --out_dir <NYMERIA_DATA_ROOT> \ --num_workers 4 --delete_raw # writes <NYMERIA_DATA_ROOT>/<track>/{ep_info.pt, images/} for every track in the manifest
It is idempotent (tracks with a complete
ep_info.ptare skipped, so the job resumes freely) and can be scoped with--match <substr>,--limit <N>, or--tracks_file <list.txt>(e.g. one shard per sbatch-array index). Add--skip_imagesto write only theep_info.pttensors.1b — single track with the per-track worker (what
wm_generate_data.pycalls under the hood) from an already-downloaded raw recording:python wm_preprocess.py \ --input_dir <NYMERIA_RAW_ROOT> \ --output_dir <NYMERIA_DATA_ROOT> \ --ep <track> [--skip_images] # writes <NYMERIA_DATA_ROOT>/<track>/{ep_info.pt, images/}
-
Camera data (optional, for skeleton/SMPL overlays in planning):
python preprocess_camera_data.py \ --ep_folder <NYMERIA_DATA_ROOT> \ --vrs_folder <NYMERIA_VRS_ROOT> \ --split nymeria/train nymeria/test # writes camera_data.pt next to each track's ep_info.pt
-
Optional fast-loading repack (slims
ep_info.ptfor the dist-8drawsetup):python scripts/repack_ep_info_dist8.py --src_base <NYMERIA_DATA_ROOT> --dst_base <SLIM_OUT_ROOT>
-
SMPL models (optional, only for skinned-mesh rendering): download SMPL-X and set
export SMPL_MODEL_DIR=/path/to/smplx/models(or pass--smpl_model_dirtoplan_cem.py).
Edit train/config/waypoint_policy.yaml and set datasets.nymeria.data_folder to your Nymeria data
root, then launch from train/:
cd train
torchrun --nproc_per_node=<NUM_GPUS> train_ddp.py --config config/waypoint_policy.yaml- Logs and checkpoints are written to
logs/waypoint_policy/<timestamp>:<run_name>/. - Weights & Biases logging is on by default (project
waypoint_policy) and logs to your authenticated W&B account. SetWANDB_ENTITYto override the entity, oruse_wandb: Falsein the config to disable.
The CEM planner optimizes goal waypoints through the PEVA world model and scores rollouts with a
DreamSIM objective. Run from train/ in the planning env:
cd train
python plan_cem.py \
--algo waypoint \
--model_dir <PATH_TO>/checkpoints/waypoint_policy \
--data_folder <NYMERIA_DATA_ROOT> \
--peva_checkpoint <PATH_TO>/checkpoints/peva/nymeria_peva.pth.tar \
--num_samples_to_plan 64 --min_dist_cat 8 --max_dist_cat 8Key arguments:
| Arg | Meaning |
|---|---|
--algo |
waypoint (optimize 2D waypoints — the main mode) or peva (optimize PEVA actions directly) |
--model_dir |
trained policy dir holding the ema_9.pth checkpoint; config defaults to in-repo config/waypoint_policy.yaml (override via --nomad_config) |
--data_folder |
Nymeria data root (required) |
--peva_checkpoint |
PEVA world-model checkpoint (required); --peva_config defaults to the in-repo config |
--camera_data_folder |
optional; enables skeleton/SMPL overlays in the saved visualizations |
--smpl_model_dir |
SMPL-X dir for skinned-mesh rendering (or set SMPL_MODEL_DIR); only with skin rendering |
--no_skin |
skip skinned-mesh rendering (faster) |
Results (metrics + visualizations) are written to logs/cem/<timestamp>:<run_name>/.
This repository is derived from visualnav-transformer (GNM / ViNT / NoMaD, Berkeley AI Research) and builds on:
- NoMaD — Goal Masking Diffusion Policies for Navigation and Exploration (the policy architecture).
- Diffusion Policy — Chi et al. (the vendored
ConditionalUnet1D; MIT-licensed, © 2023 Columbia Artificial Intelligence and Robotics Lab — repo). - PEVA — the embodied video world model used by the planner.
- Nymeria — the egocentric human-motion dataset (Meta Reality Labs / Project Aria).
If you use this code, please cite the paper:
@article{wang2026lifting,
title = {Lifting Embodied World Models for Planning and Control},
author = {Wang, Alex N. and Darrell, Trevor and Izmailov, Pavel
and Bai, Yutong and Bar, Amir},
journal = {arXiv preprint arXiv:2604.26182},
year = {2026},
url = {https://arxiv.org/abs/2604.26182}
}See LICENSE.


