ROCKET-3-1.5x
This repository contains the released ROCKET-3 1.5x policy checkpoint from the paper Scalable Multi-Task Reinforcement Learning for Generalizable Spatial Intelligence in Visuomotor Agents.
ROCKET-3 is a vision-based, multi-task reinforcement-learning policy for visuomotor interaction. It uses a current view together with a cross-view goal specification and predicts Minecraft actions. The paper studies automated task synthesis and cross-view goal conditioning for spatial reasoning and interaction, including zero-shot transfer to unseen 3D environments.
Checkpoint
- File:
rocket3.ckpt - Format: PyTorch checkpoint containing the ROCKET-3 policy weights
- Size: approximately 750 MB
- License: MIT
The checkpoint is intended to be used with the inference and training code in the CraftJarvis/ROCKET-3 repository. The loader supports this bare state dictionary and infers the view-token count and previous-action conditioning from the checkpoint.
Quick start
Clone the code repository and install its dependencies:
git clone https://github.com/CraftJarvis/ROCKET-3.git
cd ROCKET-3
python -m pip install -r requirements.txt
Download the checkpoint and load it with the repository loader:
from huggingface_hub import hf_hub_download
from rocket3.policy import load_rocket3_policy
checkpoint = hf_hub_download(
repo_id="CraftJarvis/ROCKET-3-1.5x",
filename="rocket3.ckpt",
)
policy = load_rocket3_policy(checkpoint).eval()
The policy expects the observation structure used by the repository, including a current RGB view and a cross-view goal (cross_view_image, cross_view_obj_mask, and cross_view_obj_id). It returns action-policy logits and auxiliary value, visibility, and target-location predictions. See smoke_test.py for a minimal CPU checkpoint-load and synthetic inference check.
Running with Minecraft
The online training and environment integration are provided in the GitHub repository. They use MineStudio for the policy base class, Minecraft simulator, rollout manager, and PPO trainer. A complete run requires Python 3.10+, Java 8, a rendering setup such as Xvfb or VirtualGL, and (for the default distributed configuration) a Ray cluster.
python run_online.py \
--checkpoint /path/to/rocket3.ckpt \
--ray-address localhost:9899
Update the training configuration for your hardware and task setup before launching distributed training. The repository's smoke test can validate the checkpoint without starting Minecraft or Ray:
python smoke_test.py --checkpoint /path/to/rocket3.ckpt
Scope and limitations
- This repository contains the released policy checkpoint only; the complete pretraining pipeline, dataset-generation pipeline, and the exact dependency versions used for the paper are not included.
- The checkpoint is not packaged as a Transformers model and is not compatible with
transformers.AutoModelor a generic text/image inference pipeline. - The model was developed for research in simulated 3D environments. It may produce unsafe, invalid, or ineffective actions outside the supported MineStudio/Minecraft observation and action interfaces.
- The default multi-GPU training configuration and convergence are not guaranteed on every hardware or software setup. Validate behavior in your environment before relying on results.
Citation
@misc{cai2025scalable,
title={Scalable Multi-Task Reinforcement Learning for Generalizable Spatial Intelligence in Visuomotor Agents},
author={Shaofei Cai and Zhancun Mu and Haiwen Xia and Bowei Zhang and Anji Liu and Yitao Liang},
year={2025},
eprint={2507.23698},
archivePrefix={arXiv},
primaryClass={cs.RO}
}