ROCKET-3-1.5x

This repository contains the released ROCKET-3 1.5x policy checkpoint from the paper Scalable Multi-Task Reinforcement Learning for Generalizable Spatial Intelligence in Visuomotor Agents.

ROCKET-3 is a vision-based, multi-task reinforcement-learning policy for visuomotor interaction. It uses a current view together with a cross-view goal specification and predicts Minecraft actions. The paper studies automated task synthesis and cross-view goal conditioning for spatial reasoning and interaction, including zero-shot transfer to unseen 3D environments.

Checkpoint

  • File: rocket3.ckpt
  • Format: PyTorch checkpoint containing the ROCKET-3 policy weights
  • Size: approximately 750 MB
  • License: MIT

The checkpoint is intended to be used with the inference and training code in the CraftJarvis/ROCKET-3 repository. The loader supports this bare state dictionary and infers the view-token count and previous-action conditioning from the checkpoint.

Quick start

Clone the code repository and install its dependencies:

git clone https://github.com/CraftJarvis/ROCKET-3.git
cd ROCKET-3
python -m pip install -r requirements.txt

Download the checkpoint and load it with the repository loader:

from huggingface_hub import hf_hub_download
from rocket3.policy import load_rocket3_policy

checkpoint = hf_hub_download(
    repo_id="CraftJarvis/ROCKET-3-1.5x",
    filename="rocket3.ckpt",
)
policy = load_rocket3_policy(checkpoint).eval()

The policy expects the observation structure used by the repository, including a current RGB view and a cross-view goal (cross_view_image, cross_view_obj_mask, and cross_view_obj_id). It returns action-policy logits and auxiliary value, visibility, and target-location predictions. See smoke_test.py for a minimal CPU checkpoint-load and synthetic inference check.

Running with Minecraft

The online training and environment integration are provided in the GitHub repository. They use MineStudio for the policy base class, Minecraft simulator, rollout manager, and PPO trainer. A complete run requires Python 3.10+, Java 8, a rendering setup such as Xvfb or VirtualGL, and (for the default distributed configuration) a Ray cluster.

python run_online.py \
  --checkpoint /path/to/rocket3.ckpt \
  --ray-address localhost:9899

Update the training configuration for your hardware and task setup before launching distributed training. The repository's smoke test can validate the checkpoint without starting Minecraft or Ray:

python smoke_test.py --checkpoint /path/to/rocket3.ckpt

Scope and limitations

  • This repository contains the released policy checkpoint only; the complete pretraining pipeline, dataset-generation pipeline, and the exact dependency versions used for the paper are not included.
  • The checkpoint is not packaged as a Transformers model and is not compatible with transformers.AutoModel or a generic text/image inference pipeline.
  • The model was developed for research in simulated 3D environments. It may produce unsafe, invalid, or ineffective actions outside the supported MineStudio/Minecraft observation and action interfaces.
  • The default multi-GPU training configuration and convergence are not guaranteed on every hardware or software setup. Validate behavior in your environment before relying on results.

Citation

@misc{cai2025scalable,
  title={Scalable Multi-Task Reinforcement Learning for Generalizable Spatial Intelligence in Visuomotor Agents},
  author={Shaofei Cai and Zhancun Mu and Haiwen Xia and Bowei Zhang and Anji Liu and Yitao Liang},
  year={2025},
  eprint={2507.23698},
  archivePrefix={arXiv},
  primaryClass={cs.RO}
}

Links

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Paper for CraftJarvis/ROCKET-3-1.5x