Instructions to use x-square-robot/X-Planner-9B-0916 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use x-square-robot/X-Planner-9B-0916 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="x-square-robot/X-Planner-9B-0916") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("x-square-robot/X-Planner-9B-0916") model = AutoModelForMultimodalLM.from_pretrained("x-square-robot/X-Planner-9B-0916", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use x-square-robot/X-Planner-9B-0916 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "x-square-robot/X-Planner-9B-0916" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "x-square-robot/X-Planner-9B-0916", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/x-square-robot/X-Planner-9B-0916
- SGLang
How to use x-square-robot/X-Planner-9B-0916 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "x-square-robot/X-Planner-9B-0916" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "x-square-robot/X-Planner-9B-0916", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "x-square-robot/X-Planner-9B-0916" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "x-square-robot/X-Planner-9B-0916", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use x-square-robot/X-Planner-9B-0916 with Docker Model Runner:
docker model run hf.co/x-square-robot/X-Planner-9B-0916
X-Planner is a planning front end for long-horizon robot manipulation. It combines a high-level instruction, synchronized multi-view observations, and optional execution history to predict the next action-grounded event. The resulting structured plan is passed to a downstream world-action model, making the intermediate planning state explicit and inspectable.
The released X-Planner-9B-0916 model provides the Qwen3.5-based multimodal planner. It supports two complementary interfaces: a readable event mode for structured natural-language planning states, and a unified mode that produces compact latent planning states with Staircase Decoding.
X-Planner architecture. Multi-view observations and instructions are converted into structured events or latent planning states for a downstream world-action model.
Resources
- Code and inference guide: X-Square-Robot/Xplanner
- Benchmark and video preview: x-square-robot/xplanner-benchmark
- Project page: X-Planner
- Technical report: arXiv:2609.25187 (PDF)
Release contents
| Property | Value |
|---|---|
| Release | X-Planner-9B-0916 |
| Architecture | Qwen3_5ForConditionalGeneration (Qwen3.5 9B architecture) |
| Stored parameters | 9,409,813,744 |
| Weight precision | BF16 |
| Serialization | Safetensors, sharded at 5 GB |
| Transformers version recorded by the checkpoint | 5.2.0 |
The repository includes all model weights, the model/generation configuration,
tokenizer, chat template, image/video processor configuration, and minimal
checkpoint metadata. Sharding preserves all 760 tensors bit for bit.
release_manifest.json records file hashes and the original, unsharded weights
SHA-256 for provenance. Training logs, optimizer state, and machine-specific
training paths are not required for inference and are not part of this release.
Download
python -m pip install -U huggingface_hub
hf download x-square-robot/X-Planner-9B-0916 \
--local-dir checkpoints/X-Planner-9B-0916
For reproducible runs, add --revision <commit> using the desired revision from
this repository's commit history.
Load with Transformers
Use Transformers with Qwen3.5 support (the checkpoint was saved with 5.2.0), PyTorch, and Accelerate. No custom remote model code is required.
import torch
from transformers import AutoModelForImageTextToText, AutoProcessor
checkpoint = "checkpoints/X-Planner-9B-0916"
processor = AutoProcessor.from_pretrained(checkpoint)
model = AutoModelForImageTextToText.from_pretrained(
checkpoint,
dtype=torch.bfloat16,
device_map="auto",
attn_implementation="sdpa",
).eval()
model.config.use_cache = True
model.config.text_config.use_cache = True
The weights occupy approximately 18.82 GB in BF16. Inference also needs memory for activations, visual tokens, and the generation cache; total memory use depends on the inputs and generation length.
Structured planning inference
The task-specific prompt, image preparation, history format, and output parser are defined by the X-Planner event-state runtime. After installing its dependencies and preparing an event snapshot, run from the code checkout:
python scripts/inference/run_event_planner.py \
--checkpoint checkpoints/X-Planner-9B-0916 \
--snapshot /path/to/event_snapshot \
--output-dir work_dirs/inference
The event-state CLI requires an X-Planner-compatible data backend and a prepared event snapshot. See the code repository's installation and data-preparation instructions for the backend's public availability. The benchmark's raw video manifest is not a prepared event snapshot.
Benchmark and evaluation scope
The published benchmark contains 1,500 episodes, 3,490 videos, and episode-level planning metadata, with playable multi-view videos in Dataset Preview. Full temporal scoring annotations and a fixed end-to-end evaluation protocol are separate from this media release.
This model card does not report a new evaluation of X-Planner-9B-0916. Results from the technical report or other checkpoint revisions should retain their original model and evaluation provenance.
License
Model weights are distributed under Apache 2.0, consistent with the Qwen3.5-9B architecture's upstream model release. The X-Planner source code is MIT-licensed. Benchmark data and media retain their respective upstream terms, as described in the dataset card.
Citation
@article{xplanner2026event,
title = {X-Planner: Event-Structured Task Planning for Embodied Intelligence},
author = {{X Square Robot Team}},
journal = {arXiv preprint arXiv:2609.25187},
year = {2026},
url = {https://arxiv.org/abs/2609.25187}
}
- Downloads last month
- 76