Instructions to use zeromodels/gpt-oss-120b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- ZeroModels
How to use zeromodels/gpt-oss-120b with ZeroModels:
# pip install -U zeromodels # ZeroModels is pure Keras 3, so pick a backend: "jax", "torch" or "tensorflow". import os os.environ["KERAS_BACKEND"] = "jax" from zeromodels import AutoZModel # AutoZModel reads the repo's model_type and loads the matching class. # For a task head use the matching loader, e.g. AutoZMImageClassify / AutoZMDetect / # AutoZMSemanticSegment / AutoZMTextGenerate (see zeromodels.auto). model = AutoZModel.from_weights("zeromodels/gpt-oss-120b") - Keras
How to use zeromodels/gpt-oss-120b with Keras:
# !pip install -U keras tensorflow huggingface_hub # Keras needs TensorFlow installed to read "hf://" paths, so the tensorflow backend is selected here; # "jax" and "torch" also work for computation once TensorFlow is installed. import os os.environ["KERAS_BACKEND"] = "tensorflow" import keras model = keras.saving.load_model("hf://zeromodels/gpt-oss-120b") - Notebooks
- Google Colab
- Kaggle
See our collection for all GPT-OSS versions.
Run GPT-OSS with Keras 3: JAX, PyTorch, or TensorFlow
zeromodels/gpt-oss-120b
GPT-OSS is OpenAI's open-weight mixture-of-experts LLM family: top-k routed experts, learned per-head attention sinks, alternating sliding-window / full causal attention, and YaRN-scaled rotary positions. The experts ship in MXFP4 (4-bit).
For more details on the model, please see OpenAI's original model card.
Pure-Keras 3 conversion of openai/gpt-oss-120b
for zeromodels. One implementation runs unmodified on
TensorFlow / Torch / JAX, and the MoE experts are kept in MXFP4 exactly as
OpenAI ships them (uint8 nibble blocks + e8m0 scales), matching the official
footprint and dequantized on the fly at run time on every backend.
Quick start
import os
os.environ["KERAS_BACKEND"] = "torch" # or "jax" / "tensorflow"
from zeromodels.models.gpt_oss import GptOssTextGenerate, GptOssTokenizer
model = GptOssTextGenerate.from_weights("zeromodels/gpt-oss-120b")
tokenizer = GptOssTokenizer.from_weights("zeromodels/gpt-oss-120b")
inputs = tokenizer([
{"role": "user", "content": "Explain rotary embeddings in one sentence."}
])
outputs = model.generate(**inputs, max_new_tokens=64)
print(tokenizer.decode(outputs[0]))
Load any GPT-OSS variant the same way with from_weights("zeromodels/<variant>"):
| Variant | Hub | Weights |
|---|---|---|
gpt-oss-20b |
zeromodels/gpt-oss-20b |
MXFP4 MoE |
gpt-oss-120b |
zeromodels/gpt-oss-120b |
MXFP4 MoE |
Tips
- Set
KERAS_BACKENDbefore importing Keras / zeromodels. - Community / upstream safetensors also work via the
hf:prefix, e.g.GptOssTextGenerate.from_weights("hf:openai/gpt-oss-120b"). - See the GPT-OSS docs.
Special Thanks
A huge thank you to the OpenAI GPT-OSS authors for creating and releasing these models under Apache 2.0.
License: Apache 2.0.
- Downloads last month
- 3
Model tree for zeromodels/gpt-oss-120b
Base model
openai/gpt-oss-120b