MiniCPM5-2B — FP8 (compressed-tensors)

FP8-dynamic quantization of openbmb/MiniCPM5-2B, weights and activations in float8_e4m3 (per-channel weights, per-token dynamic activations), lm_head and embeddings kept in bf16. Produced with llm-compressor 0.13.

  • 2.84 GiB on disk (bf16 base is 4.68 GiB — 39 % smaller)
  • Serves on vLLM (compressed-tensors), native FP8 tensor-core path on Ada / Blackwell
  • Near-lossless on coding benchmarks — see the eval below

This is the recommended quant for MiniCPM5-2B when you want maximum quality retention. If you need to fit closer to 2 GB, see the NVFP4 W4A16 build (smaller, ~2–5 pp coding cost).

Evaluation

lm-evaluation-harness, vLLM backend, greedy decoding, 3 draws each (the harness is non-deterministic run-to-run even at greedy — median and range are reported; a single draw is not a reproducible score).

build HumanEval-instruct MBPP (3-shot) size
bf16 base 86.59 % (85.98–86.59) 50.60 % (50.40–51.00) 4.68 GiB
FP8 (this) 84.76 % (84.15–85.37) 48.80 % (48.80–49.00) 2.84 GiB
Δ vs bf16 −1.8 pp −1.8 pp −39 %

Both deltas sit inside the eval's own run-to-run spread — FP8 is effectively lossless here. (For contrast, plain RTN NVFP4-W4A16 of the same model loses 6.7 / 9.4 pp; GPTQ-recovered NVFP4-W4A16 loses 2.4 / 4.8 pp.)

v1 — updated as more evals land (RULER long-context, agentic SWE-style, IFEval).

Usage

vllm serve Ttimms/MiniCPM5-2B-FP8 --max-model-len 32768 --kv-cache-dtype fp8
from vllm import LLM, SamplingParams
llm = LLM("Ttimms/MiniCPM5-2B-FP8", trust_remote_code=True)
print(llm.chat([{"role": "user", "content": "Write a Python LRU cache."}],
               SamplingParams(temperature=0.6, max_tokens=512))[0].outputs[0].text)

Method & provenance

  • Quantizer: llm-compressor 0.13, QuantizationModifier(scheme="FP8_DYNAMIC"), ignore lm_head + embed_tokens. No calibration data required (dynamic activations).
  • Base: openbmb/MiniCPM5-2B (LlamaForCausalLM, 2.5 B, Apache-2.0).
  • Built and evaluated on an RTX 5070 Ti (Blackwell, SM120), vLLM 0.26.
  • Full quant-format comparison + methodology: https://github.com/t-timms/blackwell-16gb-moe

License

Apache-2.0, inherited from openbmb/MiniCPM5-2B.

Downloads last month
-
Safetensors
Model size
3B params
Tensor type
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Ttimms/MiniCPM5-2B-FP8

Quantized
(36)
this model