Marco-Mini-Instruct · REAP expert-pruned GGUFs

Marco-Mini-Instruct (17.3B total, ~0.86B active, 256 experts, 8 active per token), with 5% to 70% of its experts removed by REAP. Nine files from 9.3 GB down to 3.1 GB, all Q4_K_M, all measured on one RTX 3090.

Files

File Experts kept Size (GB) PPL pp512 (t/s) tg128 (t/s)
unpruned base (not in this repo) 256 10.00 8.17 6,221 —
Marco-Mini-Instruct-REAP05-Q4_K_M.gguf 244 9.33 8.59 5,160 155¹
Marco-Mini-Instruct-REAP10-Q4_K_M.gguf 231 8.85 9.30 6,482 281
Marco-Mini-Instruct-REAP15-Q4_K_M.gguf 218 8.36 10.28 6,806 256
Marco-Mini-Instruct-REAP20-Q4_K_M.gguf 205 7.88 11.40 6,005 277
Marco-Mini-Instruct-REAP30-Q4_K_M.gguf 179 6.94 14.64 7,661 278
Marco-Mini-Instruct-REAP40-Q4_K_M.gguf 154 5.97 18.74 6,970 278
Marco-Mini-Instruct-REAP50-Q4_K_M.gguf 128 5.00 26.42 6,966 316
Marco-Mini-Instruct-REAP60-Q4_K_M.gguf 102 4.07 38.62 9,335 287
Marco-Mini-Instruct-REAP70-Q4_K_M.gguf 77 3.10 64.43 9,713 278

¹ Measured in a separate session from the other rows; treat as an outlier, not a slowdown.

KL divergence is not listed for this family yet: our two measurement sessions disagree with each other, and we'd rather show nothing than a number we can't stand behind.

Which one to pick

  • REAP10–REAP15 (8.9–8.4 GB): the sweet spot, at most 1.26× the base perplexity.
  • REAP30 (6.9 GB): the best balance of size and quality if you need room on a 24 GB card for long context or a second model.
  • Under 5 GB: don't prune Mini further. Use Marco-Nano-Instruct-REAP-GGUF instead (see below).

The first 10% is nearly free (1.14× base perplexity). Past that, Mini loses quality faster than Nano: at REAP30 it sits at 1.79× its base, where Nano sits at 1.57×.

Generation speed barely changes with pruning, because a MoE model runs 8 experts per token whatever the total. Pruning buys you size, not speed.

Pick down, don't prune down

Across our sweeps, a smaller model pruned lightly beats a bigger model pruned hard at the same file size. Both Marco models share a tokenizer (151,936 tokens), so their perplexities compare directly:

Size budget Marco-Mini, pruned hard Marco-Nano, pruned lightly
~5 GB REAP50 · 5.00 GB · PPL 26.42 REAP05 · 4.77 GB · PPL 12.07
~4 GB REAP60 · 4.07 GB · PPL 38.62 REAP20 · 4.05 GB · PPL 15.36
~3 GB REAP70 · 3.10 GB · PPL 64.43 REAP40 · 3.11 GB · PPL 22.71

Rule of thumb: choose the smallest model that fits your memory, then prune it 5–20%. Past about 30%, each extra step costs a lot more quality than the one before.

How these were made

  • Pruning: REAP (Router-weighted Expert Activation Pruning, Cerebras Research), using the reference implementation. REAP scores each expert by its router weight times the size of its output over a calibration set, then removes the lowest-scoring experts whole. It is one-shot: no retraining. Seed 42.
  • Calibration set: theblackcat102/evol-codealpaca-v1 (train split, shuffled, seed 42), 64 samples per category, batch size 1, max sequence length 2048 tokens.
  • Quantisation: Q4_K_M with llama.cpp.
  • REAPxx in a file name is the share of experts removed, e.g. REAP20 = 20% of experts removed.

How we measured

  • Perplexity (PPL): WikiText-2 raw test split, 50 chunks, context 512, llama-perplexity. Lower is better.
  • Speed: llama-bench on one RTX 3090 (24 GB), all layers on GPU: prompt processing over 512 tokens (pp512) and generation over 128 tokens (tg128), 3 runs, llama.cpp builds b8164 and b8575.
  • Size: file size in GB.

Perplexity tracks how close a pruned model stays to its base, but it is not a task benchmark. We have not run MMLU, coding or multilingual evals on these files, so test them on your own workload before relying on them.

Run it

llama.cpp (OpenAI-compatible server):

llama-server --hf-repo kueizen/Marco-Mini-Instruct-REAP-GGUF --hf-file Marco-Mini-Instruct-REAP15-Q4_K_M.gguf -ngl 99

Ollama: this repo holds several files with the same quant type, so download the one you want and point a Modelfile at it:

FROM ./Marco-Mini-Instruct-REAP15-Q4_K_M.gguf
PARAMETER num_ctx 8192
ollama create marco-mini-instruct-reap15 -f Modelfile
ollama run marco-mini-instruct-reap15

LM Studio: search for kueizen/Marco-Mini-Instruct-REAP-GGUF and pick a file.

License and credits

These files are derived from ATH-MaaS/Marco-Mini-Instruct and released under the same Apache 2.0 licence. What we changed: removed experts with REAP and quantised the result to GGUF Q4_K_M. Nothing else was modified or retrained.

REAP is by Cerebras Research: paper, code. All credit for the base model goes to its authors.

The series

Published by Kueizen.

Downloads last month
254
GGUF
Model size
16B params
Architecture
qwen3moe
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kueizen/Marco-Mini-Instruct-REAP-GGUF

Quantized
(8)
this model

Paper for kueizen/Marco-Mini-Instruct-REAP-GGUF