pplx-decider-v1.1-27b
pplx-decider-v1.1-27b is an update of our pplx-decider-v1-27b model, that boost the Decision Index score from 56.4 to 61.56, now outperforming Jev by more than 3.5 points, while using the same backbone. Most of the gains are driven by the lifting of the causal mask and training on more data, including data from tasksource.
Evaluation
| Decision Index category | Jev | pplx-decider-v1-27b | pplx-decider-v1.1-27b |
|---|---|---|---|
| Knowledge | 51.4 | 40.9 | 48.18 |
| Language | 62.0 | 63.5 | 69.45 |
| Retrieval | 55.4 | 54.9 | 61.26 |
| Tools | 75.1 | 79.3 | 78.88 |
| Arts | 37.7 | 39.4 | 44.66 |
| Overall Decision Index | 57.9 | 56.4 | 61.56 |
The overall score uses the suite's weighting, not a simple category average.
Checkpoint format and serving
This is the native decision-checkpoint layout: a Qwen3_5Model backbone
plus readout.safetensors containing a BF16 [255, 5120] decision head.
The Transformers class name does not change the base model identity, Qwen3.8-27B.
Use the included inference implementation to preserve the evaluated behavior:
- Full-attention layers use the saved noncausal attention mode. The model's linear-attention layers retain their native behavior. Default causal inference does not reproduce this checkpoint's evaluated setup.
- The supplied
DecisionModel.predictapplies the saved calibration temperature. When using logits directly, apply it exactly once and normalize over the valid candidates for the current decision. - This artifact has a separate readout, not a full-vocabulary
lm_head. A serving system requiringQwen3_5ForConditionalGenerationneeds a separate export, mapping readout rowiinto vocabulary rowdecision_config.json["token_ids"][i]. Such an export must also preserve the attention behavior. If the frontend applies temperature, use the value above.
Usage
Use Python 3.12+, authenticated Hugging Face access to this private repository,
and a CUDA GPU with room for roughly 49 GiB of weights plus working memory.
Download the repository and install its pinned requirements.txt dependencies.
import sys
from pathlib import Path
from huggingface_hub import snapshot_download
checkpoint = Path(snapshot_download("perplexity-ai/pplx-decider-v1.1-27b"))
sys.path.insert(0, str(checkpoint / "source" / "src"))
from autojev.model import DecisionModel, answer
model = DecisionModel(checkpoint, device="cuda")
question = {
"type": "choice",
"instructions": "Which team should handle this request?",
"criteria": {
"billing": "Charges and refunds",
"support": "Technical integration errors",
"sales": "Questions about buying a product",
},
}
row = {"state": "My integration keeps failing. Please help.", "question": question}
probabilities = model.predict([row])[0]
print(answer(question, probabilities))
source/src/autojev/model.py is copied byte-for-byte from the evaluated training
run. release-manifest.json records file checksums and checkpoint provenance.
Acknowledgment
A large part of the gains are thanks to the tasksource data. If you also use this data, consider citing the corresponding article
@inproceedings{sileo-2024-tasksource,
title = "tasksource: A Large Collection of {NLP} tasks with a Structured Dataset Preprocessing Framework",
author = "Sileo, Damien",
booktitle = "Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)",
month = may,
year = "2024",
address = "Torino, Italia",
publisher = "ELRA and ICCL",
url = "https://aclanthology.org/2024.lrec-main.1361",
pages = "15655--15684",
}
- Downloads last month
- 12
Model tree for perplexity-ai/pplx-decider-v1.1-27b
Base model
Qwen/Qwen3.8-27B