pplx-decider-v1.1-27b

pplx-decider-v1.1-27b is an update of our pplx-decider-v1-27b model, that boost the Decision Index score from 56.4 to 61.56, now outperforming Jev by more than 3.5 points, while using the same backbone. Most of the gains are driven by the lifting of the causal mask and training on more data, including data from tasksource.

Evaluation

Decision Index category Jev pplx-decider-v1-27b pplx-decider-v1.1-27b
Knowledge 51.4 40.9 48.18
Language 62.0 63.5 69.45
Retrieval 55.4 54.9 61.26
Tools 75.1 79.3 78.88
Arts 37.7 39.4 44.66
Overall Decision Index 57.9 56.4 61.56

The overall score uses the suite's weighting, not a simple category average.

Checkpoint format and serving

This is the native decision-checkpoint layout: a Qwen3_5Model backbone plus readout.safetensors containing a BF16 [255, 5120] decision head. The Transformers class name does not change the base model identity, Qwen3.8-27B.

Use the included inference implementation to preserve the evaluated behavior:

  • Full-attention layers use the saved noncausal attention mode. The model's linear-attention layers retain their native behavior. Default causal inference does not reproduce this checkpoint's evaluated setup.
  • The supplied DecisionModel.predict applies the saved calibration temperature. When using logits directly, apply it exactly once and normalize over the valid candidates for the current decision.
  • This artifact has a separate readout, not a full-vocabulary lm_head. A serving system requiring Qwen3_5ForConditionalGeneration needs a separate export, mapping readout row i into vocabulary row decision_config.json["token_ids"][i]. Such an export must also preserve the attention behavior. If the frontend applies temperature, use the value above.

Usage

Use Python 3.12+, authenticated Hugging Face access to this private repository, and a CUDA GPU with room for roughly 49 GiB of weights plus working memory. Download the repository and install its pinned requirements.txt dependencies.

import sys
from pathlib import Path
from huggingface_hub import snapshot_download

checkpoint = Path(snapshot_download("perplexity-ai/pplx-decider-v1.1-27b"))
sys.path.insert(0, str(checkpoint / "source" / "src"))
from autojev.model import DecisionModel, answer

model = DecisionModel(checkpoint, device="cuda")
question = {
    "type": "choice",
    "instructions": "Which team should handle this request?",
    "criteria": {
        "billing": "Charges and refunds",
        "support": "Technical integration errors",
        "sales": "Questions about buying a product",
    },
}
row = {"state": "My integration keeps failing. Please help.", "question": question}
probabilities = model.predict([row])[0]
print(answer(question, probabilities))

source/src/autojev/model.py is copied byte-for-byte from the evaluated training run. release-manifest.json records file checksums and checkpoint provenance.

Acknowledgment

A large part of the gains are thanks to the tasksource data. If you also use this data, consider citing the corresponding article

@inproceedings{sileo-2024-tasksource,
    title = "tasksource: A Large Collection of {NLP} tasks with a Structured Dataset Preprocessing Framework",
    author = "Sileo, Damien",
    booktitle = "Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)",
    month = may,
    year = "2024",
    address = "Torino, Italia",
    publisher = "ELRA and ICCL",
    url = "https://aclanthology.org/2024.lrec-main.1361",
    pages = "15655--15684",
}
Downloads last month
12
Safetensors
Model size
26B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for perplexity-ai/pplx-decider-v1.1-27b

Base model

Qwen/Qwen3.8-27B
Finetuned
(480)
this model

Spaces using perplexity-ai/pplx-decider-v1.1-27b 2