Instructions to use infercrane/Commerce-1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use infercrane/Commerce-1 with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("infercrane/Commerce-1", device_map="auto") - Notebooks
- Google Colab
- Kaggle
InferCrane Commerce-1
A 27B, self-hosted decision model for commerce agents. Give Commerce-1 one structured state and ordered Choice, Noul, or Score questions; it returns calibrated probability distributions and typed answers—not prose—through an offline, content-verified runtime.
Canonical repository: infercrane/Commerce-1
Runtime and examples: github.com/infercrane/commerce-1
What it returns
| Primitive | You provide | Commerce-1 returns |
|---|---|---|
Choice |
Named candidate actions and their criteria | Selected action plus a probability for every action |
Noul |
One yes/no decision instruction | Boolean answer plus P(true) |
Score |
Ordered labels from low to high | Expected ordinal score plus a probability for every label |
The API also returns the raw ordered distributions, immutable model and runtime identities, and measured input-token use. Probabilities help callers set review thresholds, but they do not guarantee correctness or transfer unchanged to a new workload.
Intended use
Commerce-1 is designed for structured decisions inside commerce-agent workflows, including proceed/review/block gates, tool routing, approval-policy checks, action selection, and calibrated risk or priority scoring.
It is not a chat model, search engine, live-fact retriever, payment processor, or autonomous authority for refunds, disputes, fraud, credit, safety, or legal decisions. Validate on representative traffic and retain explicit controls for consequential actions.
Run locally
Use immutable 40-character revisions from the verified release receipts. Floating source branches and model revisions are not part of the qualified path.
export COMMERCE_ONE_SOURCE_REVISION=<verified-github-revision>
export COMMERCE_ONE_MODEL_REVISION=<verified-huggingface-revision>
git clone https://github.com/infercrane/commerce-1.git
cd commerce-1
git checkout --detach "$COMMERCE_ONE_SOURCE_REVISION"
test "$(git rev-parse HEAD)" = "$COMMERCE_ONE_SOURCE_REVISION"
uv sync --frozen --python 3.12 --extra v3-runtime
uv run --frozen --python 3.12 --extra v3-runtime hf download infercrane/Commerce-1 \
--revision "$COMMERCE_ONE_MODEL_REVISION" \
--local-dir ../commerce-1-model
export COMMERCE_ONE_MODEL_DIR="$(cd ../commerce-1-model && pwd)"
uv run --frozen --python 3.12 --extra v3-runtime commerce-1-v3-serve \
--model-dir "$COMMERCE_ONE_MODEL_DIR" \
--base-dir "$COMMERCE_ONE_MODEL_DIR" \
--offline \
--host 127.0.0.1 \
--port 8000
The native request/response schema and a commerce example are in
USAGE.md. The scorer is exposed at POST /v1/systemone and
POST /api/alpha/decisions; a strict JSON-only bridge is exposed at
POST /v1/chat/completions.
Measured evidence
| Evaluation | Result | Boundary |
|---|---|---|
| Decision Index 0.2.1 complete union | 60.99 | Self-hosted reproduction of the complete public suite |
| Completed rows | 150,317 / 150,317 | Unanswered rows are not removed |
| Official Decision Index 0.3 Full Score | Pending maintainer evaluation | Not yet an official score or rank |
The initial release is bound to the complete public 0.2.1 reproduction and an exact offline native-runtime qualification. Maintainers require immutable public model, runtime, and result links before they can run private 0.3 tests, so the Full Score remains pending at initial publication.
Qualified runtime
| Field | Measured value |
|---|---|
| Hardware | 1× NVIDIA H200 |
| Precision | bfloat16 |
| Runtime | PyTorch 2.14.0+cu130, CUDA 13.0 |
| Input limit | 65,536 tokens |
| Question microbatch | 1 |
| Fixture coverage | 3 fixtures: Choice, Noul, Score |
| Native/direct parity | Exact; maximum absolute difference 0.0 |
| Release/evaluated parity | Exact; maximum absolute difference 0.0 |
| Throughput | 8.836 decisions/s |
| Latency | 340.146 ms p50, 349.026 ms p95 per complete fixture set |
| Network during qualification | Disabled |
These figures describe the exact locked H200 qualification fixture and are not general serving benchmarks. Other accelerators, batching strategies, traffic shapes, and prompt lengths can produce different results. The complete safe attestation is QUALIFICATION.json.
Architecture
- Immediate parent:
perplexity-ai/pplx-decider-v1.1-27b, immutable revision3b45dead91dfa6d95aad6b95764a606fab2bf7a6. - Ultimate base family:
Qwen/Qwen3.8-27B, immutable revision1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0. - Native sharded Transformers checkpoint with a separate decision readout.
- Noncausal full-attention decision backbone with last-position pooling.
- Per-primitive held-out temperature calibration for Choice, Noul, and Score.
- Exact 65,536-token input limit with fail-closed overflow handling.
- Content-bound offline runtime; local bytes are verified before source import or model load.
Package contents
| Path | Purpose |
|---|---|
model/ |
Sharded BF16 checkpoint, tokenizer, and decision configuration |
calibration/ |
Bound held-out temperatures for the three primitives |
public-runtime/ |
Content-verified serving source closure |
source/autojev/ |
Minimal decision-model source closure required by the checkpoint |
evidence/ |
Sanitized evaluation and lineage attestations |
training-data/lineage.json |
Content-free training-source identity receipt |
runtime.json |
Locked runtime and backend identity |
release-manifest.json |
Exact package inventory and identities |
QUALIFICATION.json |
Content-free native runtime qualification |
SHA256SUMS |
Package checksum inventory |
See ARCHITECTURE.md, LOCAL_RUN.md, API.md, and REPRODUCIBILITY.md.
Training data and provenance
The exact model lineage, required training-source identities, and source notice requirements are in PROVENANCE.json. Upstream notices are retained in THIRD_PARTY.md. These documents identify the release inputs without publishing protected training rows or private control receipts.
Release identity
| Field | Verified value |
|---|---|
| Evaluated checkpoint SHA-256 | b7fba0cceb1b57834fecacfbcb5e361d94fa185e877e891f4e43fc25c6458a1d |
| Release checkpoint SHA-256 | f97262c59b102a39ae37826ffe1e1db8ddabaa1ec95b4e73bcb44328420ace57 |
| Checkpoint transform attestation | ac7ae1330b4e912697b11ddb41770121f37ff88047e5f2046e3a338f9307d561 |
| Native runtime identity | 8a2fa16c5419da709a3cbe26ea2339f7ad4604db16da76f9c5426bd5a6984739 |
| Public evaluation attestation | 0e9432adcde227639ae7427bdcf9b3fdd3fb6b023d844d118a3cf23c3aa44749 |
| Training-data lineage | d7a3c00a5fc71c30d10ea7cf92943ddc8a62925a276680aa9a862be53f970827 |
| Public runtime qualification | 2b01ab21b12c1d1c555688f205717708ade4a1b0a5b415aec08d36148c8e3639 |
| Publication gate | INITIAL_PUBLIC_RELEASE_PENDING_OFFICIAL_V0_3 |
The release checkpoint differs from the evaluated checkpoint only by the documented safe metadata projection; qualification verifies identical decision outputs. No generic hosted-inference widget is enabled because Commerce-1 uses a typed custom runtime rather than free-form text generation.
License and contact
The project license is in LICENSE. Review LICENSE, PROVENANCE.json, and THIRD_PARTY.md before redistribution or deployment. Report runtime or model-card issues through the InferCrane Commerce-1 issue tracker.