# Inference API

> Call the models the System One Engine serves from your own code: API keys, the request and the answer, plans and limits, and errors.

Source: https://systemonemodels.tech/docs/inference

Call a System One model from your own code: send a state and a few typed
questions, and get a calibrated probability for every option back in one
request. The models are the ones with a live playground here, kept loaded by
the System One Engine.

- Base URL: `https://api.systemonemodels.tech`
- Requests, answers and errors follow the HTTP binding of [OpenDXP](https://systemonemodels.tech/docs/opendxp)
  (`POST /v1/systemone`), so a client written for an OpenDXP server works here too.

## Get an API key

Create one at [Settings → API](https://systemonemodels.tech/settings/api). A key starts with `s1_pat_`
and is shown once. It calls models and nothing else: it cannot publish, and it
cannot see your private models, so a key that leaks from an app exposes nothing
but your allowance of decisions. Keep it on your server, never in a web page.
Tokens from `systemone login` can call models too.

## Call a model from Python

```bash
pip install -U systemonemodels     # 0.4 or later
export SYSTEMONE_API_KEY=s1_pat_…
```

```python
from systemone import Client

client = Client()  # reads SYSTEMONE_API_KEY; or Client("s1_pat_…")
result = client.decide(
    "nokia/anyjev",
    "Customer: I was charged twice and want my money back.",
    {
        "intent": {
            "type": "choice",
            "instructions": "What does the customer want?",
            "criteria": ["refund", "track delivery", "cancel order"],
        },
        "angry": {"type": "noul", "instructions": "The customer is angry."},
    },
)
print(result["answers"]["intent"]["choice"])   # refund
print(result["answers"]["angry"]["noul"])      # the probability that it is true
```

`client.served_models()` lists the models you can call, and `client.usage()`
what is left of your plan.

## Over HTTP

```bash
curl https://api.systemonemodels.tech/v1/systemone \
  -H "Authorization: Bearer $SYSTEMONE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "nokia/anyjev",
       "state": "Customer: where is my parcel?",
       "questions": {"intent": {"type": "choice",
                                "instructions": "What does the customer want?",
                                "criteria": ["refund", "track delivery", "cancel order"]}}}'
```

| Field | Meaning |
| --- | --- |
| `model` | The model, as `namespace/name` |
| `state` | The situation to decide about: text, up to 8,000 characters |
| `questions` | 1 to 8 questions, each under an id you choose (letters, digits, `.`, `_`, `:`, `-`) |
| `checkpoint` | Optional: one of the model's checkpoints; the first one answering when left out |

Each question has a `type` and `instructions`:

- `choice` picks one of 2 to 20 options. `criteria` is a list of names, or
  `{"name": "description, or null"}`.
- `score` places the state on an ordered scale. `criteria` is a list of 2 to
  10 levels, lowest first.
- `noul` asks whether the statement in `instructions` is true. `criteria` is
  left out, or `{"false": "…", "true": "…"}`.

## The answer

```json
{
  "model": "nokia/anyjev",
  "version": "0.2.0-qwen3-1.7b",
  "checkpoint": "f16",
  "answers": {
    "intent": { "type": "choice", "choice": "refund",
                "probabilities": { "refund": 0.93, "track delivery": 0.02, "cancel order": 0.05 } },
    "angry": { "type": "noul", "noul": 0.71 }
  },
  "usage": { "input_tokens": 212, "output_tokens": 0, "decisions": 2 },
  "latency_ms": 840.5
}
```

A score answer carries `score`, the expected level, with its `legend` and a
probability per level. System One models answer without generating text, so
`output_tokens` is always 0.

## Which models

`GET /v1/systemone/models` lists the models you can call, with their state
(`ready`; `starting` while one loads; `offline`), checkpoints and limits. It
needs no key. A model can be called once it has a live playground, which its
owner asks for from the model's Playground tab ([how](https://systemonemodels.tech/docs/engine)).

## Plans and limits

A decision is one question answered, so a request with three questions uses
three. A request the model could not answer costs nothing. Days and months are
UTC.

The current limits are on [Settings → API](/settings/api).

Every answer says what is left, in `X-Quota-Limit`, `X-Quota-Remaining` and
`X-Quota-Reset`. `GET /v1/usage`, with your key or signed in, gives your plan
and your calls by day, model and key; [Settings → API](https://systemonemodels.tech/settings/api) shows the
same. For the Developer or Enterprise plan, write to ceo@systemonemodels.tech.

The models run on shared CPU servers today: answers take from tens of
milliseconds (Laya, Julia 1) to several seconds (language models asked once
per option order, like AnyJev).

## Errors

The body has the registry's usual `detail` and `code`, and OpenDXP's
`error` with a `type`:

```json
{ "detail": "You have used today's 500 decisions on the Free plan. …",
  "code": "quota_exceeded", "errors": [],
  "error": { "type": "request_refused", "message": "You have used today's 500 decisions on the Free plan. …" } }
```

| Status | `code` | What to do |
| --- | --- | --- |
| 400 | `validation_error` | Fix the request; `errors` names the field |
| 401 | `unauthorized` | Send a valid key as `Authorization: Bearer …` |
| 403 | `insufficient_scope` | Use an API key: this token cannot call models |
| 404 | `model_not_found`, `model_not_served` | Check the name, or pick one from `GET /v1/systemone/models` |
| 422 | `bad_question` | The model cannot take this question; the message says why |
| 429 | `rate_limited`, `quota_exceeded` | Wait `Retry-After` seconds, or until your plan renews |
| 503 | `starting`, `busy`, `timeout`, `api_paused` | Try again after `Retry-After` seconds |
