DOCUMENTATION
Inference API
Call the models the System One Engine serves from your own code: API keys, the request and the answer, plans and limits, and errors.
View as Markdown · for AI agents: llms.txt · skill.md
Call a System One model from your own code: send a state and a few typed questions, and get a calibrated probability for every option back in one request. The models are the ones with a live playground here, kept loaded by the System One Engine.
- Base URL:
https://api.systemonemodels.tech - Requests, answers and errors follow the HTTP binding of OpenDXP
(
POST /v1/systemone), so a client written for an OpenDXP server works here too.
Get an API key
Create one at Settings → API. A key starts with s1_pat_
and is shown once. It calls models and nothing else: it cannot publish, and it
cannot see your private models, so a key that leaks from an app exposes nothing
but your allowance of decisions. Keep it on your server, never in a web page.
Tokens from systemone login can call models too.
Call a model from Python
pip install -U systemonemodels # 0.4 or later
export SYSTEMONE_API_KEY=s1_pat_…
from systemone import Client
client = Client() # reads SYSTEMONE_API_KEY; or Client("s1_pat_…")
result = client.decide(
"nokia/anyjev",
"Customer: I was charged twice and want my money back.",
{
"intent": {
"type": "choice",
"instructions": "What does the customer want?",
"criteria": ["refund", "track delivery", "cancel order"],
},
"angry": {"type": "noul", "instructions": "The customer is angry."},
},
)
print(result["answers"]["intent"]["choice"]) # refund
print(result["answers"]["angry"]["noul"]) # the probability that it is true
client.served_models() lists the models you can call, and client.usage()
what is left of your plan.
Over HTTP
curl https://api.systemonemodels.tech/v1/systemone \
-H "Authorization: Bearer $SYSTEMONE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "nokia/anyjev",
"state": "Customer: where is my parcel?",
"questions": {"intent": {"type": "choice",
"instructions": "What does the customer want?",
"criteria": ["refund", "track delivery", "cancel order"]}}}'
| Field | Meaning |
|---|---|
model | The model, as namespace/name |
state | The situation to decide about: text, up to 8,000 characters |
questions | 1 to 8 questions, each under an id you choose (letters, digits, ., _, :, -) |
checkpoint | Optional: one of the model's checkpoints; the first one answering when left out |
Each question has a type and instructions:
choicepicks one of 2 to 20 options.criteriais a list of names, or{"name": "description, or null"}.scoreplaces the state on an ordered scale.criteriais a list of 2 to 10 levels, lowest first.noulasks whether the statement ininstructionsis true.criteriais left out, or{"false": "…", "true": "…"}.
The answer
{
"model": "nokia/anyjev",
"version": "0.2.0-qwen3-1.7b",
"checkpoint": "f16",
"answers": {
"intent": { "type": "choice", "choice": "refund",
"probabilities": { "refund": 0.93, "track delivery": 0.02, "cancel order": 0.05 } },
"angry": { "type": "noul", "noul": 0.71 }
},
"usage": { "input_tokens": 212, "output_tokens": 0, "decisions": 2 },
"latency_ms": 840.5
}
A score answer carries score, the expected level, with its legend and a
probability per level. System One models answer without generating text, so
output_tokens is always 0.
Which models
GET /v1/systemone/models lists the models you can call, with their state
(ready; starting while one loads; offline), checkpoints and limits. It
needs no key. A model can be called once it has a live playground, which its
owner asks for from the model's Playground tab (how).
Plans and limits
A decision is one question answered, so a request with three questions uses three. A request the model could not answer costs nothing. Days and months are UTC.
| Plan | Decisions | Requests |
|---|---|---|
| Free | 500 a day | 60 a minute |
| Developer | 100,000 a month | 600 a minute |
| Enterprise | By agreement | 3,000 a minute |
Every answer says what is left, in X-Quota-Limit, X-Quota-Remaining and
X-Quota-Reset. GET /v1/usage, with your key or signed in, gives your plan
and your calls by day, model and key; Settings → API shows the
same. For the Developer or Enterprise plan, write to [email protected].
The models run on shared CPU servers today: answers take from tens of milliseconds (Laya, Julia 1) to several seconds (language models asked once per option order, like AnyJev).
Errors
The body has the registry's usual detail and code, and OpenDXP's
error with a type:
{ "detail": "You have used today's 500 decisions on the Free plan. …",
"code": "quota_exceeded", "errors": [],
"error": { "type": "request_refused", "message": "You have used today's 500 decisions on the Free plan. …" } }
| Status | code | What to do |
|---|---|---|
| 400 | validation_error | Fix the request; errors names the field |
| 401 | unauthorized | Send a valid key as Authorization: Bearer … |
| 403 | insufficient_scope | Use an API key: this token cannot call models |
| 404 | model_not_found, model_not_served | Check the name, or pick one from GET /v1/systemone/models |
| 422 | bad_question | The model cannot take this question; the message says why |
| 429 | rate_limited, quota_exceeded | Wait Retry-After seconds, or until your plan renews |
| 503 | starting, busy, timeout, api_paused | Try again after Retry-After seconds |