DOCUMENTATION
Run on the System One Engine
How a decision model gets live inference on System One Models: the layouts the System One Engine runs, publishing, a live playground and analytics.
View as Markdown · for AI agents: llms.txt · skill.md
The System One Engine is our inference engine for System One models. It keeps approved models loaded on our servers, on the best device each can use, and answers in the System One request format. Every live playground on this site runs on it.
A model that runs on the engine gets:
- A live playground here. Anyone signed in can try it, answered at once by a model that is already loaded.
- Analytics. Requests served and failed, answer times, and the people who tried it, on the model's Analytics page.
- An inference API. Anyone with an API key can call it from their own code, within their plan: see the Inference API.
Make your model run on it
The engine runs a model when its files match a layout it knows. Each layout has a runtime: fixed code in the engine that turns those files into answers. Nothing that comes with a model is ever executed.
Laya checkpoints
Everything System One Studio publishes has this layout already.
rl_agent_config.json the decision head and its calibration
model.safetensors the weights
encoder/config.json a ModernBERT encoder config
tokenizer/tokenizer.json
tokenizer/tokenizer_config.json
A repository can hold several checkpoints, one per folder
(multilingual/rl_agent_config.json and so on). Each becomes a choice in the
playground.
Julia-style encoders
julia_config.json format_version 1, architecture JuliaDecisionModel, float32
model.safetensors
encoder/config.json
tokenizer/tokenizer.json
tokenizer/tokenizer_config.json
Letter-reading language models (Decider)
A language model that answers by the probability of each option's letter at
an answer slot. Set architecture: decider in systemone.yaml and publish:
your-model-Q8_0.gguf a GGUF build; Q8_0 is preferred over Q6_K, Q5_K_M, Q4_K_M
tokenizer.json
tokenizer_config.json
decider_config.json the prompt layout and a temperature per question type
Any other model: an OpenDXP package
A model of any other architecture runs from an OpenDXP
package: portable weights (ONNX or GGUF), a declarative template, calibration
and a conformance file, in an odxp/ folder of the version. The engine needs
no code written for the model, and checks the package against its conformance
file once it is serving; a model that passes shows OpenDXP compatible on
its page.
What every layout needs
- Files published here. The engine runs only files System One Models
stores.
systemone pushuploads them and records each file's SHA-256; a version that points at Hugging Face is listed but does not run. - Safetensors, GGUF or ONNX weights. No pickles, and no configs that name
code to import (
auto_map). - A size a playground can serve. Our servers run models up to about 2.5B parameters on the CPU today; bigger ones need GPUs, which are coming with the inference API.
- The System One answer format. Choice answers carry the chosen option and a probability per option, score answers the expected level and a probability per level, and noul answers the probability that the statement holds.
Get a live playground
- Publish the model with
systemone push. - Open its Playground tab and choose Request a live playground, with a note for the reviewer if you like.
- Once it is approved, the engine downloads and loads the model, usually within a minute. The playground starts answering, and the model's Analytics page (next to Edit on its page) shows how it is doing.
Launching a new model and want its playground live on launch day? Write to [email protected].