LESSON 3 OF 9 · 7 MIN
How they decide
What happens inside a System One model: a score for every option instead of written text, and a temperature that makes the probabilities honest.
No writing, just scores
A language model writes its answer one token at a time, and every token needs
another pass through the model. A System One model reads the request and
returns a score for every option. It generates nothing: in the inference API's
answer, output_tokens is always 0.
Models get there in two main ways. The open standard OpenDXP calls them profiles.
Encoders that score markers
Laya is built on ModernBERT, and Julia 1 on mmBERT. Both are encoders: models that read their whole input at once, in both directions.
The question, each option and the state are packed into one sequence, and each option gets a marker, a special token placed next to it. One pass through the encoder gives every marker a score. Adding an option adds a few tokens, not another pass.
OpenDXP calls this profile encoder-markers.
Language models that read letters
Decider and AnyJev start from a language model. The prompt lists the options with a letter each (A, B, C) and ends where the answer would begin. The model does not write the answer. Instead, the runtime reads the probability the model gives to each letter at that point.
OpenDXP calls this profile causal-letters. AnyJev asks each question once
for every order of its options and combines the results, so no option wins
because it came first. That takes more passes: on the shared CPU servers behind
the inference API, AnyJev answers in seconds, where Laya and Julia 1 take tens
of milliseconds.
From scores to probabilities
Either way, the model ends with one raw score per option. The scores are divided by a temperature, and a softmax turns them into probabilities that add up to 1.
The temperature is fitted after training, on examples the model did not train
on. It is what makes 0.9 mean "right about 9 times in 10". A model's
temperature for each question type is part of its package, in
calibration.json.
A score question works the same way over its levels, and reports the expected level. A noul question has two answers, true and false, and reports the probability of true.
Questions and decisions
One request can ask up to 8 questions about the same state. The inference API counts each question answered as one decision.
Try it
Watch the probabilities share out
In the Julia 1 playground, ask a choice question. Then add an option that fits the message better, and run it again. The probabilities always add up to 1, so the new option takes its share from the others.
Ask a score question
Add a score question, such as "How urgent is this?", with the levels
low,mediumandhigh. The answer is the expected level, with a probability for each level.
Read more
Every lesson is open to read. Sign in to tick them off and claim the certificate at the end.