On an outside benchmark, Julia 1 says 96% and is right 72% of the time. Three numbers, one temperature per question type, fitted to 50 labelled requests (250 labelled questions), brought its confidence down to its accuracy without touching its weights or changing a single decision. Temperature scaling divides a model's scores by a number before the softmax, and in an OpenDXP package you can fit and serve it without retraining.
Six models, before and after
We scored six OpenDXP packages on typed-decisions (Apache-2.0), 2,000 decisions whose gold answers come from a teacher model: Julia 1 (Supersonic Labs), three Laya checkpoints (Convai Innovations), Decider 2B (Mark Marosi; our unofficial OpenDXP package, source: github.com/Mapika/decider) and our package of AnyJev L0 (Nokia's method) on Qwen3-1.7B. Each was fitted with right-answer labels on half of the test split and scored on the other half:
| Model | Accuracy | Confidence before -> after | ECE before -> after |
|---|---|---|---|
| Julia 1 | 0.721 | 0.956 -> 0.728 | 0.236 -> 0.030 |
| Laya (main) | 0.363 | 0.536 -> 0.370 | 0.173 -> 0.015 |
| Laya (multilingual) | 0.347 | 0.663 -> 0.344 | 0.317 -> 0.024 |
| Laya (typed-decisions)* | 0.766 | 0.553 -> 0.774 | 0.213 -> 0.021 |
| Decider 2B | 0.615 | 0.619 -> 0.585 | 0.072 -> 0.051 |
| AnyJev L0 (Qwen3-1.7B) | 0.486 | 0.934 -> 0.491 | 0.448 -> 0.032 |
* Convai's typed-decisions checkpoint. Laya's main checkpoint scores as expected: Convai Innovations' own card reports about the same for it on this suite, and offers the typed-decisions checkpoint for this kind of task.
ECE, the expected calibration error, is the gap between confidence and accuracy averaged over ten confidence bins; 0 is perfect. The overconfidence is the model's, not the package's: our rebuild of Julia 1's inference code, run in float32 on a CPU, scores the same as its package to the third decimal.
How many labels?
Julia 1, fitted on random subsets of labelled requests and scored on held-out ones:
| Labelled requests (5 questions each) | Julia 1 confidence | Accuracy | Calibration error (ECE) |
|---|---|---|---|
| none (as shipped) | 0.956 | 0.721 | 0.236 |
| 5 | 0.761 | 0.721 | 0.089 |
| 10 | 0.731 | 0.721 | 0.063 |
| 20 | 0.725 | 0.721 | 0.058 |
| 50 | 0.718 | 0.721 | 0.044 |
| 200 | 0.728 | 0.721 | 0.041 |
Fifty labelled requests bring Julia 1's confidence to its accuracy. Ten get most of the way.
What it means in practice
- No decision changes. Dividing every score by the same number never changes their order. (AnyJev combines several rotated prompts, where that isn't guaranteed; in all our runs its accuracy didn't move either.)
- It can't make a model more accurate. Laya's general checkpoint is right 36% of the time here before and after. It just stops claiming otherwise: a router can act on a calibrated 40%, but not on an uncalibrated 90%.
- Sometimes the fix is more confidence. Convai's typed-decisions checkpoint hedges like the teacher: it says 55% on average while being right 77% of the time. Its temperatures came out below 1 (0.43 to 0.57), sharpening its answers until its confidence matched its accuracy.
- Label with what was right. Fitted to the teacher's distributions instead, Julia 1 ended at 0.586 confidence and an ECE of 0.135, against 0.728 and 0.030 with right answers: it learned to hedge like the teacher. Label with a big model's answers only if you want a small model to reproduce its probabilities.
- Don't calibrate on data the model has seen. Julia 1 is right on 95.7% of the benchmark's training questions, against 72.1% on the test questions. Fitted on the training split, it still said 87% on the test split (ECE 0.154). Calibrate on a labelled sample of your own traffic.
- Some answers carry almost no information. On these labels, which come from a teacher model, the best
temperature for the general Laya checkpoints' yes/no questions was the highest the search allows, so the
calibrated model answers close to 50/50.
opendxp calibrateprints a note when that happens.
Do it with your own labels
The temperatures sit in the package's calibration.json, next to the weights, so you can fit them to your own
labelled examples and serve the result (opendxp 0.3.0 or later):
opendxp calibrate packages/julia-1 labelled.jsonl --test held-out.jsonl --out julia-cal.json
opendxp serve packages/julia-1 --calibration julia-cal.json
Each line of the labelled file is an ordinary System One request plus a label per question:
{"state": "I was charged twice.", "questions": {"intent": {"type": "choice", "criteria": ["refund", "cancel"]}}, "labels": {"intent": "refund"}}
A label can be an option, a distribution over the options, or a whole answer object from another model; add
--hard-labels to fit distribution labels to their leading option. calibrate fits one temperature per question
type (choice, score, yes/no) and prints confidence, accuracy, KL, Brier and ECE before and after. Fit on 50 or more
labelled requests the model wasn't trained on, and check on others with --test. The package stays untouched and
still passes its conformance check. --calibration also works with opendxp run, mcp and bench, and in Python
as opendxp.load(path, calibration="julia-cal.json").
Browse decision models at systemonemodels.tech, or read the OpenDXP docs.