LESSON 7 OF 9 · 8 MIN
Check, measure and calibrate with OpenDXP
Why run a model as an OpenDXP package: prove it answers like its own code, measure it on your hardware, and fit its confidence to your own data without retraining.
Why a package, and not the model's own code
A model's own code runs that one model. An OpenDXP package runs in every OpenDXP runtime, and it brings four things the model's own code does not:
- Proof. The package carries the answers the model's own code gave on a
fixed set of requests.
opendxp checkreplays them on your machine and tells you whether every decision is the same and every probability within 0.01, or exactly where they differ. - One runtime for every model. Encoders and language models alike answer the same request, over the same HTTP API and as the same MCP tool, and nothing in a package is executed.
- Speed where it counts. On an NVIDIA A40 GPU, in the same server, Laya's package answered 1.8 times the decisions per second of the laya package, and 2.1 times with requests batched. On a laptop CPU the gain depends on input length (faster on short requests, slower on long ones), so measure it on yours.
- Confidence you can fix. The temperatures that turn a model's scores into
probabilities are data in
calibration.json, so they can be refitted to your requests without touching the weights.
Check it
opendxp check odxp/
The report says whether the package is compatible: how many cases passed,
how many decisions were the same, and the largest difference in probability.
Run it on the machine and device you will serve from: a GPU's default
arithmetic can move probabilities, and --precision exact asks it for
float32 instead.
Measure it
opendxp bench odxp/ --device cpu --rounds 3
bench answers the package's requests one at a time and reports decisions
per second (a decision is one question answered) and latency percentiles.
Give it your machine's price, --usd-per-hour 0.51, and it adds the cost per
1,000 decisions. Point it at a server's URL instead of a folder to measure a
running server under load.
Calibrate it
A calibrated model's confidence means what it says: of the answers it gives at 90%, nine in ten are right. Many models are surer than they are right. On an outside benchmark, Julia 1 answered with a mean confidence of 96% and was right 72% of the time.
The fix is one temperature per question type, fitted to labelled requests. Each line of the labelled file is a request with the right answer to each question:
{"state": "Customer: I was charged twice and want my money back.", "questions": {"intent": {"type": "choice", "instructions": "What does the customer want?", "criteria": ["refund", "track delivery", "cancel order"]}}, "labels": {"intent": "refund"}}
opendxp calibrate odxp/ labelled.jsonl --test held-out.jsonl --out my-calibration.json
opendxp serve odxp/ --calibration my-calibration.json
calibrate prints confidence, accuracy and calibration error before and after,
and writes my-calibration.json. Fitted to 50 labelled requests, Julia 1's
confidence came to 72%, its accuracy, and no decision changed: the order of
the options stays the same.
Three rules make it work:
- Label with what was right, one option per question. Labels that are probabilities, such as another model's answers, fit the model to their spread instead.
- Use requests the model has not seen. On its own training data a model is right more often than it will be on yours, and the fit believes it.
- Keep some labels back (
--test) to see the result on requests the fit did not use.
The package itself does not change, so it still passes its check. A server
answering with a fitted file names it in /v1/models by its SHA-256.
Try it
Check a package
With a package in
odxp/(lesson 6), replay its conformance file on your machine.pip install -U "opendxp[onnx]" opendxp check odxp/
Measure it
Decisions per second and latency on your hardware. Add
--usd-per-hourwith your machine's price for the cost per 1,000 decisions.opendxp bench odxp/ --device cpu --rounds 3Calibrate it on your own labels
Write a few dozen requests you know the answers to into
labelled.jsonl, one per line, and a few more intoheld-out.jsonl. Then fit, and answer with the fitted file.opendxp calibrate odxp/ labelled.jsonl --test held-out.jsonl --out my-calibration.json opendxp serve odxp/ --calibration my-calibration.json
Read more
Every lesson is open to read. Sign in to tick them off and claim the certificate at the end.