In about ten minutes you can give Claude Desktop, Claude Code, Cursor or any other client of the Model Context Protocol an open decision model as a tool. It answers typed questions in one pass, in well under 100 milliseconds on an Apple M4's CPU (the first call after starting included), with a probability for every option that the agent can act on. You need opendxp 0.3.0 or later.
Why bother: an agent handling support tickets makes the same small decisions over and over. What does this customer want? Are they angry? Should a person see this one? You can ask the agent's LLM each time, but that costs a generation, the answer comes back as text, and "I'm fairly sure it's a refund" is not a number your code can act on.
1. Get a model and package it
We'll use Julia 1, an open decision model from Supersonic Labs that scores a marker token for each option. First download it, then convert it into an OpenDXP package: an ONNX graph plus its template and calibration as data, with nothing executable.
pip install systemonemodels "opendxp[export,onnx]>=0.3.0"
systemone pull supersonic-labs/julia-1 --dest ./models
opendxp export julia ./models/julia-1 ./julia-odxp
The export took about 20 seconds on an Apple M4. It checks the graph against the PyTorch model it came from before
it writes julia-odxp/odxp.json.
2. Add it to your MCP client
opendxp mcp serves one or more packages as MCP tools over stdio. In Claude Desktop, add it to the configuration
file:
{
"mcpServers": {
"julia-1": {
"command": "opendxp",
"args": ["mcp", "/absolute/path/to/julia-odxp"]
}
}
}
In Claude Code, add it with one command:
claude mcp add julia-1 -- opendxp mcp /absolute/path/to/julia-odxp
3. What the agent sees
The server offers one tool, decide, described to the agent as a way to ask the model typed questions about a
state: pick one option (choice), place something on a scale (score), or say whether something holds (noul).
The agent calls it like this:
{
"state": "I was charged twice for the same order. Please give me my money back.",
"questions": {
"intent": {"type": "choice", "instructions": "What does the customer want?",
"criteria": ["refund", "track delivery", "cancel order"]},
"angry": {"type": "noul", "instructions": "The customer is angry."}
}
}
and gets back structured output:
{
"answers": {
"intent": {"type": "choice", "choice": "refund",
"probabilities": {"refund": 0.9087, "track delivery": 0.0, "cancel order": 0.0913},
"confidence": 0.9087},
"angry": {"type": "noul", "noul": 0.3771}
},
"usage": {"input_tokens": 70, "output_tokens": 0}
}
4. Use the probability
The point of a calibrated answer is that it can be thresholded. Tell the agent what to do at each level, for example in its instructions:
- above 0.9: act on the decision (queue the refund);
- between 0.6 and 0.9: act, but mention the uncertainty;
- below 0.6: ask the customer, or hand over to a person.
With the answer above, "refund" at 0.91 clears the bar and "angry" at 0.38 doesn't, so the agent queues a refund without escalating. The LLM stays in charge of the conversation, and the repeated, measurable decisions go to a model built for them.
Thresholds only work if 0.9 means right nine times in ten, so check that on requests like yours before you rely on them. As shipped, Julia 1 said 96% on an outside benchmark and was right 72% of the time; fitting its temperatures to 50 labelled requests brought its confidence down to its accuracy (how, with the numbers). The MCP server can answer at the fitted calibration:
opendxp calibrate /absolute/path/to/julia-odxp labelled.jsonl --test held-out.jsonl --out julia-cal.json
claude mcp remove julia-1
claude mcp add julia-1 -- opendxp mcp /absolute/path/to/julia-odxp --calibration /absolute/path/to/julia-cal.json
Why a package, not a script
You could wrap Julia 1's own Python code in an MCP server yourself. An OpenDXP package gets you three things that script wouldn't:
- Any runtime can serve it. The same folder runs under
opendxp serveover HTTP, underopendxp mcp, and in our hosted engine (the System One Engine, which is proprietary). In the same server on an A40, Julia 1's package answered 1.6 times the decisions per second of our rebuild of its inference code, both taking one request at a time. On a laptop CPU the speed-up holds for short inputs, not long ones: on an Apple M4, the package was 1.11 times as fast on requests under 500 tokens but 0.63 times overall. - Nothing in it executes. The package holds weights, a template and a calibration. There is no model code to trust.
- It can prove it's still the same model. A package published with a conformance file (the answers the model's
own code gave to 52 fixed requests) can be checked on any machine with
opendxp check.
Browse more decision models at systemonemodels.tech, or read the OpenDXP docs.