On one rented NVIDIA A40 at $0.51 an hour, a million decisions cost $0.46 with Julia 1 and $11.79 with our package of AnyJev L0 (Nokia's method) on Qwen3-1.7B. A decision is one question answered, such as "which of these five teams should get this ticket?", with a calibrated probability for every option. Each model ran in our hosted engine (the System One Engine, which is proprietary), with 1 to 128 clients calling at once.
One model on one GPU
| Model | Served as | One client: decisions/s | Latency (p50) | Peak decisions/s | $ per 1M decisions |
|---|---|---|---|---|---|
| Julia 1 | OpenDXP package (ONNX Runtime) | 155 | 7.5 ms | 310 | 0.46 |
| Julia 1 | our rebuild of its inference code (PyTorch) | 104 | 15.6 ms | 164 | 0.86 |
| Laya | OpenDXP package (ONNX Runtime) | 90 | 12.5 ms | 147 | 0.96 |
| Laya | the laya package (its own code) | 32 | 32.9 ms | 70 | 2.02 |
| Decider | our rebuild of its own code (llama.cpp) | 34 | 35.5 ms | 35 | 4.08 |
| Decider | OpenDXP package (llama.cpp) | 30 | 33.5 ms | 33 | 4.25 |
| AnyJev | OpenDXP package (llama.cpp) | 11.6 | 69 ms | 12 | 11.79 |
Julia 1 (Supersonic Labs) and Laya (Convai Innovations) are ModernBERT encoders. Decider 2B (Mark Marosi; our unofficial OpenDXP package, source: github.com/Mapika/decider) is a causal model with 8-bit weights. Cost per million decisions is the hourly price divided by the decisions answered in an hour at peak, with batching on wherever it helps.
Many models on one GPU
Six packages (Julia 1, three Laya checkpoints, Decider and AnyJev) take about a third of the A40's memory, 16 of 45 GiB. Memory isn't the limit. Mixing light and heavy models costs the light ones most of their speed, because the heavy ones keep the GPU busy, so group models by weight:
| One A40 holding | Decisions/s, all models | Decisions per day |
|---|---|---|
| Julia 1 alone | 310 | 26.8 million |
| four encoders | 184-195 | 16-17 million |
| Decider and AnyJev | 19-22 | 1.7-1.9 million |
| all six | 68-80 | 5.8-6.9 million |
Per day assumes the load never stops. Real traffic has peaks, so plan for a third of it: about 5 million decisions a day from one $0.51-an-hour GPU running four encoder models, which is 5,000 users making 1,000 decisions each. More in Six models on one GPU.
What it means
The package can be faster than the model's own code. In the same server on an A40, Julia 1's package answered 1.6 times the decisions per second of our rebuild of its inference code, both taking one request at a time (192 against 117). Laya's package answered 1.8 times the laya package one request at a time, and 2.1 times at peak. Every decision stays the same, and every probability within 0.007 of the reference answers. An OpenDXP package is data, so any runtime that speaks the standard can run it, including a faster one. On a laptop CPU the speed-up holds for short inputs, not long ones: on an Apple M4, Julia 1's package was 1.11 times as fast on requests under 500 tokens, but 0.63 times overall.
Causal models pay for every token. Decider and AnyJev read their whole prompt for every question, and AnyJev reads it again for each rotation of the options. That removes position bias, and it's why AnyJev costs about 25 times as much as Julia 1 per decision. A prompt cache gives half of it back.
Per token, it's already cheaper than any API we found. At full load Julia 1 costs about $0.003 per million input tokens and Laya $0.008. Hosted embedding APIs list $0.005 to $0.02 per million input tokens, and small LLM APIs $0.02 to $1.00 (prices on 30 September 2026). Decider and AnyJev, at $0.03 to $0.05, sit around the cheapest LLM APIs.
Speed means nothing without exactness. On the A40, the encoder packages keep every decision of the model's own
code, with every probability within 0.01. The llama.cpp packages keep every decision but move probabilities by up
to 0.05. Float32 arithmetic (--precision exact) brings AnyJev, loaded from its BF16 weights, back within 0.01 at
about half the speed. Why.
Measure your own
pip install opendxp "onnxruntime-gpu==1.23.*" # encoder packages on CUDA 12 (the [onnx] extra installs the CPU build)
CMAKE_ARGS="-DGGML_CUDA=on" pip install llama-cpp-python # GGUF packages: llama.cpp built for CUDA
opendxp bench ./package --device cuda --usd-per-hour 0.51 # the package, in this process
opendxp serve ./package --device cuda & # an OpenDXP server on port 8790
opendxp bench http://127.0.0.1:8790 --model supersonic-labs/julia-1 \
--concurrency 1,4,16,64 --usd-per-hour 0.51 # any OpenDXP server; --model is a name /v1/models lists
opendxp check ./package --device cuda # and prove it's still the same model
opendxp serve has no GPU batching, so expect its numbers to sit near the one-request-at-a-time figures above.
Or skip the GPU: Julia 1, Laya and Decider can be called through the System One Models API.