A decision model's product is its probability: it feeds a threshold, routes a ticket or decides whether to escalate to a larger model. Moved from a Mac to an NVIDIA A40, the encoder packages we tested keep their answers, and the llama.cpp ones keep every decision but move probabilities by up to 0.05. Two settings remove most of that drift. Four things can change the answers, or where they are computed, without raising an error.
We tested six OpenDXP packages built from four open models: Julia 1 (Supersonic Labs), three Laya checkpoints (Convai Innovations), Decider 2B (Mark Marosi; our unofficial OpenDXP package, source: github.com/Mapika/decider) and our package of AnyJev L0 (Nokia's method) on Qwen3-1.7B. Each was replayed against the answers its model's own code recorded on an Apple M4: 52 requests per model. A package passes if every decision is the same and every probability is within 0.01.
On an NVIDIA A40
| Package | Runtime on the A40 | Cases passed | Decisions the same | Largest probability difference |
|---|---|---|---|---|
| Julia 1 | onnxruntime, CUDA | 52/52 | 89/89 | 0.0067 |
| Laya (3 checkpoints) | onnxruntime, CUDA | 52/52 | 91/91 | 0.0009-0.0017 |
| Decider (Q8_0) | llama.cpp, CUDA | 43/52 | 91/91 | 0.031 |
| AnyJev (F16) | llama.cpp, CUDA | 45/52 | 90/90 | 0.049 |
| AnyJev (Q8_0) | llama.cpp, CUDA | 36/52 | 88/90 | 0.585 |
The encoder packages pass. The two causal models keep every decision but drift past the tolerance, and it isn't noise: four runs gave identical numbers. Quantized to 8 bits, AnyJev flips 2 of its 90 decisions. A quantized model is a different model, and it needs its own check.
Two settings remove most of the drift
Both GPU runtimes trade precision for speed by default, and both let you turn that off.
ONNX Runtime uses TF32 for float32 matrix products on Ampere and newer NVIDIA GPUs. Setting the CUDA provider's
use_tf32 option to 0 brings every encoder package back to the CPU's numbers:
| Package | Largest difference with TF32 (default) | With TF32 off | Speed with TF32 off |
|---|---|---|---|
| Julia 1 | 0.0067 | 0.00005 | 99% |
| Laya (main) | 0.0014 | 0.00005 | 67% |
llama.cpp adds up the products of F16 and BF16 weights in 16-bit numbers on CUDA. With
GGML_CUDA_CUBLAS_COMPUTE_TYPE=f32 it adds them up in float32, and AnyJev, loaded from its original BF16 weights,
passes: 52 of 52 cases, largest difference 0.0082, at half the speed. The 8-bit models don't come back, because
quantized weights go through llama.cpp's own integer kernels.
OpenDXP 0.3.0 exposes this as one setting, precision: fast (each runtime's default) or exact (float32 on the
GPU). Check the package at the setting you'll serve with:
opendxp check ./anyjev-bf16 --device cuda --precision exact
# cases 52/52 passed, questions 90, max |dp| 8.20e-03, argmax 90/90, errors matched 1/1
On an x86 server CPU
On an Intel Xeon Gold 6342, the encoders matched the Mac to the fifth decimal and AnyJev passed (0.0073). Decider missed by 0.056 and flipped 1 of 91 decisions: its reference answers were recorded with llama.cpp on the Mac's ARM CPU, so the reference itself carries ARM's arithmetic. AnyJev's reference, recorded in float32 with its authors' anyjev package, travels. Record reference answers in float32.
Four changes that raise no error
- A GPU run on the CPU. onnxruntime-gpu 1.30 is built for CUDA 13. On a CUDA 12.8 machine its CUDA provider
can't load
libcublasLt.so.13, so onnxruntime logs one line and runs everything on the CPU. Checksession.get_providers(), and match the build to your CUDA version (on CUDA 12,onnxruntime-gpu==1.23.2). Since 0.3.0,opendxprefuses to run when a provider you named fails to load. - A CPU run on the GPU. A CUDA build of llama.cpp hands large matrix products to the GPU even with no layers
offloaded (its
op_offloaddefault). OpenDXP 0.3.0 turns that off for a CPU run. - A different transformers version. llama.cpp's GGUF converter pins
transformers<5, so installing it into the same environment downgrades transformers to 4.57.6, which computes ModernBERT differently. Julia 1 then passed only 7 of its 52 cases and Laya multilingual 2, with probability errors up to 0.96, on a CPU too. The exporter's own check can't catch this: the graph and the PyTorch model it came from are wrong in the same way. Keep converters in their own environment, and pin the libraries your reference answers depend on. - GPU precision, above. Pick the tolerance before the hardware: a 0.05 change in probability doesn't matter to an argmax, but it does matter to a 0.9 confidence threshold.
Check on every machine
Record reference answers once, then replay them everywhere. OpenDXP packages carry theirs as a conformance file,
and opendxp check replays it in under 2 seconds for Julia 1 on the A40. On an NVIDIA GPU, install the CUDA builds
of the runtimes first:
pip install opendxp "onnxruntime-gpu==1.23.*" # encoder packages on CUDA 12 (the [onnx] extra installs the CPU build)
CMAKE_ARGS="-DGGML_CUDA=on" pip install llama-cpp-python # GGUF packages: llama.cpp built for CUDA
opendxp check ./package --device cuda
Related: the OpenDXP docs, What is an AI decision model?.