On an NVIDIA A40, 4-bit weights changed 14 of a decision model's 90 decisions, 8-bit weights changed 2, and neither made it faster. The model was our package of AnyJev L0 (Nokia's method) on Qwen3-1.7B, whose Q4_K_M weights take 1.3 GB against 4.1 GB at F16. A decision model's output is a probability that something acts on, such as a router, a threshold or an escalation, so a quantized variant has to be checked before it replaces the original.
Results
AnyJev asks a causal language model a multiple-choice question and reads the letter probabilities once for every rotation of the options, so option order doesn't bias the answer. We ran the package in four precisions on the A40 with llama.cpp, one request at a time, and compared every answer with the ones AnyJev's own code recorded: 52 requests, 90 questions.
| Weights | Decisions the same as the model's own | Largest probability change | Mean change | Decisions per second |
|---|---|---|---|---|
| F16 | 90 of 90 | 0.049 | 0.003 | 10.0 |
| BF16 | 90 of 90 | 0.101 | 0.004 | 10.5 |
| Q8_0 | 88 of 90 | 0.585 | 0.028 | 10.3 |
| Q4_K_M | 76 of 90 | 0.999 | 0.153 | 10.2 |
At 4 bits, the model changes its answer on 14 of the 90 questions. The mean probability moves by 0.15, so a question it used to answer at 0.8 might now come out at 0.65 or 0.95. Even at 8 bits, two decisions flip.
Why it wasn't faster
On a GPU, quantization mostly saves memory bandwidth, which is what limits text generation: every new token reads the whole model again. A decision model doesn't generate. It reads each prompt of a few hundred tokens in a single pass and takes the next-token probabilities. That pass is limited by compute, not memory, so on this GPU smaller weights didn't make it faster: 10.0 to 10.5 decisions per second for every precision. This is one model on one A40; on a CPU, a phone or another GPU the balance can differ.
Quantization still has a use: a 4-bit model fits on a smaller GPU, or on a phone. But then you're running a different model, and you need to know how different.
Measure it before you ship it
Every OpenDXP package carries a conformance file: the answers the original code gave to a fixed set of requests. Any variant, whether quantized, converted or moved to new hardware, can be checked against it in seconds:
CMAKE_ARGS="-DGGML_CUDA=on" pip install llama-cpp-python # llama.cpp built for CUDA
pip install opendxp
opendxp check ./anyjev-q4_k_m --device cuda
# cases 31/52 passed, questions 90, max |dp| 9.99e-01, argmax 76/90, errors matched 1/1
For this package on NVIDIA GPUs, only the BF16 weights with float32 arithmetic (--precision exact) pass the
check: 52 of 52 cases, largest difference 0.0082. F16 passes on CPUs but misses one case on CUDA even at exact
precision. Q8_0 flips 2 of 90 decisions, so we'd offer it only as a variant under its own name, with its own
check, for machines that can't hold the 4 GB file. We wouldn't ship Q4_K_M as the same model.
Related: What a decision costs, Same model, different answers?.