Ask a language model "Which team should handle this: A) billing, B) shipping, C) returns?" and part of the answer depends on the letters: move "returns" from C to A and its probability changes. Asking once per rotation of the options removes that bias, but it multiplies the compute. On an NVIDIA A40, a prompt cache gives half of it back, decoding only a third of the tokens, with every decision unchanged.
What rotations cost
AnyJev, Nokia's method for asking an instruction-tuned language model a typed question, asks each question once per rotation of its options, so every option appears in every position, and combines the answers. What's left is the model's view of the options, not of the letters.
For our package of AnyJev L0 (Nokia's method) on Qwen3-1.7B, on the 52 test requests (90 questions) of the OpenDXP request set:
- 286 prompts for 90 questions, about 3.2 per question: one for every rotation of each choice and yes/no question, and one for each score question, which this package doesn't rotate.
- 10 decisions per second on the A40, reading one prompt at a time; 0.28 on an Apple M4's CPU.
- $11.79 per million decisions at $0.51 an hour through our hosted engine (the System One Engine, which is proprietary), against $0.46 for Julia 1 (Supersonic Labs), a ModernBERT decision model that scores every option in one pass.
A prompt cache halves it
The rotations of one question share almost everything: the chat template, the ticket text and the question. They differ only in the order of the options, near the end. A prompt cache keeps the last prompt in the model's memory and decodes only the tokens where the next prompt differs.
| AnyJev on the A40 | Time for the 51 answered cases | Tokens decoded | Decisions unchanged | Largest probability change |
|---|---|---|---|---|
| Every prompt read in full | 9.2 s | 71,399 | 90 / 90 | 0.049 |
| With a prompt cache | 4.5 s | 22,915 (32%) | 90 / 90 | 0.060 |
That's twice as fast, with every decision the same. The probabilities move slightly more, by up to 0.060 against 0.049, because the numbers a GPU computes depend a little on how the work is split up.
A prompt cache needs a model whose cache can drop the tokens after a given position. For Decider 2B (Mark Marosi; our unofficial OpenDXP package, source: github.com/Mapika/decider), llama.cpp couldn't cut the cache back to a shared prefix, so the cache saved nothing.
Reading rotations together helps less
A request's rotations are independent prompts, so a GPU can read several of them in one pass. OpenDXP calls this
batch_rows. In OpenDXP's reference runtime on the A40, one request at a time:
| Rows per pass | Cache | Decisions per second | Speed-up | Decisions unchanged | Largest probability change |
|---|---|---|---|---|---|
| 1 | 10.0 | 1.00x | 90 / 90 | 0.049 | |
| 4 | one shared | 11.0 | 1.10x | 90 / 90 | 0.044 |
| 4 | one per row | 12.1 | 1.21x | 89 / 90 | 0.094 |
| 16 | one shared | 10.7 | 1.07x | 90 / 90 | 0.068 |
That's a modest gain. Each prompt is already a few hundred tokens, which keeps the GPU's matrix units busy.
When to pay for rotations
Rotations are worth it when the options' order carries no meaning and you can't afford the bias: routing, triage, anything with a threshold. They cost several times the compute, and a prompt cache recovers about half of that. Models built to score every option in one pass, like Julia 1 and Laya (Convai Innovations), don't pay for rotations, and on the same GPU they were 12 to 26 times cheaper per decision. We didn't measure their position bias.
Check it yourself
Every setting above is checked the same way, against the answers the model's own code recorded:
CMAKE_ARGS="-DGGML_CUDA=on" pip install llama-cpp-python # llama.cpp built for CUDA
pip install opendxp
opendxp check ./anyjev --device cuda --batch-rows 4
The prompt cache is a script in the OpenDXP repository.
Related: What a decision costs, Batching made our CPU slower.