Batching, running several requests through a model in one pass, is supposed to be free throughput. For decision models it made an Apple M4's CPU slower. On an NVIDIA A40 it paid off once each pass held only rows of similar length: then Julia 1's package answered 310 decisions per second at peak, against 192 without batching. For causal models, batching a request's own rows helps, and batching requests together doesn't.
On a CPU, don't batch
On an Apple M4 (8 threads), Julia 1 (Supersonic Labs), a ModernBERT decision model run through our rebuild of its inference code, answered its 52 test requests in 2.13 seconds one at a time and in 4.72 seconds batched: 0.45× the speed. For our package of AnyJev L0 (Nokia's method) on Qwen3-1.7B, reading eight rows together was 2.1× slower (13.6 seconds a request on average, against 6.5). A single request already keeps every core busy, so a batch adds no parallelism. It only adds padding, because every sequence in a batch is padded to the longest one.
On a GPU, group rows by length
On an A40, a pass of Julia 1's OpenDXP package takes about 6 ms whether it holds 1 short row or 16, so batching short rows makes each one up to 15× cheaper. A pass filled up to a token budget throws that away. When a 1,689-token row joins the pass, every row in it is padded to 1,689 tokens, and attention cost grows with the square of the padded length: a 100-token row then costs about 285 times as much attention as it needs.
The rule that works: a pass holds only rows within 1.25× of the shortest row's length (plus 16 tokens), and at most 32 rows. Past 32 rows, a pass got slower per row on this GPU (0.44 ms per row at 32, 0.65 ms at 64).
| All 89 of Julia 1's test rows on the A40 | Time |
|---|---|
| One at a time | 625-630 ms |
| Batched by token budget | 666 ms |
| Batched by similar length | 243 ms |
Through our hosted engine (the System One Engine, which is proprietary) on the A40, with 16 clients calling:
| Julia 1 as an OpenDXP package | Decisions/s at 16 clients | p50 latency at 16 clients | Peak decisions/s (clients) |
|---|---|---|---|
| Batching off | 192 | 117 ms | 192 (16) |
| Batched by similar length | 279 | 69 ms | 310 (128) |
Causal models: batch a request's rows, not requests
Decider 2B (Mark Marosi; our unofficial OpenDXP package, source: github.com/Mapika/decider) and AnyJev read a prompt of a few hundred tokens for each question. That's already a large matrix: one prompt keeps the GPU's matrix units busy. Batching requests together changed nothing we could measure (Decider: 33.0 against 33.4 decisions per second).
Reading a single request's rows together did help. A request with several questions, or one question asked in several option orders, produces prompts short enough to share a pass. In OpenDXP's reference runtime on the A40, one request at a time:
| Package on the A40 | Rows per pass | Speed-up | Decisions unchanged |
|---|---|---|---|
| Decider | 16, shared cache | 1.38× | 91 / 91 |
| AnyJev | 4, per-row cache | 1.21× | 89 / 90 |
| AnyJev | 4, shared cache | 1.10× | 90 / 90 |
The per-row cache is faster for AnyJev, but it flips one decision. The shared cache keeps all 90. That's why this setting is chosen per machine, with a check, and isn't baked into the model.
What to do
- Measure on the device you serve on. Batching helped a GPU and hurt a CPU with the same code.
- Sort by length and group by it. A token budget alone lets one long input slow down every short one.
- Check the answers, not only the speed. Every batched configuration above was checked against the model's own recorded answers.
opendxp serve answers one request at a time per model, so it doesn't batch requests from many clients. For a
causal package such as Decider or AnyJev, OpenDXP's reference runtime reads a request's rows together with
--batch-rows:
CMAKE_ARGS="-DGGML_CUDA=on" pip install llama-cpp-python # llama.cpp built for CUDA
pip install opendxp
opendxp bench ./package --device cuda --batch-rows 16 # how fast
opendxp check ./package --device cuda --batch-rows 16 # and whether it's still the same model
Related: What a decision costs.