Six decision-model packages take 16 GiB of an NVIDIA A40's 45 GiB, and served together they answer up to 80 decisions per second. How well they share the GPU depends on two things: whether a slow model can hold up the fast ones, and which models you put together. We measured both in our hosted engine (the System One Engine, which is proprietary).
The six are OpenDXP packages of four open models. Julia 1 (Supersonic Labs) and three Laya checkpoints (Convai Innovations) are ModernBERT encoders that take 6 to 15 ms a request on their own. Decider 2B (Mark Marosi; our unofficial OpenDXP package, source: github.com/Mapika/decider) and our package of AnyJev L0 (Nokia's method) on Qwen3-1.7B are causal models of 2B and 1.7B parameters that take 30 to 110 ms. All six were called at once, with 1, 4 or 16 clients per model.
Memory is not the limit
All six took 16 GiB of the A40's 45 GiB once loaded, and up to 21 GiB under load. The four encoders alone took 6.4 GiB. On memory alone, this GPU could hold about two copies of the set.
Don't let a slow model hold up the fast ones
With one serving thread per GPU, the models take turns. Every model got the same 7 decisions a second, 44 in all, and the fast ones waited behind the slow ones' prompts:
| 1 client per model, one thread per GPU | Decisions/s | p50 latency | Alone on the GPU |
|---|---|---|---|
| Julia 1 | 7.2 | 129 ms | 8 ms |
| Laya (main) | 7.3 | 124 ms | 12 ms |
| Decider | 7.3 | 128 ms | 34 ms |
| AnyJev | 7.3 | 123 ms | 69 ms |
Giving each model a serving thread of its own, with up to four models computing on the GPU at once, raised the total:
| All six at once | 1 client each | 4 clients each | 16 clients each |
|---|---|---|---|
| One thread per GPU | 43.9 | 43.3 | 45.1 decisions/s |
| A thread per model, four at once | 68.8 | 67.6 | 80.3 decisions/s |
That's 56 to 78 % more in total. Julia 1 doubled (7.2 to 14.5 decisions per second at one client), and its p50 dropped from 129 to 94 ms. The heavy models slowed a little, because the GPU is shared more fairly: Decider's p50 went from 128 to 174 ms. The four encoders on their own gained too: 135 to 161 decisions per second with one thread, 184 to 195 with four.
Keep light and heavy models apart
Even with concurrency, the light models ran at a tenth of their solo speed next to the causal ones, whose prompts keep the GPU full. Mixing loads is fine for memory and bad for latency, so plan for separate GPUs:
| One A40 holding | Decisions/s in all | Decisions per day at that rate |
|---|---|---|
| four encoders | 184-195 | 16-17 million |
| Decider and AnyJev | 19-22 | 1.7-1.9 million |
| all six | 68-80 | 5.8-6.9 million |
Steady over half an hour
Serving the four encoders for 30 minutes, 4 clients each, the server handled 199,748 requests at 192.8 decisions per second on average, about 347,000 decisions. It answered all but 2,436 of them, the requests Julia 1 refuses by design (3.85 % of its requests, as its conformance file expects), and had no other errors. GPU memory averaged 9,560 MiB over the first five minutes and 9,556 MiB over the last five, and the server's resident memory grew by 4 MiB. Each model's 99th-percentile latency was between 347 and 470 ms.
Try it
You can load-test any OpenDXP server the same way, with one client process per model. opendxp serve holds
several packages on port 8790 and names each by its package name. It answers one request at a time per model, with
no GPU batching and no limit on how many models compute at once, so expect different numbers from ours:
opendxp serve ./julia-1 ./anyjev --device cuda &
curl -s http://127.0.0.1:8790/v1/models # the name of each package the server holds
opendxp bench http://127.0.0.1:8790 --model supersonic-labs/julia-1 --concurrency 4 --duration 30 &
opendxp bench http://127.0.0.1:8790 --model "$ANYJEV" --concurrency 4 --duration 30 &
wait
Set ANYJEV to the name /v1/models lists for the AnyJev package: a package is served under the name in its
odxp.json.
Related: What a decision costs.