Metask Lab's calibrated decision model in 16 languages. Qwen3.5-4B with a merged LoRA trained with the Nimble candidate-logit objective; one forward pass and a softmax over at most 26 answer-letter logits, with one temperature per question type.
Respan's behaviour-scoring model for evals, guardrails and monitoring. For each plain-language behaviour you define, it reads a conversation or agent trace and returns the probability the behaviour is present, absent or not observable, in one forward pass.
Decides
choice, score, noul, classify, route
noul, classify
Architecture
metask-jev
span
Fine-tuned from
qwen/qwen3.5-4b
—
License
apache-2.0
proprietary
Availability
Open weights
Hosted API
Hosted by
—
Respan
Input price
—
$0.020/MTok
Decision accuracy
80.1%
—
Calibration error
—
—
Valid action rate
—
—
Median latency
62.8 ms
—
p95 latency
—
—
Figures are from each model’s manifest; accuracy and latency are what the publishers report, on their own suites and hardware. Add a third model.
Questions
What is the difference between metask-jev and span-01?
metask-jev is from Metask Lab and span-01 from Respan. metask-jev has open weights you can download and run; span-01 is only available as a hosted API. Both answer noul and classify questions. Only metask-jev answers choice, score and route. metask-jev is licensed apache-2.0; span-01, proprietary.
Which is more accurate, metask-jev or span-01?
Only metask-jev publishes an accuracy figure (80.1% on JevBench v1.2 public set (231 items) at 4,096 tokens, maker's run); span-01 does not, so there is no comparison to make without your own test.
Which is cheaper, metask-jev or span-01?
metask-jev: Free (open weights). span-01: $0.02 / $0 per 1M. Open weights cost nothing per call beyond your own hardware.
Can I run metask-jev or span-01 locally?
metask-jev yes — systemone pull metask-lab/metask-jev downloads its weights. The other is only served as a hosted API.
Evaluation suite
JevBench v1.2 public set (231 items) at 4,096 tokens, maker's run