Hanno Labs' small calibrated decision model. A LoRA on Qwen3-1.7B plus trained decision-token embeddings with stable slots, returning the full distribution over up to 255 caller-defined choices and a null slot for Choice, Score and Noul questions.
Respan's behaviour-scoring model for evals, guardrails and monitoring. For each plain-language behaviour you define, it reads a conversation or agent trace and returns the probability the behaviour is present, absent or not observable, in one forward pass.
Decides
choice, score, noul, classify, route
noul, classify
Architecture
bosun
span
Fine-tuned from
qwen/qwen3-1.7b
—
License
apache-2.0
proprietary
Availability
Open weights
Hosted API
Hosted by
—
Respan
Input price
—
$0.020/MTok
Decision accuracy
84.9%
—
Calibration error
0.050
—
Valid action rate
—
—
Median latency
—
—
p95 latency
—
—
Evaluation suite
Figures are from each model’s manifest; accuracy and latency are what the publishers report, on their own suites and hardware. Add a third model.
Questions
What is the difference between bosun and span-01?
bosun is from Hanno Labs and span-01 from Respan. bosun has open weights you can download and run; span-01 is only available as a hosted API. Both answer noul and classify questions. Only bosun answers choice, score and route. bosun is licensed apache-2.0; span-01, proprietary.
Which is more accurate, bosun or span-01?
Only bosun publishes an accuracy figure (84.9% on DecisionBench (Hanno-Labs/decision-bench, 23,900 rows; seen task families)); span-01 does not, so there is no comparison to make without your own test.
Which is cheaper, bosun or span-01?
bosun: Free (open weights). span-01: $0.02 / $0 per 1M. Open weights cost nothing per call beyond your own hardware.
Can I run bosun or span-01 locally?
bosun yes — systemone pull hanno-labs/bosun downloads its weights. The other is only served as a hosted API.
DecisionBench (Hanno-Labs/decision-bench, 23,900 rows; seen task families)