Hanno Labs' small calibrated decision model. A LoRA on Qwen3-1.7B plus trained decision-token embeddings with stable slots, returning the full distribution over up to 255 caller-defined choices and a null slot for Choice, Score and Noul questions.
An embedding model turned decision model. Qwen3-Embedding-4B with a merged LoRA, driven by the open JevEmbed toolkit, which scores Choice, Score and Noul questions from embedding similarity and returns a distribution without generating text.
Decides
choice, score, noul, classify, route
choice, score, noul, classify
Architecture
bosun
jevembed
Fine-tuned from
qwen/qwen3-1.7b
qwen/qwen3-embedding-4b
License
apache-2.0
apache-2.0
Availability
Open weights
Open weights
Hosted by
—
—
Input price
—
—
Decision accuracy
84.9%
85.9%
Calibration error
0.050
—
Valid action rate
—
—
Median latency
—
—
p95 latency
—
—
Figures are from each model’s manifest; accuracy and latency are what the publishers report, on their own suites and hardware. Add a third model.
Questions
What is the difference between bosun and jevembed?
bosun is from Hanno Labs and jevembed from HIT-TMG (Lychee Team). Both have open weights you can download and run. Both answer choice, score, noul and classify questions. Only bosun answers route. bosun is the smaller model, at 2.0B parameters to 4.0B.
Which is more accurate, bosun or jevembed?
They report on different suites — bosun 84.9% on DecisionBench (Hanno-Labs/decision-bench, 23,900 rows; seen task families), jevembed 85.9% on JevEmbed-Data test split (64,110 hard-label questions), final LoRA checkpoint before merging — so the numbers do not rank them. Test both on your own labelled examples.
Which is cheaper, bosun or jevembed?
bosun: Free (open weights). jevembed: Free (open weights). Open weights cost nothing per call beyond your own hardware.
Can I run bosun or jevembed locally?
Yes, both: systemone pull hanno-labs/bosun and systemone pull hit-tmg/jevembed download the weights.
Evaluation suite
DecisionBench (Hanno-Labs/decision-bench, 23,900 rows; seen task families)
JevEmbed-Data test split (64,110 hard-label questions), final LoRA checkpoint before merging