A rank-16 LoRA on Qwen3.5-4B built for the JevBench setting. A document, a policy and a question go in; a calibrated distribution over the option letters comes out of one forward pass. Research and demo use only.
The first System One model. Reads a state, answers typed Choice, Score and Noul questions in one call with calibrated probabilities, and generates no text. Closed weights, served by TypeSafe AI.
Decides
choice, score, noul, classify
choice, score, noul, classify, route
Architecture
hopper
jev
Fine-tuned from
qwen/qwen3.5-4b
—
License
Research and demo use only (training data includes RACE, non-commercial); serving code Apache-2.0
proprietary
Availability
Open weights
Hosted API
Hosted by
—
TypeSafe AI
Input price
—
$0.042/MTok
Decision accuracy
68.5%
—
Calibration error
0.102
—
Valid action rate
—
—
Median latency
—
—
p95 latency
—
—
Figures are from each model’s manifest; accuracy and latency are what the publishers report, on their own suites and hardware. Add a third model.
Questions
What is the difference between hopper and jev?
hopper is from HopitAI and jev from TypeSafe AI. hopper has open weights you can download and run; jev is only available as a hosted API. Both answer choice, score, noul and classify questions. Only jev answers route. hopper is licensed other; jev, proprietary.
Which is more accurate, hopper or jev?
Only hopper publishes an accuracy figure (68.5% on JevBench public hard tier (111 items, measured by the authors)); jev does not, so there is no comparison to make without your own test.
Which is cheaper, hopper or jev?
hopper: Free (open weights). jev: $0.042 / $0 per 1M. Open weights cost nothing per call beyond your own hardware.
Can I run hopper or jev locally?
hopper yes — systemone pull hopit-ai/hopper downloads its weights. The other is only served as a hosted API.
Evaluation suite
JevBench public hard tier (111 items, measured by the authors)