Metask-Jev-4B
Metask Lab's calibrated decision model in 16 languages. Qwen3.5-4B with a merged LoRA trained with the Nimble candidate-logit objective; one forward pass and a softmax over at most 26 answer-letter logits, with one temperature per question type.
Trained on 60.9k decisions: human-labelled public data, 16.1k MASSIVE routing items in 14 locales, and 390 synthetic policy items that deliberately mirror JevBench's hard-tier families (disclosed). The maker reports 80.1% on JevBench v1.2's 231 public items at 4,096 tokens, a 78.9% macro on 13 human-labelled subsets, 85.9% on held-out MASSIVE in 14 locales, and ECE 0.114 falling to 0.040 after temperature fitting. Its claim that it would rank first is its own estimate: the official JevBench v1.4.2 run placed it 12th (79.7% public, 27.6% sealed). Validated at 4,096 tokens; the base supports 262,144.
What it decides
- choice — picks one option from a set
- score — places the input on an ordered scale
- noul — answers a yes/no question with one calibrated probability
- classify — assigns a category from a fixed taxonomy
- route — sends the input to one of several destinations
At a glance
| Parameters | 4.5B |
| Base model | qwen/qwen3.5-4b |
| Maker | Metask Lab |
| Released | 2026-09-21 |
| License | apache-2.0 |
| Reported accuracy | 80.1% |
| Reported latency | 62.8 ms p50 on an RTX 4090 |
Get the weights
pip install systemonemodels
systemone pull metask-lab/metask-jev
The files are served from the maker's Hugging Face repository, wayfind/metask-jev-4b-policy-mix, and verified against the checksums recorded here.