All models

COMPARE

3 models, side by side.

What each one decides, how well calibrated it is, how fast it answers and what it costs to pull. Up to 4 at a time; the better value in each row is marked.

Propertybespoke-labs/bespoke-nimble-9balibi-serikbay/jevk5wfzyx/von
SummaryAn open Jev-style LoRA on Qwen3.5-9B from Bespoke Labs, trained on 2,676 contrastively curated examples to score the allowed answer tokens directly for enums, booleans and rubric levels. Recipe, data and a public benchmark suite are released with it.Qwen3.5-4B with a merged rank-16 LoRA distilled from 17,408 teacher questions and 30,000 public training rows, read out as a temperature-scaled softmax over answer-letter logits. Also in 9B, 2B and GGUF.A ModernBERT-large encoder with an option-marker head. Premise and options are packed into one sequence and each option's marker is scored in a single bidirectional pass; version 1.2 makes the scoring order-invariant.
Decideschoice, noul, score, classify, routechoice, score, noul, classify, routechoice, score, noul, classify
Architecturenimblejevk5von
Fine-tuned fromqwen/qwen3.5-9bqwen/qwen3.5-4banswerdotai/modernbert-large
Licenseapache-2.0apache-2.0apache-2.0
AvailabilityOpen weights + hosted APIOpen weightsOpen weights
Hosted byBespoke Labs——
Input price———
Decision accuracy90.1%78.4%63.9%
Calibration error0.0540.0350.045
Valid action rate———
Median latency106 ms13.2 ms18 ms
p95 latency———
Evaluation suiteBespoke held-out set (324 examples)JevBench public hard tierJevBench public standard tier
Latest version2026.090.3.01.2.0
VariantsSHA256SUMSSHA256SUMS—
Size of latest version184.3 MB7.9 GB2.9 GB
Files1082
Downloads000
Stars000
Tagssystem-one, qwen, lora, curated-data, 9bsystem-one, qwen, lora, distilled, 4bsystem-one, encoder, modernbert, 395m
UpdatedSep 25, 2026Sep 25, 2026Sep 25, 2026