jevbench: six classifiers compared, calibration included
Jev against GPT-5-mini, Claude Sonnet 5, fine-tuned DistilBERT, BART NLI and Laya on SST-2, AG News and more, reporting accuracy, macro-F1, ECE, latency, throughput and cost.
BUILDS
Agents, routers, scorers and games — built on models that decide. Every entry links to the models it uses, so you can pull the same one and start from there.
9 builds
Jev against GPT-5-mini, Claude Sonnet 5, fine-tuned DistilBERT, BART NLI and Laya on SST-2, AG News and more, reporting accuracy, macro-F1, ECE, latency, throughput and cost.
Abhijay stress-tested Jev, SemIf and Laya on adapted exam questions: Jev scored 83.7%, SemIf 61.6% and Laya 31.2%.
Describe a past Claude Code, Codex or OpenCode session in plain words; Chat Seek searches the local histories and reranks the matches with Laya.
Collects vacancies from jobnet.dk and jobindex.dk, filters them against a candidate profile and ranks the matches locally.
Leonardo Stenico moved the Needle extension from hosted Jev to local Laya, so finding passages by meaning needs no API key.
Astrid's Postgres extension, written in Rust on Candle, runs the model in the database so a typed decision is a SQL call.
GLiNER, GLiClass, Laya, Von, Jev and more in one web UI, with demos, benchmarks, model sizes and licences.
A frozen protocol on four datasets with raw predictions, calibration and latency: Laya ties Jev on one and trails by 10–33 points on the others.
Shows across 400 documents that the decision head's attention falls off after roughly 200 tokens, then provides a chunking harness that restores accuracy on long inputs.
Many of the first builds here were collected by madewithlaya.com. Every entry links to its creator’s own post, repository or site.