laya-zh-eval: Chinese skill routing across all three checkpoints
Twenty real Chinese requests on every checkpoint, measuring accuracy, latency and whether the model's confidence drops when it is wrong.
BUILDS
Agents, routers, scorers and games — built on models that decide. Every entry links to the models it uses, so you can pull the same one and start from there.
10 builds
Twenty real Chinese requests on every checkpoint, measuring accuracy, latency and whether the model's confidence drops when it is wrong.
A local MLX run of Laya against Jev on real support tickets in two languages. Laya answered faster; Jev was far more accurate, 90–100% against 30–80% zero-shot.
A reproducible zero-shot benchmark on informal Darija reviews written in both Arabic script and Arabizi.
A frozen benchmark of synthetic Feishu scenarios with fixed inputs, prompts and labels. Laya answered in 151 ms against Jev's 253 ms but got far fewer right.
Forty questions on automotive, embedded, enterprise software and product decisions, three trials each, published with expected answers, raw responses and an HTML report.
An Arabic checkpoint trained on MASSIVE and XNLI for language understanding, with a companion ranking model for retrieval.
A Myanmar-language checkpoint trained on SIB-200, published with a Gradio demo.
Google's LiteRT community conversion of the multilingual checkpoint, for on-device text classification.
Compares Jev, open-weight Laya, cross-encoders and GPT models as rerankers over a Postgres full-text and pgvector first stage, with cost per 1,000 queries.
GGUF conversions of the English and multilingual models for llama.cpp-style local runtimes.
Many of the first builds here were collected by madewithlaya.com. Every entry links to its creator’s own post, repository or site.