Laya benchmarks
The official Laya README publishes detailed numbers, and a growing set of independent repositories now test those claims. The overall picture: Laya is much faster when run locally and free to self-host, while Jev is more accurate out of the box on most shared tests.
228 projects
jev-kev-laya-selfhost
suarify
Provides a Dockerized, self-hosted HTTP API for Laya typed decisions with agent-skill documentation
laya-vs-dgpl-arena
vk-alto-none
Stages AI decision models in a game arena, including local Laya and hosted Jev
laya-zeroshot-agent-guard-eval
Zephyr4772
An evaluation measures zero-shot Laya checkpoints as guards between agent steps
Smart-Support-Ticket-Router-Using-Laya
affan1311
Routes support tickets with Laya and uses Gemini to draft replies, with a benchmark against a Gemini-only pipeline
Smart-Support-Ticket-Classifier-Using-Laya
affanhyder-diggit
Routes support tickets with Laya and benchmarks the results against a Gemini-only pipeline
privatemode-decisions-benchmark
edgelesssys
A benchmark compares speed, accuracy, and cost for Jev, Laya, and Privatemode decisions
jev-equivalent-research
hjl1045
Evaluates Jev, Laya, and a GPT-5.6 Luna baseline on synthetic auto-claims classification
MacJev-322M-4K-Laya-MLX
lawrence3699
Runs the MacJev typed-decision model natively on Apple silicon with MLX and compares it with Laya
jev-bench
brandonrc
Benchmarks Jev, Laya, Claude Haiku, and CLM on package-registry triage tasks
decision-lab
Amine-LG
Provides a visual playground to build decisions and compare Jev, OpenJEV, and local Laya
KVev
AdudodlaVarish
Serves cached long-document chat and uses Laya for offline routing experiments
decision-model-arena
sathik11
Compares Jev, Laya, and Microsoft Foundry for typed decisions in enterprise incident triage
tool-output-pruning-lab
Dymyt-ry
Benchmarks Laya and other selectors for pruning tool output in coding-agent sessions
kime-compat
tamnd
Checks whether kime supports TypeSafe and Laya APIs, clients, SDKs, and tests
narde
anliang0306
Reimplements the Laya decision engine and verifies behavioral parity against upstream
jev-agent-risk-gate
jayeshvpatil
Compares Jev, Claude, fine-tuned Qwen, and open Laya in experiments on agent shell-command risk gates
d3code-calibration
gkastanis
Compares Jev and open-weights Laya probability calibration against human ratings
jev-context-rot
casperkwok
Reproduces an experiment comparing context sensitivity in Jev, Laya and Decider models
jev-vs-laya
Cognition-Forge
Compares Jev and Laya variants across typed-decision tasks
Laya-Finetune
zamax14
Fine-tunes and evaluates multilingual Laya for support-ticket classification
laya-games
BM-Popcorn
Demonstrates Laya-driven turn-based games and fine-tunes the model on engine-labeled positions
laya_bio
maris205
Adapts the Laya encoder into a shared candidate-scoring model for biological sequence tasks
laya_demo
itsvrushabh
Demonstrates and benchmarks Laya’s typed decision engine across interactive use cases
laya-deploy
iintothewind
Deploys and benchmarks Laya Serve on Windows with Docker and a temporary public tunnel
Official numbers
The core README and its BENCHMARKS.md report latency on a Tesla T4 (32.8 ms for one question with laya-multilingual, 39.5 ms with laya) and accuracy across 51 languages. Its Jev figures are third-party published, not measured by the Laya authors.
Independent head-to-heads
- sysone-bench: byte-identical inputs across 751 states. Jev leads on triage, guardrails, moderation, Banking77 and multilingual intent. Laya leads on AG News and MNLI.
- jev-laya-benchmark: 1,470 synthetic items. Jev 92.9% vs Laya 65.3% of judgments correct. For a single question, Laya was 42 ms locally on an M3 Pro and Jev was 136 ms. Jev was faster above 3 to 4 questions per request.
- laya-jev-lab: 40 Chinese support tickets. Jev 78% at 588 ms, Laya 57% at 7.6 ms on an M4 Max. A cascade that escalates below 0.60 confidence matched Jev's accuracy at 1.8x its speed.
- JevBench: a composite board. In v1.4.2, Laya (English, CPU) ranks 41st with a score of 30.25 and Jev 1.13.0 ranks 2nd with 63.29.
- laya-mcp eval: key accuracy 32.5% for Laya vs 98.8% for Jev, with median latency of 26 ms vs 138 ms.
CPU latency
laya-cpu-benchmark measured laya-multilingual in steady state on an Intel i9-14900HX. One question took 50 ms for a 71-token state, 307 ms for 261 tokens, and 1,315 ms for 927 tokens. Batching gave little gain on CPU. Its conclusion: a realistic CPU call is 0.3 to 2 seconds, not 33 ms.
Run your own
git clone https://github.com/glukicov/laya_router && cd laya_router
uv sync --all-extras
uv run laya-router eval run --backend laya --device mps
Or use jevbench (dhruvmehra) to compare Jev, Laya, LLMs and BERT models on public datasets.
Reading these numbers
Most of these suites are small (40 to 1,500 items), synthetic or in one domain, and their authors say so. Latency depends on hardware, state length and question count. The core README's own advice: treat Laya as a fast base to specialise, and fine-tune for accuracy.
More ways to use Laya