Laya benchmarks

The official Laya README publishes detailed numbers, and a growing set of independent repositories now test those claims. The overall picture: Laya is much faster when run locally and free to self-host, while Jev is more accurate out of the box on most shared tests.

228 projects

jev-kev-laya-selfhost

suarify

Provides a Dockerized, self-hosted HTTP API for Laya typed decisions with agent-skill documentation

0·GitHub repo

laya-vs-dgpl-arena

vk-alto-none

Stages AI decision models in a game arena, including local Laya and hosted Jev

0·GitHub repo

laya-zeroshot-agent-guard-eval

Zephyr4772

An evaluation measures zero-shot Laya checkpoints as guards between agent steps

0·GitHub repo

Smart-Support-Ticket-Router-Using-Laya

affan1311

Routes support tickets with Laya and uses Gemini to draft replies, with a benchmark against a Gemini-only pipeline

0·GitHub repo

Smart-Support-Ticket-Classifier-Using-Laya

affanhyder-diggit

Routes support tickets with Laya and benchmarks the results against a Gemini-only pipeline

0·GitHub repo

privatemode-decisions-benchmark

edgelesssys

A benchmark compares speed, accuracy, and cost for Jev, Laya, and Privatemode decisions

0·GitHub repo

jev-equivalent-research

hjl1045

Evaluates Jev, Laya, and a GPT-5.6 Luna baseline on synthetic auto-claims classification

0·GitHub repo

MacJev-322M-4K-Laya-MLX

lawrence3699

Runs the MacJev typed-decision model natively on Apple silicon with MLX and compares it with Laya

0·GitHub repo

jev-bench

brandonrc

Benchmarks Jev, Laya, Claude Haiku, and CLM on package-registry triage tasks

0·GitHub repo

decision-lab

Amine-LG

Provides a visual playground to build decisions and compare Jev, OpenJEV, and local Laya

0·GitHub repo

KVev

AdudodlaVarish

Serves cached long-document chat and uses Laya for offline routing experiments

0·GitHub repo

decision-model-arena

sathik11

Compares Jev, Laya, and Microsoft Foundry for typed decisions in enterprise incident triage

0·GitHub repo

tool-output-pruning-lab

Dymyt-ry

Benchmarks Laya and other selectors for pruning tool output in coding-agent sessions

0·GitHub repo

kime-compat

tamnd

Checks whether kime supports TypeSafe and Laya APIs, clients, SDKs, and tests

0·GitHub repo

narde

anliang0306

Reimplements the Laya decision engine and verifies behavioral parity against upstream

0·GitHub repo

jev-agent-risk-gate

jayeshvpatil

Compares Jev, Claude, fine-tuned Qwen, and open Laya in experiments on agent shell-command risk gates

0·GitHub repo

d3code-calibration

gkastanis

Compares Jev and open-weights Laya probability calibration against human ratings

0·GitHub repo

jev-context-rot

casperkwok

Reproduces an experiment comparing context sensitivity in Jev, Laya and Decider models

0·GitHub repo

jev-vs-laya

Cognition-Forge

Compares Jev and Laya variants across typed-decision tasks

0·GitHub repo

Laya-Finetune

zamax14

Fine-tunes and evaluates multilingual Laya for support-ticket classification

0·GitHub repo

laya-games

BM-Popcorn

Demonstrates Laya-driven turn-based games and fine-tunes the model on engine-labeled positions

0·GitHub repo

laya_bio

maris205

Adapts the Laya encoder into a shared candidate-scoring model for biological sequence tasks

0·GitHub repo

laya_demo

itsvrushabh

Demonstrates and benchmarks Laya’s typed decision engine across interactive use cases

0·GitHub repo

laya-deploy

iintothewind

Deploys and benchmarks Laya Serve on Windows with Docker and a temporary public tunnel

0·GitHub repo

Official numbers

The core README and its BENCHMARKS.md report latency on a Tesla T4 (32.8 ms for one question with laya-multilingual, 39.5 ms with laya) and accuracy across 51 languages. Its Jev figures are third-party published, not measured by the Laya authors.

Independent head-to-heads

  • sysone-bench: byte-identical inputs across 751 states. Jev leads on triage, guardrails, moderation, Banking77 and multilingual intent. Laya leads on AG News and MNLI.
  • jev-laya-benchmark: 1,470 synthetic items. Jev 92.9% vs Laya 65.3% of judgments correct. For a single question, Laya was 42 ms locally on an M3 Pro and Jev was 136 ms. Jev was faster above 3 to 4 questions per request.
  • laya-jev-lab: 40 Chinese support tickets. Jev 78% at 588 ms, Laya 57% at 7.6 ms on an M4 Max. A cascade that escalates below 0.60 confidence matched Jev's accuracy at 1.8x its speed.
  • JevBench: a composite board. In v1.4.2, Laya (English, CPU) ranks 41st with a score of 30.25 and Jev 1.13.0 ranks 2nd with 63.29.
  • laya-mcp eval: key accuracy 32.5% for Laya vs 98.8% for Jev, with median latency of 26 ms vs 138 ms.

CPU latency

laya-cpu-benchmark measured laya-multilingual in steady state on an Intel i9-14900HX. One question took 50 ms for a 71-token state, 307 ms for 261 tokens, and 1,315 ms for 927 tokens. Batching gave little gain on CPU. Its conclusion: a realistic CPU call is 0.3 to 2 seconds, not 33 ms.

Run your own

git clone https://github.com/glukicov/laya_router && cd laya_router
uv sync --all-extras
uv run laya-router eval run --backend laya --device mps

Or use jevbench (dhruvmehra) to compare Jev, Laya, LLMs and BERT models on public datasets.

Reading these numbers

Most of these suites are small (40 to 1,500 items), synthetic or in one domain, and their authors say so. Latency depends on hardware, state length and question count. The core README's own advice: treat Laya as a fast base to specialise, and fine-tune for accuracy.

More ways to use Laya