Laya benchmarks

The official Laya README publishes detailed numbers, and a growing set of independent repositories now test those claims. The overall picture: Laya is much faster when run locally and free to self-host, while Jev is more accurate out of the box on most shared tests.

228 projects

layar

nullean

A .NET port of Laya provides ONNX and TorchSharp inference backends and CLI tools

0·GitHub repo

LAYA-RLCD

shyamsridhar123

Teaches RLCD and trains Laya pilots through a tactical game, lessons, and benchmarks

0·GitHub repo

laya-test

JVMoreiraD

Evaluates the Laya model's decision-making capabilities

0·GitHub repo

laya-serve

c4bbage

Serves Laya inference with Go, dynamic batching, and TensorRT or CUDA execution

0·GitHub repo

laya-mlx

Gates-456

Runs open-weight Laya typed decisions locally on Apple Silicon using MLX

0·GitHub repo

Laya-Experiments

alperunlu07

Collects reproducible experiments using Laya, including a Pong controller comparison

0·GitHub repo

laya-cpp

samiul000

Implements native C++ Laya inference with ONNX Runtime, benchmarking, and INT8 quantization

0·GitHub repo

laya-doom

dylanbstorey

Runs Laya in a real-time control loop to play Doom on Apple Silicon

0·GitHub repo

laya-dino

that-daniel

Controls Chrome's dinosaur game with Laya through a real-browser decision loop

0·GitHub repo

poc-laya

MayukhCars24

Measures Laya classification latency with a FastAPI backend, web console, and observability tools

0·GitHub repo

laya-ts

mikeboe

Serves the Laya model over an API and benchmarks it against the hosted Jev API

0·GitHub repo

migration-laya

ViniCarvalhoDados

Evaluates Laya for triaging legacy SQL queries before database migration

0·GitHub repo

Laya-Navigator

Vann4799

Adapts Laya decisions to workflow-state classification and next-action prediction

0·GitHub repo

jev-laya-explore

NachiketKandari

Evaluates the local Laya decision model alongside TypeSafe's hosted Jev

0·GitHub repo

laya-local-lab

torresnicolas0

Reproduces local evaluations of three Laya variants across support, guardrails, RAG, and model selection

0·GitHub repo

laya-mlx-benchmarks

jayluxferro

Benchmarks Laya typed-decision models running through the MLX inference port

0·GitHub repo

laya-medical-finetune

Priyanshu-5257

Provides RLCD fine-tuning recipes and evaluations for Laya on medical decision tasks

0·GitHub repo

laya-multilingual-dml

minicom365

Runs and benchmarks Laya Multilingual on AMD hardware using DirectML

0·GitHub repo

Jev-vs-Laya

aarush-dhingra

Compares Jev and Laya as chess decision-makers in a local browser arena

0·GitHub repo

laya-fit-check

JhouCode

A Python kit reproduces Laya's published benchmark and compares it with free baselines on custom data

0·GitHub repo

jev-v-laya

rnunley

A held-out MMLU-Pro benchmark compares answer routing with local Laya and hosted Jev

0·GitHub repo

laya-mlx-windows-cuda

oboroge0

Documents running the MLX Laya model on Windows with an NVIDIA GPU

0·GitHub repo

kime-bench

tamnd

Benchmarks kime against Laya, laya-mlx, laya-coreml, and Jev on matched hardware

0·GitHub repo

ewm-laya-bonsai-lab

alexmy21

Uses notebooks to inspect a Rust Laya daemon's interactions and compare Laya with Jev and a mock

0·GitHub repo

Official numbers

The core README and its BENCHMARKS.md report latency on a Tesla T4 (32.8 ms for one question with laya-multilingual, 39.5 ms with laya) and accuracy across 51 languages. Its Jev figures are third-party published, not measured by the Laya authors.

Independent head-to-heads

  • sysone-bench: byte-identical inputs across 751 states. Jev leads on triage, guardrails, moderation, Banking77 and multilingual intent. Laya leads on AG News and MNLI.
  • jev-laya-benchmark: 1,470 synthetic items. Jev 92.9% vs Laya 65.3% of judgments correct. For a single question, Laya was 42 ms locally on an M3 Pro and Jev was 136 ms. Jev was faster above 3 to 4 questions per request.
  • laya-jev-lab: 40 Chinese support tickets. Jev 78% at 588 ms, Laya 57% at 7.6 ms on an M4 Max. A cascade that escalates below 0.60 confidence matched Jev's accuracy at 1.8x its speed.
  • JevBench: a composite board. In v1.4.2, Laya (English, CPU) ranks 41st with a score of 30.25 and Jev 1.13.0 ranks 2nd with 63.29.
  • laya-mcp eval: key accuracy 32.5% for Laya vs 98.8% for Jev, with median latency of 26 ms vs 138 ms.

CPU latency

laya-cpu-benchmark measured laya-multilingual in steady state on an Intel i9-14900HX. One question took 50 ms for a 71-token state, 307 ms for 261 tokens, and 1,315 ms for 927 tokens. Batching gave little gain on CPU. Its conclusion: a realistic CPU call is 0.3 to 2 seconds, not 33 ms.

Run your own

git clone https://github.com/glukicov/laya_router && cd laya_router
uv sync --all-extras
uv run laya-router eval run --backend laya --device mps

Or use jevbench (dhruvmehra) to compare Jev, Laya, LLMs and BERT models on public datasets.

Reading these numbers

Most of these suites are small (40 to 1,500 items), synthetic or in one domain, and their authors say so. Latency depends on hardware, state length and question count. The core README's own advice: treat Laya as a fast base to specialise, and fine-tune for accuracy.

More ways to use Laya