Laya benchmarks

The official Laya README publishes detailed numbers, and a growing set of independent repositories now test those claims. The overall picture: Laya is much faster when run locally and free to self-host, while Jev is more accurate out of the box on most shared tests.

228 projects

laya-scripts

muck-stump

Shell scripts test authentication, server behavior, concurrency, and routing for Laya-compatible APIs

0·GitHub repo

laya-steering-lab

Hantlowt

Tests methods for specializing frozen Laya models and benchmarks their accuracy and runtime

0·GitHub repo

apm-laya-triage

danielmeppiel

Benchmarks local Laya issue classification against labels in the Microsoft APM corpus

0·GitHub repo

laya-training-log

ChenneyZhuang

Documents training and benchmark results for a fine-tuned Laya browser model

0·GitHub repo

laya-first-look

cvranjith

Explores base Laya checkpoints locally on Apple M4 and reports informal behavior and latency measurements

0·GitHub repo

Tenstorrent.Blackhole-convaiinnovations_laya

Thatch-cloud

Develops a Tenstorrent Blackhole inference backend and service integration for Laya

0·GitHub repo

jev_and_laya_benchmarking

pavanjava

Benchmarks Jev and Laya accuracy and inference speed on typed-decision tasks

0·GitHub repo

jev-laya-openai-comparison

amansahani

Benchmarks Laya and Jev against OpenAI models on financial regulatory decisions

0·GitHub repo

decision-model-bench

SaiNarayana-B

Tests the accuracy and calibration of Laya decision-model confidence scores

0·GitHub repo
E

jev-arena

eliot5566

Lets users create text-driven fighting bots piloted by Jev, Laya, or other models

0·npm package
B

laya-calibration-lab

BunsDev

Fits and evaluates probability calibration for Laya typed-decision models

0·HF Space
X

laya-fp16-parity

xauberer93

Checks FP16 parity for the Laya model on a ZeroGPU lane

0·HF Space

Official numbers

The core README and its BENCHMARKS.md report latency on a Tesla T4 (32.8 ms for one question with laya-multilingual, 39.5 ms with laya) and accuracy across 51 languages. Its Jev figures are third-party published, not measured by the Laya authors.

Independent head-to-heads

  • sysone-bench: byte-identical inputs across 751 states. Jev leads on triage, guardrails, moderation, Banking77 and multilingual intent. Laya leads on AG News and MNLI.
  • jev-laya-benchmark: 1,470 synthetic items. Jev 92.9% vs Laya 65.3% of judgments correct. For a single question, Laya was 42 ms locally on an M3 Pro and Jev was 136 ms. Jev was faster above 3 to 4 questions per request.
  • laya-jev-lab: 40 Chinese support tickets. Jev 78% at 588 ms, Laya 57% at 7.6 ms on an M4 Max. A cascade that escalates below 0.60 confidence matched Jev's accuracy at 1.8x its speed.
  • JevBench: a composite board. In v1.4.2, Laya (English, CPU) ranks 41st with a score of 30.25 and Jev 1.13.0 ranks 2nd with 63.29.
  • laya-mcp eval: key accuracy 32.5% for Laya vs 98.8% for Jev, with median latency of 26 ms vs 138 ms.

CPU latency

laya-cpu-benchmark measured laya-multilingual in steady state on an Intel i9-14900HX. One question took 50 ms for a 71-token state, 307 ms for 261 tokens, and 1,315 ms for 927 tokens. Batching gave little gain on CPU. Its conclusion: a realistic CPU call is 0.3 to 2 seconds, not 33 ms.

Run your own

git clone https://github.com/glukicov/laya_router && cd laya_router
uv sync --all-extras
uv run laya-router eval run --backend laya --device mps

Or use jevbench (dhruvmehra) to compare Jev, Laya, LLMs and BERT models on public datasets.

Reading these numbers

Most of these suites are small (40 to 1,500 items), synthetic or in one domain, and their authors say so. Latency depends on hardware, state length and question count. The core README's own advice: treat Laya as a fast base to specialise, and fine-tune for accuracy.

More ways to use Laya