Laya benchmarks

The official Laya README publishes detailed numbers, and a growing set of independent repositories now test those claims. The overall picture: Laya is much faster when run locally and free to self-host, while Jev is more accurate out of the box on most shared tests.

228 projects

jev-zen

loongWoong

Reconstructs and validates Laya inference locally using NumPy and a browser-based interface

0·GitHub repo

jev-eval-ja

unirt

Benchmarks Jev and Laya on Japanese business-style decision tasks

0·GitHub repo

laya

harisathees

Tests Laya predictions, multilingual behavior, CPU performance, and interactive runs through a local web console

0·GitHub repo

product_search_bench

Wanke15

Compares BM25, Jev, Qwen, and Laya for product search and reranking

0·GitHub repo

laya-t-rex-runner

10086ggqq

A browser-based T-Rex runner uses typed-decision models and distilled agents to choose actions

0·GitHub repo

laya-ROCm

don-milsey-miller

Provides an AMD ROCm runtime and benchmarks for running the Laya decision model

0·GitHub repo

laya-vk

sulistta

Ports the Laya decision runtime to a Vulkan/IREE encoder backend for AMD and Qualcomm GPUs

0·GitHub repo

laya-tictactoe

Kasa-Harendra

A terminal Tic-Tac-Toe demo runs the Laya model locally through Core ML on Apple Silicon

0·GitHub repo

laya-onnx

Geoking2104

Runs Laya decision models with ONNX Runtime and provides export, inference, and benchmark tools

0·GitHub repo

laya-agent

adhishthite

Benchmarks ConvAI Laya against TypeSafe Jev with and without live web grounding

0·GitHub repo

laya-example

lim6112j

Demonstrates typed-question routing with Laya checkpoints and startup optimizations

0·GitHub repo

laya-dino

JakkNaj

A fine-tuned Laya model plays Chrome Dino using a Python backend and TypeScript frontend

0·GitHub repo

laya.cpp

kyr0

Implements native C++ inference and a JEV-compatible server for Laya typed decisions

0·GitHub repo

laya-minesweeper

Animal2404

Demonstrates and evaluates Laya’s mine-risk judgments against a Minesweeper constraint solver

0·GitHub repo

laya-snapdragon

piffie

Runs Laya typed decisions on Snapdragon X NPUs through ONNX Runtime and Qualcomm QNN

0·GitHub repo

laya-tetris

kuchris

Runs Laya as a Tetris placement decision-maker with a heuristic safety guard

0·GitHub repo

Laya_Playground

jonas050210

Runs local Laya decision demos, games, and a raw playground through a browser interface

0·GitHub repo

laya-burmese

aungthuhein2005

Studies Laya zero-shot transfer, calibration, and fine-tuning for Burmese topic classification

0·GitHub repo

laya-todo

firede

Classifies to-do items locally with Laya and compares its predictions with an optional Kev backend

0·GitHub repo

laya-codex-bench

pbrehmer-ai

Benchmarks Laya-assisted Codex workflows for call reduction, latency, and decision quality

0·GitHub repo

laya-gfx1030

Thotheris

Runs and benchmarks the Laya decision model on Windows with an AMD Radeon RX 6900 XT

0·GitHub repo

laya-hexagon-npu

EricYu123456

Deploys Laya inference on Qualcomm Hexagon NPU using ONNX Runtime

0·GitHub repo

exp-laya-router

zheyar-ltd

Demonstrates a CUDA-based Laya policy router in real-time Snake and Tetris games

0·GitHub repo

laya-capability-business

matrix-air

Presents experiments measuring Laya’s capabilities, limitations, fine-tuning, and runtime migration

0·GitHub repo

Official numbers

The core README and its BENCHMARKS.md report latency on a Tesla T4 (32.8 ms for one question with laya-multilingual, 39.5 ms with laya) and accuracy across 51 languages. Its Jev figures are third-party published, not measured by the Laya authors.

Independent head-to-heads

  • sysone-bench: byte-identical inputs across 751 states. Jev leads on triage, guardrails, moderation, Banking77 and multilingual intent. Laya leads on AG News and MNLI.
  • jev-laya-benchmark: 1,470 synthetic items. Jev 92.9% vs Laya 65.3% of judgments correct. For a single question, Laya was 42 ms locally on an M3 Pro and Jev was 136 ms. Jev was faster above 3 to 4 questions per request.
  • laya-jev-lab: 40 Chinese support tickets. Jev 78% at 588 ms, Laya 57% at 7.6 ms on an M4 Max. A cascade that escalates below 0.60 confidence matched Jev's accuracy at 1.8x its speed.
  • JevBench: a composite board. In v1.4.2, Laya (English, CPU) ranks 41st with a score of 30.25 and Jev 1.13.0 ranks 2nd with 63.29.
  • laya-mcp eval: key accuracy 32.5% for Laya vs 98.8% for Jev, with median latency of 26 ms vs 138 ms.

CPU latency

laya-cpu-benchmark measured laya-multilingual in steady state on an Intel i9-14900HX. One question took 50 ms for a 71-token state, 307 ms for 261 tokens, and 1,315 ms for 927 tokens. Batching gave little gain on CPU. Its conclusion: a realistic CPU call is 0.3 to 2 seconds, not 33 ms.

Run your own

git clone https://github.com/glukicov/laya_router && cd laya_router
uv sync --all-extras
uv run laya-router eval run --backend laya --device mps

Or use jevbench (dhruvmehra) to compare Jev, Laya, LLMs and BERT models on public datasets.

Reading these numbers

Most of these suites are small (40 to 1,500 items), synthetic or in one domain, and their authors say so. Latency depends on hardware, state length and question count. The core README's own advice: treat Laya as a fast base to specialise, and fine-tune for accuracy.

More ways to use Laya