system-one-bench

by alilibx

Benchmarks typed-decision models, including Laya, across labelled datasets for accuracy, calibration, latency and cost

System One Bench: typed-decision models (TypeSafe Jev, OpenAI Decisions, open-weight Laya) on real labelled datasets, with accuracy, calibration, latency and cost, plus a live race

Platforms

On X

Related projects