system-one-benchmark

by yanng981

Compares accuracy, calibration and latency across System One models on multilingual datasets

Zero-shot accuracy, calibration and latency of Jev, Kev, Laya, Von and GLiNER2.5-Decide on SST-2, TREC, Banking77 and MASSIVE in 8 languages. Code and raw predictions.

Related projects