system-one-benchmark
by yanng981
Compares accuracy, calibration and latency across System One models on multilingual datasets
Zero-shot accuracy, calibration and latency of Jev, Kev, Laya, Von and GLiNER2.5-Decide on SST-2, TREC, Banking77 and MASSIVE in 8 languages. Code and raw predictions.
Platforms
Use cases