d3code-calibration

by gkastanis

Compares Jev and open-weights Laya probability calibration against human ratings

Checking whether a model's probability means what it says: TypeSafe Jev and open-weights Laya against 150,000 human ratings, with stdlib tools to run the same check on your own data.

Related projects