d3code-calibration
by gkastanis
Compares Jev and open-weights Laya probability calibration against human ratings
Checking whether a model's probability means what it says: TypeSafe Jev and open-weights Laya against 150,000 human ratings, with stdlib tools to run the same check on your own data.
Platforms
Use cases