jev-websearch-eval
by anweat
An offline web-search evidence evaluation compares hosted Jev with local Laya and a rule baseline
Independent offline evaluation of the Bocha Jev decision model in a web-search evidence pipeline (vs. rule baseline and Laya): tasks, metadata-only snapshots, LLM-draft labels, judge outputs, reports
Platforms
Use cases