base-model-evals/off-the-shelf-model-evals
Babel Tower evaluation archive Shared filesystem for the base-model-evals project. This repository stores the full result for each evaluated question: model responses, available scores, original/rephrased variants, and the settings needed to interpret each run. The archive is initialized and ready for runs. No measured model results have been added yet. The downloadable toolkit includes a separately labeled synthetic demo; those demonstration records are not published as… See the full description on the dataset page: https://huggingface.co/datasets/base-model-evals/off-the-shelf-model-evals.
066
Add item-level response comparisons in toolkit 1.2.0
Fix toolkit download link
Initialize Babel Tower evaluation filesystem
initial commit
