Tincan0325/smoea-15ood-benchmark
SMoEA fixed 15-OOD benchmark data The evaluation data for scripts/run_rejection_benchmark.py. Not needed for interactive or batch mode — only the benchmark reads it. Getting it scripts/setup_workspace.sh downloads this for you into dataset/ood_data/, which is what --benchmark-root defaults to. To fetch it on its own: python scripts/fetch_benchmark.py --repo Tincan0325/smoea-15ood-benchmark That verifies the per-group counts on arrival. The plain hf command works… See the full description on the dataset page: https://huggingface.co/datasets/Tincan0325/smoea-15ood-benchmark.
flat layout: SHA256SUMS
flat layout: README.md
drop legacy nested layout: dataset/
drop legacy nested layout: ood/
drop legacy nested layout: prompts/
add flat layout: task933_test.json
add flat layout: task476_test.json
add flat layout: task1670_test.json
add flat layout: task1622_test.json
add flat layout: task149_test.json
add flat layout: ni_task_descriptions.json
add flat layout: mmlu_pro_test.json
add flat layout: bbh_test.json
SMoEA fixed 15-OOD benchmark data (4159 answer-free prompts)
initial commit
