datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
extracted-untestedgpt-oss-120b-untested
gpt-oss-120b-untested
Repo: tytodd/gpt-oss-120b-untested
Local path: datasets/gpt-oss-120b-untested
Config: configs/datasets/untested/untested.yaml
benchmark
train
val
ood
all
mmlu
95.00%
95.00%
arc_challenge
100.00%
100.00%
gpqa_diamond
90.00%
90.00%
mmlu_pro
85.00%
85.00%
musr_murder_mysteries
70.00%
70.00%
musr_object_placements
55.00%
55.00%
musr_team_allocation
55.00%
55.00%
argument_quality_ranking
10.00%
10.00%… See the full description on the dataset page: https://huggingface.co/datasets/tytodd/gpt-oss-120b-untested.
