datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
untested-nekoextracted-untestedgpt-oss-120b-untested
gpt-oss-120b-untested
Repo: tytodd/gpt-oss-120b-untested
Local path: datasets/gpt-oss-120b-untested
Config: configs/datasets/untested/untested.yaml
benchmark
train
val
ood
all
mmlu
95.00%
95.00%
arc_challenge
100.00%
100.00%
gpqa_diamond
90.00%
90.00%
mmlu_pro
85.00%
85.00%
musr_murder_mysteries
70.00%
70.00%
musr_object_placements
55.00%
55.00%
musr_team_allocation
55.00%
55.00%
argument_quality_ranking
10.00%
10.00%… See the full description on the dataset page: https://huggingface.co/datasets/tytodd/gpt-oss-120b-untested.untested-WiP-dpoDreadPoor__UNTESTED-VENN_1.2-8B-Model_Stock-details
Dataset Card for Evaluation run of DreadPoor/UNTESTED-VENN_1.2-8B-Model_Stock
Dataset automatically created during the evaluation run of model DreadPoor/UNTESTED-VENN_1.2-8B-Model_Stock
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__UNTESTED-VENN_1.2-8B-Model_Stock-details.
