datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
repro-organic-data-72BQwen__Qwen2.5-72B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2.5-72B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-72B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-72B-Instruct-details.abacusai__Dracarys-72B-Instruct-details
Dataset Card for Evaluation run of abacusai/Dracarys-72B-Instruct
Dataset automatically created during the evaluation run of model abacusai/Dracarys-72B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abacusai__Dracarys-72B-Instruct-details.rubenroy__Gilgamesh-72B-details
Dataset Card for Evaluation run of rubenroy/Gilgamesh-72B
Dataset automatically created during the evaluation run of model rubenroy/Gilgamesh-72B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/rubenroy__Gilgamesh-72B-details.EVA-UNIT-01__EVA-Qwen2.5-72B-v0.2-details
Dataset Card for Evaluation run of EVA-UNIT-01/EVA-Qwen2.5-72B-v0.2
Dataset automatically created during the evaluation run of model EVA-UNIT-01/EVA-Qwen2.5-72B-v0.2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EVA-UNIT-01__EVA-Qwen2.5-72B-v0.2-details.anthracite-org__magnum-v2-72b-details
Dataset Card for Evaluation run of anthracite-org/magnum-v2-72b
Dataset automatically created during the evaluation run of model anthracite-org/magnum-v2-72b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/anthracite-org__magnum-v2-72b-details.Qwen__Qwen2-72B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2-72B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2-72B-Instruct
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2-72B-Instruct-details.MaziyarPanahi__calme-2.1-qwen2.5-72b-details
Dataset Card for Evaluation run of MaziyarPanahi/calme-2.1-qwen2.5-72b
Dataset automatically created during the evaluation run of model MaziyarPanahi/calme-2.1-qwen2.5-72b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/MaziyarPanahi__calme-2.1-qwen2.5-72b-details.MaziyarPanahi__calme-2.2-qwen2.5-72b-details
Dataset Card for Evaluation run of MaziyarPanahi/calme-2.2-qwen2.5-72b
Dataset automatically created during the evaluation run of model MaziyarPanahi/calme-2.2-qwen2.5-72b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/MaziyarPanahi__calme-2.2-qwen2.5-72b-details.newsbang__Homer-v1.0-Qwen2.5-72B-details
Dataset Card for Evaluation run of newsbang/Homer-v1.0-Qwen2.5-72B
Dataset automatically created during the evaluation run of model newsbang/Homer-v1.0-Qwen2.5-72B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/newsbang__Homer-v1.0-Qwen2.5-72B-details.migtissera__Tess-v2.5.2-Qwen2-72B-details
Dataset Card for Evaluation run of migtissera/Tess-v2.5.2-Qwen2-72B
Dataset automatically created during the evaluation run of model migtissera/Tess-v2.5.2-Qwen2-72B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/migtissera__Tess-v2.5.2-Qwen2-72B-details.huihui-ai__Qwen2.5-72B-Instruct-abliterated-details
Dataset Card for Evaluation run of huihui-ai/Qwen2.5-72B-Instruct-abliterated
Dataset automatically created during the evaluation run of model huihui-ai/Qwen2.5-72B-Instruct-abliterated
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/huihui-ai__Qwen2.5-72B-Instruct-abliterated-details.Sakalti__ultiima-72B-v1.5-details
Dataset Card for Evaluation run of Sakalti/ultiima-72B-v1.5
Dataset automatically created during the evaluation run of model Sakalti/ultiima-72B-v1.5
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sakalti__ultiima-72B-v1.5-details.MaziyarPanahi__calme-2.1-qwen2-72b-details
Dataset Card for Evaluation run of MaziyarPanahi/calme-2.1-qwen2-72b
Dataset automatically created during the evaluation run of model MaziyarPanahi/calme-2.1-qwen2-72b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/MaziyarPanahi__calme-2.1-qwen2-72b-details.Qwen__Qwen2.5-72B-details
Dataset Card for Evaluation run of Qwen/Qwen2.5-72B
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-72B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-72B-details.Qwen__Qwen2.5-Math-72B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2.5-Math-72B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-Math-72B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-Math-72B-Instruct-details.Qwen__Qwen2-Math-72B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2-Math-72B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2-Math-72B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2-Math-72B-Instruct-details.zetasepic__Qwen2.5-72B-Instruct-abliterated-details
Dataset Card for Evaluation run of zetasepic/Qwen2.5-72B-Instruct-abliterated
Dataset automatically created during the evaluation run of model zetasepic/Qwen2.5-72B-Instruct-abliterated
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/zetasepic__Qwen2.5-72B-Instruct-abliterated-details.nvidia__AceMath-72B-Instruct-details
Dataset Card for Evaluation run of nvidia/AceMath-72B-Instruct
Dataset automatically created during the evaluation run of model nvidia/AceMath-72B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/nvidia__AceMath-72B-Instruct-details.nvidia__AceInstruct-72B-details
Dataset Card for Evaluation run of nvidia/AceInstruct-72B
Dataset automatically created during the evaluation run of model nvidia/AceInstruct-72B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/nvidia__AceInstruct-72B-details.Qwen__Qwen2-72B-details
Dataset Card for Evaluation run of Qwen/Qwen2-72B
Dataset automatically created during the evaluation run of model Qwen/Qwen2-72B
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2-72B-details.qwen2-72b-magpie-enQwen__Qwen2-VL-72B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2-VL-72B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2-VL-72B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2-VL-72B-Instruct-details.cognitivecomputations__dolphin-2.9.2-qwen2-72b-details
Dataset Card for Evaluation run of cognitivecomputations/dolphin-2.9.2-qwen2-72b
Dataset automatically created during the evaluation run of model cognitivecomputations/dolphin-2.9.2-qwen2-72b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cognitivecomputations__dolphin-2.9.2-qwen2-72b-details.abacusai__Smaug-72B-v0.1-details
Dataset Card for Evaluation run of abacusai/Smaug-72B-v0.1
Dataset automatically created during the evaluation run of model abacusai/Smaug-72B-v0.1
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abacusai__Smaug-72B-v0.1-details.abacusai__Smaug-Qwen2-72B-Instruct-details
Dataset Card for Evaluation run of abacusai/Smaug-Qwen2-72B-Instruct
Dataset automatically created during the evaluation run of model abacusai/Smaug-Qwen2-72B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abacusai__Smaug-Qwen2-72B-Instruct-details.nvidia__AceMath-72B-RM-details
Dataset Card for Evaluation run of nvidia/AceMath-72B-RM
Dataset automatically created during the evaluation run of model nvidia/AceMath-72B-RM
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/nvidia__AceMath-72B-RM-details.dfurman__Qwen2-72B-Orpo-v0.1-details
Dataset Card for Evaluation run of dfurman/Qwen2-72B-Orpo-v0.1
Dataset automatically created during the evaluation run of model dfurman/Qwen2-72B-Orpo-v0.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/dfurman__Qwen2-72B-Orpo-v0.1-details.davidkim205__Rhea-72b-v0.5-details
Dataset Card for Evaluation run of davidkim205/Rhea-72b-v0.5
Dataset automatically created during the evaluation run of model davidkim205/Rhea-72b-v0.5
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/davidkim205__Rhea-72b-v0.5-details.pankajmathur__orca_mini_v7_72b-details
Dataset Card for Evaluation run of pankajmathur/orca_mini_v7_72b
Dataset automatically created during the evaluation run of model pankajmathur/orca_mini_v7_72b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/pankajmathur__orca_mini_v7_72b-details.
