datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Qwen__Qwen2.5-7B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2.5-7B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-7B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-7B-Instruct-details.Qwen__Qwen2.5-72B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2.5-72B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-72B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-72B-Instruct-details.JungZoona__T3Q-qwen2.5-14b-v1.0-e3-details
Dataset Card for Evaluation run of JungZoona/T3Q-qwen2.5-14b-v1.0-e3
Dataset automatically created during the evaluation run of model JungZoona/T3Q-qwen2.5-14b-v1.0-e3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/JungZoona__T3Q-qwen2.5-14b-v1.0-e3-details.Qwen__Qwen2.5-32B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2.5-32B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-32B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-32B-Instruct-details.Qwen__Qwen2-1.5B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2-1.5B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2-1.5B-Instruct
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2-1.5B-Instruct-details.Qwen__Qwen2-0.5B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2-0.5B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2-0.5B-Instruct
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2-0.5B-Instruct-details.Qwen__Qwen2.5-Math-1.5B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2.5-Math-1.5B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-Math-1.5B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-Math-1.5B-Instruct-details.Qwen__Qwen2.5-0.5B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2.5-0.5B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-0.5B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-0.5B-Instruct-details.Qwen__Qwen2-7B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2-7B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2-7B-Instruct
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2-7B-Instruct-details.Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8-details
Dataset Card for Evaluation run of Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8
Dataset automatically created during the evaluation run of model Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8-details.Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8.9-details
Dataset Card for Evaluation run of Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8.9
Dataset automatically created during the evaluation run of model Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8.9
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8.9-details.Qwen__Qwen2.5-7B-details
Dataset Card for Evaluation run of Qwen/Qwen2.5-7B
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-7B-details.EVA-UNIT-01__EVA-Qwen2.5-72B-v0.2-details
Dataset Card for Evaluation run of EVA-UNIT-01/EVA-Qwen2.5-72B-v0.2
Dataset automatically created during the evaluation run of model EVA-UNIT-01/EVA-Qwen2.5-72B-v0.2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EVA-UNIT-01__EVA-Qwen2.5-72B-v0.2-details.Qwen__Qwen2-7B-details
Dataset Card for Evaluation run of Qwen/Qwen2-7B
Dataset automatically created during the evaluation run of model Qwen/Qwen2-7B
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2-7B-details.YOYO-AI__Qwen2.5-14B-YOYO-V4-details
Dataset Card for Evaluation run of YOYO-AI/Qwen2.5-14B-YOYO-V4
Dataset automatically created during the evaluation run of model YOYO-AI/Qwen2.5-14B-YOYO-V4
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/YOYO-AI__Qwen2.5-14B-YOYO-V4-details.Qwen__Qwen2-72B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2-72B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2-72B-Instruct
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2-72B-Instruct-details.ZeroXClem__Qwen2.5-7B-Qandora-CySec-details
Dataset Card for Evaluation run of ZeroXClem/Qwen2.5-7B-Qandora-CySec
Dataset automatically created during the evaluation run of model ZeroXClem/Qwen2.5-7B-Qandora-CySec
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ZeroXClem__Qwen2.5-7B-Qandora-CySec-details.MaziyarPanahi__calme-2.1-qwen2.5-72b-details
Dataset Card for Evaluation run of MaziyarPanahi/calme-2.1-qwen2.5-72b
Dataset automatically created during the evaluation run of model MaziyarPanahi/calme-2.1-qwen2.5-72b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/MaziyarPanahi__calme-2.1-qwen2.5-72b-details.tanliboy__lambda-qwen2.5-32b-dpo-test-details
Dataset Card for Evaluation run of tanliboy/lambda-qwen2.5-32b-dpo-test
Dataset automatically created during the evaluation run of model tanliboy/lambda-qwen2.5-32b-dpo-test
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tanliboy__lambda-qwen2.5-32b-dpo-test-details.ssmits__Qwen2.5-95B-Instruct-details
Dataset Card for Evaluation run of ssmits/Qwen2.5-95B-Instruct
Dataset automatically created during the evaluation run of model ssmits/Qwen2.5-95B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ssmits__Qwen2.5-95B-Instruct-details.CombinHorizon__huihui-ai-abliterated-Qwen2.5-32B-Inst-BaseMerge-TIES-details
Dataset Card for Evaluation run of CombinHorizon/huihui-ai-abliterated-Qwen2.5-32B-Inst-BaseMerge-TIES
Dataset automatically created during the evaluation run of model CombinHorizon/huihui-ai-abliterated-Qwen2.5-32B-Inst-BaseMerge-TIES
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/CombinHorizon__huihui-ai-abliterated-Qwen2.5-32B-Inst-BaseMerge-TIES-details.zetasepic__Qwen2.5-32B-Instruct-abliterated-v2-details
Dataset Card for Evaluation run of zetasepic/Qwen2.5-32B-Instruct-abliterated-v2
Dataset automatically created during the evaluation run of model zetasepic/Qwen2.5-32B-Instruct-abliterated-v2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/zetasepic__Qwen2.5-32B-Instruct-abliterated-v2-details.Qwen__Qwen2.5-3B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2.5-3B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-3B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-3B-Instruct-details.MaziyarPanahi__calme-2.2-qwen2.5-72b-details
Dataset Card for Evaluation run of MaziyarPanahi/calme-2.2-qwen2.5-72b
Dataset automatically created during the evaluation run of model MaziyarPanahi/calme-2.2-qwen2.5-72b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/MaziyarPanahi__calme-2.2-qwen2.5-72b-details.newsbang__Homer-v1.0-Qwen2.5-72B-details
Dataset Card for Evaluation run of newsbang/Homer-v1.0-Qwen2.5-72B
Dataset automatically created during the evaluation run of model newsbang/Homer-v1.0-Qwen2.5-72B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/newsbang__Homer-v1.0-Qwen2.5-72B-details.ZeroXClem__Qwen2.5-7B-CelestialHarmony-1M-details
Dataset Card for Evaluation run of ZeroXClem/Qwen2.5-7B-CelestialHarmony-1M
Dataset automatically created during the evaluation run of model ZeroXClem/Qwen2.5-7B-CelestialHarmony-1M
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ZeroXClem__Qwen2.5-7B-CelestialHarmony-1M-details.JayHyeon__Qwen2.5-0.5B-SFT-2e-5-2ep-DPO_3e-7-3ep_0alp_0lam-details
Dataset Card for Evaluation run of JayHyeon/Qwen2.5-0.5B-SFT-2e-5-2ep-DPO_3e-7-3ep_0alp_0lam
Dataset automatically created during the evaluation run of model JayHyeon/Qwen2.5-0.5B-SFT-2e-5-2ep-DPO_3e-7-3ep_0alp_0lam
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/JayHyeon__Qwen2.5-0.5B-SFT-2e-5-2ep-DPO_3e-7-3ep_0alp_0lam-details.bunnycore__Qwen2.5-7B-Instruct-Merge-Stock-v0.1-details
Dataset Card for Evaluation run of bunnycore/Qwen2.5-7B-Instruct-Merge-Stock-v0.1
Dataset automatically created during the evaluation run of model bunnycore/Qwen2.5-7B-Instruct-Merge-Stock-v0.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bunnycore__Qwen2.5-7B-Instruct-Merge-Stock-v0.1-details.migtissera__Tess-v2.5.2-Qwen2-72B-details
Dataset Card for Evaluation run of migtissera/Tess-v2.5.2-Qwen2-72B
Dataset automatically created during the evaluation run of model migtissera/Tess-v2.5.2-Qwen2-72B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/migtissera__Tess-v2.5.2-Qwen2-72B-details.CoolSpring__Qwen2-0.5B-Abyme-merge2-details
Dataset Card for Evaluation run of CoolSpring/Qwen2-0.5B-Abyme-merge2
Dataset automatically created during the evaluation run of model CoolSpring/Qwen2-0.5B-Abyme-merge2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/CoolSpring__Qwen2-0.5B-Abyme-merge2-details.
