datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nla-av-responses-llama-70b-layer53nla-av-ar-attribution-llama-70b-layer53Llama-3.3-70B-Instruct-eval-logs-and-scoresLlama-3-Taiwan-70B-Instruct-eval-logs-and-scoresmeta-llama__Meta-Llama-3-70B-Instruct-details
Dataset Card for Evaluation run of meta-llama/Meta-Llama-3-70B-Instruct
Dataset automatically created during the evaluation run of model meta-llama/Meta-Llama-3-70B-Instruct
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/meta-llama__Meta-Llama-3-70B-Instruct-details.rootxhacker__Apollo-70B-details
Dataset Card for Evaluation run of rootxhacker/Apollo-70B
Dataset automatically created during the evaluation run of model rootxhacker/Apollo-70B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/rootxhacker__Apollo-70B-details.allenai__Llama-3.1-Tulu-3-70B-details
Dataset Card for Evaluation run of allenai/Llama-3.1-Tulu-3-70B
Dataset automatically created during the evaluation run of model allenai/Llama-3.1-Tulu-3-70B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/allenai__Llama-3.1-Tulu-3-70B-details.mlabonne__Hermes-3-Llama-3.1-70B-lorablated-details
Dataset Card for Evaluation run of mlabonne/Hermes-3-Llama-3.1-70B-lorablated
Dataset automatically created during the evaluation run of model mlabonne/Hermes-3-Llama-3.1-70B-lorablated
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mlabonne__Hermes-3-Llama-3.1-70B-lorablated-details.tenyx__Llama3-TenyxChat-70B-details
Dataset Card for Evaluation run of tenyx/Llama3-TenyxChat-70B
Dataset automatically created during the evaluation run of model tenyx/Llama3-TenyxChat-70B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tenyx__Llama3-TenyxChat-70B-details.VAGOsolutions__Llama-3.1-SauerkrautLM-70b-Instruct-details
Dataset Card for Evaluation run of VAGOsolutions/Llama-3.1-SauerkrautLM-70b-Instruct
Dataset automatically created during the evaluation run of model VAGOsolutions/Llama-3.1-SauerkrautLM-70b-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/VAGOsolutions__Llama-3.1-SauerkrautLM-70b-Instruct-details.gbueno86__Meta-LLama-3-Cat-Smaug-LLama-70b-details
Dataset Card for Evaluation run of gbueno86/Meta-LLama-3-Cat-Smaug-LLama-70b
Dataset automatically created during the evaluation run of model gbueno86/Meta-LLama-3-Cat-Smaug-LLama-70b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/gbueno86__Meta-LLama-3-Cat-Smaug-LLama-70b-details.cloudyu__Llama-3-70Bx2-MOE-details
Dataset Card for Evaluation run of cloudyu/Llama-3-70Bx2-MOE
Dataset automatically created during the evaluation run of model cloudyu/Llama-3-70Bx2-MOE
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cloudyu__Llama-3-70Bx2-MOE-details.failspy__Meta-Llama-3-70B-Instruct-abliterated-v3.5-details
Dataset Card for Evaluation run of failspy/Meta-Llama-3-70B-Instruct-abliterated-v3.5
Dataset automatically created during the evaluation run of model failspy/Meta-Llama-3-70B-Instruct-abliterated-v3.5
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/failspy__Meta-Llama-3-70B-Instruct-abliterated-v3.5-details.Wildchat-RIP-Filtered-by-70b-LlamaRIP is a method for perference data filtering. The core idea is that low-quality input prompts lead to high variance and low-quality responses. By measuring the quality of rejected responses and the reward gap between chosen and rejected preference pairs, RIP effectively filters prompts to enhance dataset quality.
We release 4k data that filtered from 20k Wildchat prompts. For each prompt, we provide 32 responses from Llama-3.3-70B-Instruct and their corresponding rewards obtained from ArmoRM.… See the full description on the dataset page: https://huggingface.co/datasets/facebook/Wildchat-RIP-Filtered-by-70b-Llama.152334H__miqu-1-70b-sf-details
Dataset Card for Evaluation run of 152334H/miqu-1-70b-sf
Dataset automatically created during the evaluation run of model 152334H/miqu-1-70b-sf
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/152334H__miqu-1-70b-sf-details.meta-llama__Llama-3.1-70B-Instruct-details
Dataset Card for Evaluation run of meta-llama/Llama-3.1-70B-Instruct
Dataset automatically created during the evaluation run of model meta-llama/Llama-3.1-70B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/meta-llama__Llama-3.1-70B-Instruct-details.deepseek-ai__DeepSeek-R1-Distill-Llama-70B-details
Dataset Card for Evaluation run of deepseek-ai/DeepSeek-R1-Distill-Llama-70B
Dataset automatically created during the evaluation run of model deepseek-ai/DeepSeek-R1-Distill-Llama-70B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/deepseek-ai__DeepSeek-R1-Distill-Llama-70B-details.pankajmathur__orca_mini_v9_2_70b-details
Dataset Card for Evaluation run of pankajmathur/orca_mini_v9_2_70b
Dataset automatically created during the evaluation run of model pankajmathur/orca_mini_v9_2_70b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/pankajmathur__orca_mini_v9_2_70b-details.MaziyarPanahi__calme-2.3-llama3.1-70b-details
Dataset Card for Evaluation run of MaziyarPanahi/calme-2.3-llama3.1-70b
Dataset automatically created during the evaluation run of model MaziyarPanahi/calme-2.3-llama3.1-70b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/MaziyarPanahi__calme-2.3-llama3.1-70b-details.pankajmathur__orca_mini_v8_1_70b-details
Dataset Card for Evaluation run of pankajmathur/orca_mini_v8_1_70b
Dataset automatically created during the evaluation run of model pankajmathur/orca_mini_v8_1_70b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/pankajmathur__orca_mini_v8_1_70b-details.nbeerbower__Llama3.1-Gutenberg-Doppel-70B-details
Dataset Card for Evaluation run of nbeerbower/Llama3.1-Gutenberg-Doppel-70B
Dataset automatically created during the evaluation run of model nbeerbower/Llama3.1-Gutenberg-Doppel-70B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/nbeerbower__Llama3.1-Gutenberg-Doppel-70B-details.nbeerbower__llama3.1-kartoffeldes-70B-details
Dataset Card for Evaluation run of nbeerbower/llama3.1-kartoffeldes-70B
Dataset automatically created during the evaluation run of model nbeerbower/llama3.1-kartoffeldes-70B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/nbeerbower__llama3.1-kartoffeldes-70B-details.KSU-HW-SEC__Llama3-70b-SVA-FT-1415-details
Dataset Card for Evaluation run of KSU-HW-SEC/Llama3-70b-SVA-FT-1415
Dataset automatically created during the evaluation run of model KSU-HW-SEC/Llama3-70b-SVA-FT-1415
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/KSU-HW-SEC__Llama3-70b-SVA-FT-1415-details.KSU-HW-SEC__Llama3-70b-SVA-FT-final-details
Dataset Card for Evaluation run of KSU-HW-SEC/Llama3-70b-SVA-FT-final
Dataset automatically created during the evaluation run of model KSU-HW-SEC/Llama3-70b-SVA-FT-final
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/KSU-HW-SEC__Llama3-70b-SVA-FT-final-details.m42-health__Llama3-Med42-70B-details
Dataset Card for Evaluation run of m42-health/Llama3-Med42-70B
Dataset automatically created during the evaluation run of model m42-health/Llama3-Med42-70B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/m42-health__Llama3-Med42-70B-details.Steelskull__L3.3-Nevoria-R1-70b-details
Dataset Card for Evaluation run of Steelskull/L3.3-Nevoria-R1-70b
Dataset automatically created during the evaluation run of model Steelskull/L3.3-Nevoria-R1-70b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Steelskull__L3.3-Nevoria-R1-70b-details.ehristoforu__della-70b-test-v1-details
Dataset Card for Evaluation run of ehristoforu/della-70b-test-v1
Dataset automatically created during the evaluation run of model ehristoforu/della-70b-test-v1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ehristoforu__della-70b-test-v1-details.meta-llama__Llama-2-70b-chat-hf-details
Dataset Card for Evaluation run of meta-llama/Llama-2-70b-chat-hf
Dataset automatically created during the evaluation run of model meta-llama/Llama-2-70b-chat-hf
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/meta-llama__Llama-2-70b-chat-hf-details.meta-llama__Meta-Llama-3-70B-details
Dataset Card for Evaluation run of meta-llama/Meta-Llama-3-70B
Dataset automatically created during the evaluation run of model meta-llama/Meta-Llama-3-70B
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/meta-llama__Meta-Llama-3-70B-details.nvidia__Llama-3.1-Nemotron-70B-Instruct-HF-details
Dataset Card for Evaluation run of nvidia/Llama-3.1-Nemotron-70B-Instruct-HF
Dataset automatically created during the evaluation run of model nvidia/Llama-3.1-Nemotron-70B-Instruct-HF
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/nvidia__Llama-3.1-Nemotron-70B-Instruct-HF-details.
