datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hpltv2-llama33-edu-annotation
HPLT version 2.0 educational annotations
This dataset contains annotations derived from HPLT v2 cleaned samples.
There are 500,000 annotations for each language if the source contains at least 500,000 samples.
We prompt Llama-3.3-70B-Instruct to score web pages based on their educational value following FineWeb-Edu classifier.
Note 1: The dataset contains the prompt (using the first 1500 characters of the text sample), the scores, and the full Llama 3 generation. The column "idx"… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/hpltv2-llama33-edu-annotation.lm-eval-results-hkust-nlp-dart-math-llama3-8b-prop2diff-private
Dataset Card for Evaluation run of hkust-nlp/dart-math-llama3-8b-prop2diff
Dataset automatically created during the evaluation run of model hkust-nlp/dart-math-llama3-8b-prop2diff
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-hkust-nlp-dart-math-llama3-8b-prop2diff-private.fineweb-edu-pretokenized-llama3-100b
FineWeb-Edu Pretokenized with Llama 3.1 (100B)
This repository contains the sample/100BT subset of FineWeb-Edu, pretokenized for Megatron-LM-style training with meta-llama/Meta-Llama-3.1-8B.
It is a derived, pretokenized version of the upstream data, not an official Hugging Face FineWeb release.
Dataset summary
140 indexed shards
97,270,686 non-empty documents
97,458,793,013 tokens
English web text from FineWeb-Edu sample/100BT
Source dataset revision:… See the full description on the dataset page: https://huggingface.co/datasets/ZhuofengLi/fineweb-edu-pretokenized-llama3-100b.Llama-3.3-70B-Instruct-eval-logs-and-scoresllama-3.2-3B-f1-instruct-eval-logs-and-scoresLlama-3.1-8B-Instruct-eval-logs-and-scoresLlama-3-Taiwan-70B-Instruct-eval-logs-and-scoresLlama-3.2-3B-Instruct-eval-logs-and-scoresLlama-3.1-Taiwan-8B-Instruct-eval-logs-and-scoresmultimodality-poc-llama31-ruler16k
Multimodality PoC corpus — Llama-3.1-8B-Instruct on RULER-16K
Raw pre-RoPE query and hidden-state tensors captured during prefill, used
to study whether the per-(layer, kv_head) query distribution is unimodal
Gaussian (the assumption underpinning Expected Attention's MGF closed-form
in kvpress).
What's in here
65 .npz files, one per (RULER task, prompt_index) pair (13 tasks × 5
prompts).
Each file (~414 MB) contains:
field
dtype
shape
meaning
hidden
float16… See the full description on the dataset page: https://huggingface.co/datasets/June30916/multimodality-poc-llama31-ruler16k.llama-3.1-medprm-reward-training-set
Med-PRM-Reward (Version 1.0)
🚀 Med-PRM-Reward is among the first Process Reward Models (PRMs) specifically designed for the medical domain. Unlike conventional PRMs, it enhances its verification capabilities by integrating clinical knowledge through retrieval-augmented generation (RAG). Med-PRM-Reward demonstrates exceptional performance in scaling-test-time computation, particularly outperforming majority‐voting ensembles on complex medical reasoning tasks. Moreover, its… See the full description on the dataset page: https://huggingface.co/datasets/dmis-lab/llama-3.1-medprm-reward-training-set.llama-3.1-awesome-chatgpt-prompts
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/llama-3.1-awesome-chatgpt-prompts.lm-eval-results-cognitivecomputations-dolphin-2.9-llama3-8b-private
Dataset Card for Evaluation run of cognitivecomputations/dolphin-2.9-llama3-8b
Dataset automatically created during the evaluation run of model cognitivecomputations/dolphin-2.9-llama3-8b
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 5 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-cognitivecomputations-dolphin-2.9-llama3-8b-private.lm-eval-results-ntnhan-Llama3-8B-MetaMath-private
Dataset Card for Evaluation run of ntnhan/Llama3-8B-MetaMath
Dataset automatically created during the evaluation run of model ntnhan/Llama3-8B-MetaMath
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-ntnhan-Llama3-8B-MetaMath-private.Yuma42__Llama3.1-SuperHawk-8B-details
Dataset Card for Evaluation run of Yuma42/Llama3.1-SuperHawk-8B
Dataset automatically created during the evaluation run of model Yuma42/Llama3.1-SuperHawk-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Yuma42__Llama3.1-SuperHawk-8B-details.akhadangi__Llama3.2.1B.0.01-First-details
Dataset Card for Evaluation run of akhadangi/Llama3.2.1B.0.01-First
Dataset automatically created during the evaluation run of model akhadangi/Llama3.2.1B.0.01-First
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/akhadangi__Llama3.2.1B.0.01-First-details.tenyx__Llama3-TenyxChat-70B-details
Dataset Card for Evaluation run of tenyx/Llama3-TenyxChat-70B
Dataset automatically created during the evaluation run of model tenyx/Llama3-TenyxChat-70B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tenyx__Llama3-TenyxChat-70B-details.akhadangi__Llama3.2.1B.0.01-Last-details
Dataset Card for Evaluation run of akhadangi/Llama3.2.1B.0.01-Last
Dataset automatically created during the evaluation run of model akhadangi/Llama3.2.1B.0.01-Last
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/akhadangi__Llama3.2.1B.0.01-Last-details.Danielbrdz__Barcenas-Llama3-8b-ORPO-details
Dataset Card for Evaluation run of Danielbrdz/Barcenas-Llama3-8b-ORPO
Dataset automatically created during the evaluation run of model Danielbrdz/Barcenas-Llama3-8b-ORPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Danielbrdz__Barcenas-Llama3-8b-ORPO-details.lm-eval-results-HenryJJ-llama3-8B-lima-private
Dataset Card for Evaluation run of HenryJJ/llama3-8B-lima
Dataset automatically created during the evaluation run of model HenryJJ/llama3-8B-lima
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-HenryJJ-llama3-8B-lima-private.netcat420__MFANN-llama3.1-abliterated-SLERP-v3-details
Dataset Card for Evaluation run of netcat420/MFANN-llama3.1-abliterated-SLERP-v3
Dataset automatically created during the evaluation run of model netcat420/MFANN-llama3.1-abliterated-SLERP-v3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/netcat420__MFANN-llama3.1-abliterated-SLERP-v3-details.mobiuslabsgmbh__DeepSeek-R1-ReDistill-Llama3-8B-v1.1-details
Dataset Card for Evaluation run of mobiuslabsgmbh/DeepSeek-R1-ReDistill-Llama3-8B-v1.1
Dataset automatically created during the evaluation run of model mobiuslabsgmbh/DeepSeek-R1-ReDistill-Llama3-8B-v1.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mobiuslabsgmbh__DeepSeek-R1-ReDistill-Llama3-8B-v1.1-details.MaziyarPanahi__calme-2.3-llama3.1-70b-details
Dataset Card for Evaluation run of MaziyarPanahi/calme-2.3-llama3.1-70b
Dataset automatically created during the evaluation run of model MaziyarPanahi/calme-2.3-llama3.1-70b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/MaziyarPanahi__calme-2.3-llama3.1-70b-details.sabersalehk__Llama3-001-300-details
Dataset Card for Evaluation run of sabersalehk/Llama3-001-300
Dataset automatically created during the evaluation run of model sabersalehk/Llama3-001-300
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sabersalehk__Llama3-001-300-details.sequelbox__Llama3.1-8B-PlumCode-details
Dataset Card for Evaluation run of sequelbox/Llama3.1-8B-PlumCode
Dataset automatically created during the evaluation run of model sequelbox/Llama3.1-8B-PlumCode
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sequelbox__Llama3.1-8B-PlumCode-details.nbeerbower__Llama3.1-Gutenberg-Doppel-70B-details
Dataset Card for Evaluation run of nbeerbower/Llama3.1-Gutenberg-Doppel-70B
Dataset automatically created during the evaluation run of model nbeerbower/Llama3.1-Gutenberg-Doppel-70B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/nbeerbower__Llama3.1-Gutenberg-Doppel-70B-details.nbeerbower__llama3.1-kartoffeldes-70B-details
Dataset Card for Evaluation run of nbeerbower/llama3.1-kartoffeldes-70B
Dataset automatically created during the evaluation run of model nbeerbower/llama3.1-kartoffeldes-70B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/nbeerbower__llama3.1-kartoffeldes-70B-details.agentlans__Llama3.1-Daredevilish-Instruct-details
Dataset Card for Evaluation run of agentlans/Llama3.1-Daredevilish-Instruct
Dataset automatically created during the evaluation run of model agentlans/Llama3.1-Daredevilish-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/agentlans__Llama3.1-Daredevilish-Instruct-details.KSU-HW-SEC__Llama3-70b-SVA-FT-1415-details
Dataset Card for Evaluation run of KSU-HW-SEC/Llama3-70b-SVA-FT-1415
Dataset automatically created during the evaluation run of model KSU-HW-SEC/Llama3-70b-SVA-FT-1415
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/KSU-HW-SEC__Llama3-70b-SVA-FT-1415-details.KSU-HW-SEC__Llama3-70b-SVA-FT-final-details
Dataset Card for Evaluation run of KSU-HW-SEC/Llama3-70b-SVA-FT-final
Dataset automatically created during the evaluation run of model KSU-HW-SEC/Llama3-70b-SVA-FT-final
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/KSU-HW-SEC__Llama3-70b-SVA-FT-final-details.
