datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hack-ignition-benchmark
hack-ignition benchmark — data, v0.1.6
Training trajectories of reinforcement-learning runs on exploitable graders, for studying and predicting when RL
comes to produce exploits. Each family is a set of GRPO runs over configurations of (start model, prompt,
training set, grader / reward structure, recipe), with one or more seeds per configuration. Every family stores
what its training logs contain — per-step exploit, task and reward rates, the item × step exploit record… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/hack-ignition-benchmark.pile_val_test
The Pile: Validation and Test Splits
This repo contains the validation and test splits of The Pile, an 825 GiB English text dataset designed for training large language models.
Files
File
Split
Size
val.jsonl
Validation
1.4 GB
test.jsonl
Test
1.3 GB
Format
Each line is a JSON object with two fields:
{"text": "The document text...", "meta": {"pile_set_name": "Pile-CC"}}
The meta.pile_set_name field indicates which of the 22 constituent… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/pile_val_test.EleutherAI__gpt-j-6b-details
Dataset Card for Evaluation run of EleutherAI/gpt-j-6b
Dataset automatically created during the evaluation run of model EleutherAI/gpt-j-6b
The dataset is composed of 111 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 12 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__gpt-j-6b-details.EleutherAI__pythia-1.4b-details
Dataset Card for Evaluation run of EleutherAI/pythia-1.4b
Dataset automatically created during the evaluation run of model EleutherAI/pythia-1.4b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__pythia-1.4b-details.EleutherAI__gpt-neox-20b-details
Dataset Card for Evaluation run of EleutherAI/gpt-neox-20b
Dataset automatically created during the evaluation run of model EleutherAI/gpt-neox-20b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__gpt-neox-20b-details.EleutherAI__pythia-160m-details
Dataset Card for Evaluation run of EleutherAI/pythia-160m
Dataset automatically created during the evaluation run of model EleutherAI/pythia-160m
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__pythia-160m-details.muInstructμInstruct is a dataset of 1600 instruction-response pairs collected from highly-rated Stack Exchange answers, the Khan Academy subset of AMPS, and the MATH training set. All training examples are valid Markdown have been manually reviewed by a human for quality.
The μInstruct dataset is most useful when mixed in with larger instruction or chat datasets, such as OpenHermes. Because μInstruct is especially high-quality, you may consider oversampling it in your training mixture.
μInstruct was… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/muInstruct.EleutherAI__pythia-12b-details
Dataset Card for Evaluation run of EleutherAI/pythia-12b
Dataset automatically created during the evaluation run of model EleutherAI/pythia-12b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__pythia-12b-details.EleutherAI__gpt-neo-125m-details
Dataset Card for Evaluation run of EleutherAI/gpt-neo-125m
Dataset automatically created during the evaluation run of model EleutherAI/gpt-neo-125m
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__gpt-neo-125m-details.EleutherAI__gpt-neo-2.7B-details
Dataset Card for Evaluation run of EleutherAI/gpt-neo-2.7B
Dataset automatically created during the evaluation run of model EleutherAI/gpt-neo-2.7B
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__gpt-neo-2.7B-details.EleutherAI__pythia-6.9b-details
Dataset Card for Evaluation run of EleutherAI/pythia-6.9b
Dataset automatically created during the evaluation run of model EleutherAI/pythia-6.9b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__pythia-6.9b-details.EleutherAI__pythia-410m-details
Dataset Card for Evaluation run of EleutherAI/pythia-410m
Dataset automatically created during the evaluation run of model EleutherAI/pythia-410m
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__pythia-410m-details.EleutherAI__pythia-2.8b-details
Dataset Card for Evaluation run of EleutherAI/pythia-2.8b
Dataset automatically created during the evaluation run of model EleutherAI/pythia-2.8b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__pythia-2.8b-details.EleutherAI__gpt-neo-1.3B-details
Dataset Card for Evaluation run of EleutherAI/gpt-neo-1.3B
Dataset automatically created during the evaluation run of model EleutherAI/gpt-neo-1.3B
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__gpt-neo-1.3B-details.EleutherAI__pythia-1b-details
Dataset Card for Evaluation run of EleutherAI/pythia-1b
Dataset automatically created during the evaluation run of model EleutherAI/pythia-1b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__pythia-1b-details.lm-eval-EleutherAI_deep-ignorance-unfiltered
Dataset Card for Evaluation run of EleutherAI/deep-ignorance-unfiltered
Dataset automatically created during the evaluation run of model EleutherAI/deep-ignorance-unfiltered
The dataset is composed of 1 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/lm-eval-EleutherAI_deep-ignorance-unfiltered.lm-eval-EleutherAI_tampered-deep-ignorance-random-init-fp-adversarial-20251104_051748
Dataset Card for Evaluation run of EleutherAI/tampered-deep-ignorance-random-init-fp-adversarial-20251104_051748
Dataset automatically created during the evaluation run of model EleutherAI/tampered-deep-ignorance-random-init-fp-adversarial-20251104_051748
The dataset is composed of 2 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 11 run(s). Each run can be found as a specific split in each configuration, the split being… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/lm-eval-EleutherAI_tampered-deep-ignorance-random-init-fp-adversarial-20251104_051748.youtube-cclm-eval-EleutherAI_deep-ignorance-random-init
Dataset Card for Evaluation run of EleutherAI/deep-ignorance-random-init
Dataset automatically created during the evaluation run of model EleutherAI/deep-ignorance-random-init
The dataset is composed of 0 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/lm-eval-EleutherAI_deep-ignorance-random-init.
