CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01EleutherAI /hack-ignition-benchmark hack-ignition benchmark — data, v0.1.6 Training trajectories of reinforcement-learning runs on exploitable graders, for studying and predicting when RL comes to produce exploits. Each family is a set of GRPO runs over configurations of (start model, prompt, training set, grader / reward structure, recipe), with one or more seeds per configuration. Every family stores what its training logs contain — per-step exploit, task and reward rates, the item × step exploit record… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/hack-ignition-benchmark.tabular100K<n<1M1 likes1.1k downloads6d agoHugging Face02EleutherAI /pile_val_test The Pile: Validation and Test Splits This repo contains the validation and test splits of The Pile, an 825 GiB English text dataset designed for training large language models. Files File Split Size val.jsonl Validation 1.4 GB test.jsonl Test 1.3 GB Format Each line is a JSON object with two fields: {"text": "The document text...", "meta": {"pile_set_name": "Pile-CC"}} The meta.pile_set_name field indicates which of the 22 constituent… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/pile_val_test.texttext-generation100K<n<1M0 likes1k downloads7mo agoHugging Face03open-llm-leaderboard /EleutherAI__gpt-j-6b-detailsgated Dataset Card for Evaluation run of EleutherAI/gpt-j-6b Dataset automatically created during the evaluation run of model EleutherAI/gpt-j-6b The dataset is composed of 111 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 12 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__gpt-j-6b-details.tabular10K<n<100K0 likes79 downloads2y agoHugging Face04open-llm-leaderboard /EleutherAI__pythia-1.4b-detailsgated Dataset Card for Evaluation run of EleutherAI/pythia-1.4b Dataset automatically created during the evaluation run of model EleutherAI/pythia-1.4b The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__pythia-1.4b-details.tabular10K<n<100K0 likes70 downloads2y agoHugging Face05open-llm-leaderboard /EleutherAI__gpt-neox-20b-detailsgated Dataset Card for Evaluation run of EleutherAI/gpt-neox-20b Dataset automatically created during the evaluation run of model EleutherAI/gpt-neox-20b The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__gpt-neox-20b-details.tabular10K<n<100K0 likes60 downloads2y agoHugging Face06open-llm-leaderboard /EleutherAI__pythia-160m-detailsgated Dataset Card for Evaluation run of EleutherAI/pythia-160m Dataset automatically created during the evaluation run of model EleutherAI/pythia-160m The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__pythia-160m-details.tabular10K<n<100K0 likes56 downloads2y agoHugging Face07EleutherAI /muInstructμInstruct is a dataset of 1600 instruction-response pairs collected from highly-rated Stack Exchange answers, the Khan Academy subset of AMPS, and the MATH training set. All training examples are valid Markdown have been manually reviewed by a human for quality. The μInstruct dataset is most useful when mixed in with larger instruction or chat datasets, such as OpenHermes. Because μInstruct is especially high-quality, you may consider oversampling it in your training mixture. μInstruct was… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/muInstruct.text1K<n<10K3 likes52 downloads3y agoHugging Face08open-llm-leaderboard /EleutherAI__pythia-12b-detailsgated Dataset Card for Evaluation run of EleutherAI/pythia-12b Dataset automatically created during the evaluation run of model EleutherAI/pythia-12b The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__pythia-12b-details.tabular10K<n<100K0 likes49 downloads2y agoHugging Face09open-llm-leaderboard /EleutherAI__gpt-neo-125m-detailsgated Dataset Card for Evaluation run of EleutherAI/gpt-neo-125m Dataset automatically created during the evaluation run of model EleutherAI/gpt-neo-125m The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__gpt-neo-125m-details.tabular10K<n<100K0 likes48 downloads2y agoHugging Face10open-llm-leaderboard /EleutherAI__gpt-neo-2.7B-detailsgated Dataset Card for Evaluation run of EleutherAI/gpt-neo-2.7B Dataset automatically created during the evaluation run of model EleutherAI/gpt-neo-2.7B The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__gpt-neo-2.7B-details.tabular10K<n<100K0 likes46 downloads2y agoHugging Face11open-llm-leaderboard /EleutherAI__pythia-6.9b-detailsgated Dataset Card for Evaluation run of EleutherAI/pythia-6.9b Dataset automatically created during the evaluation run of model EleutherAI/pythia-6.9b The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__pythia-6.9b-details.tabular10K<n<100K0 likes40 downloads2y agoHugging Face12open-llm-leaderboard /EleutherAI__pythia-410m-detailsgated Dataset Card for Evaluation run of EleutherAI/pythia-410m Dataset automatically created during the evaluation run of model EleutherAI/pythia-410m The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__pythia-410m-details.tabular10K<n<100K0 likes39 downloads2y agoHugging Face13open-llm-leaderboard /EleutherAI__pythia-2.8b-detailsgated Dataset Card for Evaluation run of EleutherAI/pythia-2.8b Dataset automatically created during the evaluation run of model EleutherAI/pythia-2.8b The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__pythia-2.8b-details.tabular10K<n<100K0 likes39 downloads2y agoHugging Face14open-llm-leaderboard /EleutherAI__gpt-neo-1.3B-detailsgated Dataset Card for Evaluation run of EleutherAI/gpt-neo-1.3B Dataset automatically created during the evaluation run of model EleutherAI/gpt-neo-1.3B The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__gpt-neo-1.3B-details.tabular10K<n<100K0 likes36 downloads2y agoHugging Face15open-llm-leaderboard /EleutherAI__pythia-1b-detailsgated Dataset Card for Evaluation run of EleutherAI/pythia-1b Dataset automatically created during the evaluation run of model EleutherAI/pythia-1b The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__pythia-1b-details.tabular10K<n<100K0 likes34 downloads2y agoHugging Face16EleutherAI /lm-eval-EleutherAI_deep-ignorance-unfiltered Dataset Card for Evaluation run of EleutherAI/deep-ignorance-unfiltered Dataset automatically created during the evaluation run of model EleutherAI/deep-ignorance-unfiltered The dataset is composed of 1 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/lm-eval-EleutherAI_deep-ignorance-unfiltered.tabular1K<n<10K0 likes22 downloads11mo agoHugging Face17EleutherAI /lm-eval-EleutherAI_tampered-deep-ignorance-random-init-fp-adversarial-20251104_051748 Dataset Card for Evaluation run of EleutherAI/tampered-deep-ignorance-random-init-fp-adversarial-20251104_051748 Dataset automatically created during the evaluation run of model EleutherAI/tampered-deep-ignorance-random-init-fp-adversarial-20251104_051748 The dataset is composed of 2 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 11 run(s). Each run can be found as a specific split in each configuration, the split being… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/lm-eval-EleutherAI_tampered-deep-ignorance-random-init-fp-adversarial-20251104_051748.tabular1K<n<10K0 likes19 downloads11mo agoHugging Face18EleutherAI /youtube-cctext1M<n<10M2 likes13 downloads1y agoHugging Face19EleutherAI /lm-eval-EleutherAI_deep-ignorance-random-init Dataset Card for Evaluation run of EleutherAI/deep-ignorance-random-init Dataset automatically created during the evaluation run of model EleutherAI/deep-ignorance-random-init The dataset is composed of 0 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/lm-eval-EleutherAI_deep-ignorance-random-init.tabular1K<n<10K0 likes9 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.