CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01EleutherAI /headqatabular10K<n<100K0 likes109k downloads5mo agoHugging Face02EleutherAI /filtering-pretraining-mix-arrow-formattabular100M<n<1B0 likes3k downloads2y agoHugging Face03EleutherAI /dclm-dedup_20250227-004105tabular100M<n<1B1 likes2.9k downloads2y agoHugging Face04EleutherAI /hack-ignition-benchmark hack-ignition benchmark — data, v0.1.6 Training trajectories of reinforcement-learning runs on exploitable graders, for studying and predicting when RL comes to produce exploits. Each family is a set of GRPO runs over configurations of (start model, prompt, training set, grader / reward structure, recipe), with one or more seeds per configuration. Every family stores what its training logs contain — per-step exploit, task and reward rates, the item × step exploit record… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/hack-ignition-benchmark.tabular100K<n<1M1 likes1.1k downloads5d agoHugging Face05EleutherAI /pythia-memorized-evals Pythia Memorized Evals This dataset contains the results of memorization evaluations for all Pythia models. For each model, the dataset lists every training sequence that the fully trained model has memorized. A training sequence is considered memorized if, when prompted with the first 32 tokens of the sequence, the model's greedy continuation exactly matches the next 32 tokens. This is evaluated over all ~146M training sequences in the Pile. This dataset was generated for the paper… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/pythia-memorized-evals.tabular10M<n<100M4 likes989 downloads7mo agoHugging Face06EleutherAI /LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1007 Retrain bank: WikiText-2 / GPT-2, random halves, seed 1007 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1007.tabular10K<n<100K0 likes957 downloads27d agoHugging Face07EleutherAI /LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1004 Retrain bank: WikiText-2 / GPT-2, random halves, seed 1004 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1004.tabular10K<n<100K0 likes910 downloads27d agoHugging Face08EleutherAI /LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1006 Retrain bank: WikiText-2 / GPT-2, random halves, seed 1006 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1006.tabular10K<n<100K0 likes893 downloads27d agoHugging Face09EleutherAI /LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1008 Retrain bank: WikiText-2 / GPT-2, random halves, seed 1008 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1008.tabular10K<n<100K0 likes872 downloads27d agoHugging Face10EleutherAI /deep-ignorance-pretraining-mix Deep Ignorance Model Suite We explore an intuitive yet understudied question: Can we prevent LLMs from learning unsafe technical capabilities (such as CBRN) by filtering out enough of the relevant pretraining data before we begin training a model? Research into this question resulted in the Deep Ignorance Suite. In our experimental setup, we find that filtering pretraining data prevents undesirable knowledge, doesn't sacrifice general performance, and results in models that are… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/deep-ignorance-pretraining-mix.tabular100M<n<1B4 likes867 downloads1y agoHugging Face11EleutherAI /LDS-retrain-bank-muon-N16k-bs16tabular1K<n<10K0 likes781 downloads1mo agoHugging Face12EleutherAI /LDS-retrain-bank-adamw-N4k-bs256 Retrain bank: plan_adam_eps1e17_4k_bs256 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on the same 4,000-document corpus with a different random 1% (40 documents) held out, from the same seed and the same data order as the base model in retrained/base. Retraining is deterministic within one environment, so the models differ only by the documents removed. That is the expensive part of any leave-k-out… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-N4k-bs256.tabular1K<n<10K0 likes760 downloads1mo agoHugging Face13EleutherAI /dclm-dedup-25BThe first 25,000,003,108 tokens of Zyphra/dclm-dedup. Tokenized using the NeoX tokenizer (EleutherAI/gpt-neox-20b). tabular10M<n<100M2 likes691 downloads2y agoHugging Face14EleutherAI /LDS-retrain-bank-muon-N16k-bs128tabular1K<n<10K0 likes691 downloads1mo agoHugging Face15EleutherAI /LDS-retrain-bank-adamw-N16k-bs256 Retrain bank: sm_adamw_eps1e17_16k_bs256 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on the same 16,000-document corpus with a different random 1% (160 documents) held out, from the same seed and the same data order as the base model in retrained/base. Retraining is deterministic within one environment, so the models differ only by the documents removed. That is the expensive part of any leave-k-out… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-N16k-bs256.tabular1K<n<10K0 likes680 downloads1mo agoHugging Face16EleutherAI /LDS-retrain-bank-muon-N32k-bs256tabular1K<n<10K0 likes676 downloads1mo agoHugging Face17EleutherAI /LDS-retrain-bank-adamw-N16k-bs256-scale0.25 Retrain bank: plan_adam_eps1e17_16k_scale0.25 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on the same 16,000-document corpus with a different random 1% (160 documents) held out, from the same seed and the same data order as the base model in retrained/base. Retraining is deterministic within one environment, so the models differ only by the documents removed. That is the expensive part of any leave-k-out… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-N16k-bs256-scale0.25.tabular1K<n<10K0 likes666 downloads1mo agoHugging Face18EleutherAI /LDS-retrain-bank-muon-N16k-bs256tabular1K<n<10K0 likes595 downloads1mo agoHugging Face19EleutherAI /LDS-retrain-bank-adamw-N16k-bs256-ep4tabular1K<n<10K0 likes507 downloads1mo agoHugging Face20EleutherAI /LDS-retrain-bank-adamw-N32k-bs256tabular1K<n<10K0 likes492 downloads1mo agoHugging Face21EleutherAI /LDS-retrain-bank-adamw-N16k-bs256-gpt2-mediumtabular1K<n<10K0 likes487 downloads1mo agoHugging Face22EleutherAI /PARTIAL_LDS-retrain-bank-gpt2medium-16k-bs32tabular1K<n<10K0 likes479 downloads24d agoHugging Face23EleutherAI /LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1005 Retrain bank: WikiText-2 / GPT-2, random halves, seed 1005 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1005.tabular10K<n<100K0 likes469 downloads27d agoHugging Face24EleutherAI /PARTIAL_LDS-retrain-bank-london16k-bs256-adamwtabular1K<n<10K0 likes452 downloads24d agoHugging Face25EleutherAI /LDS-retrain-bank-adamw-N8k-bs256 Retrain bank: plan_adam_eps1e17_8k_bs256 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on the same 8,000-document corpus with a different random 1% (80 documents) held out, from the same seed and the same data order as the base model in retrained/base. Retraining is deterministic within one environment, so the models differ only by the documents removed. That is the expensive part of any leave-k-out… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-N8k-bs256.tabular1K<n<10K0 likes375 downloads1mo agoHugging Face26EleutherAI /LDS-retrain-bank-muon-N4k-bs256 Retrain bank: plan_muon_eps1e17_4k_bs256 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on the same 4,000-document corpus with a different random 1% (40 documents) held out, from the same seed and the same data order as the base model in retrained/base. Retraining is deterministic within one environment, so the models differ only by the documents removed. That is the expensive part of any leave-k-out… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-muon-N4k-bs256.tabular1K<n<10K0 likes359 downloads1mo agoHugging Face27EleutherAI /transformer-reasoning-bios-dataset-250000tabular100M<n<1B0 likes344 downloads2y agoHugging Face28EleutherAI /LDS-retrain-bank-muon-N8k-bs256 Retrain bank: plan_muon_eps1e17_8k_bs256 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on the same 8,000-document corpus with a different random 1% (80 documents) held out, from the same seed and the same data order as the base model in retrained/base. Retraining is deterministic within one environment, so the models differ only by the documents removed. That is the expensive part of any leave-k-out… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-muon-N8k-bs256.tabular1K<n<10K0 likes338 downloads1mo agoHugging Face29EleutherAI /transformer-reasoning-bios-dataset-25000tabular10M<n<100M0 likes329 downloads2y agoHugging Face30EleutherAI /LDS-retrain-bank-adamw-N16k-bs64tabular1K<n<10K0 likes305 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.