datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pythia-deduped-stats-rawThis dataset has been created as an artefact of the paper Causal Estimation of Memorisation Profiles (Lesci et al., 2024).
More info about this dataset in the related collection Memorisation-Profiles.
Collection of data statistics computed using the intermediate checkpoints (step0, step1000, ..., step143k) of all Pythia deduped versions.
This folder contains the model evaluations (or "stats") for each model size included in the study. This is the "raw" version where we have stats at the… See the full description on the dataset page: https://huggingface.co/datasets/pietrolesci/pythia-deduped-stats-raw.data_inference_pythia_6_9bpile-standard-pythia-preshuffledpythia-training-metrics Dataset for storing training metrics of pythia modelspythia_pile_idxmapspythia-semantic-memorization-perplexities
Dataset Card for "pythia-semantic-memorization-perplexities"
More Information needed
pythia_deduped_pile_idxmapspythia-training-evalspythia-massive-activations
Hidden Dynamics of Massive Activations in Transformer Training
Dataset Description
This dataset contains comprehensive analysis data for the paper "Hidden Dynamics of Massive Activations in Transformer Training". It provides detailed measurements and mathematical characterizations of massive activation emergence patterns across the Pythia model family during training.
Massive activations are scalar values in transformer hidden states that achieve values orders of… See the full description on the dataset page: https://huggingface.co/datasets/Aimpoint-Digital/pythia-massive-activations.processed_pythia_dataset_2048
Dataset Card for "processed_pythia_dataset_2048"
More Information needed
pythia-training-metrics-40m-qk-layernorm Dataset for storing training metrics of pythia modelspythia-memorized-evals
Pythia Memorized Evals
This dataset contains the results of memorization evaluations for all Pythia models. For each model, the dataset lists every training sequence that the fully trained model has memorized.
A training sequence is considered memorized if, when prompted with the first 32 tokens of the sequence, the model's greedy continuation exactly matches the next 32 tokens. This is evaluated over all ~146M training sequences in the Pile.
This dataset was generated for the paper… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/pythia-memorized-evals.details_EleutherAI__pythia-12b
Dataset Card for Evaluation run of EleutherAI/pythia-12b
Dataset Summary
Dataset automatically created during the evaluation run of model EleutherAI/pythia-12b on the Open LLM Leaderboard.
The dataset is composed of 122 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_EleutherAI__pythia-12b.pile-deduped-pythia-preshuffledpythia-training-metrics-40m-base Dataset for storing training metrics of pythia modelsall-the-news-2-pythia-tfidf-wordleveldetails_EleutherAI__pythia-12b-deduped
Dataset Card for Evaluation run of EleutherAI/pythia-12b-deduped
Dataset Summary
Dataset automatically created during the evaluation run of model EleutherAI/pythia-12b-deduped on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_EleutherAI__pythia-12b-deduped.all-the-news-2-pythia-tfidf-topic-stratified-v1-articlesoptim_policy_pretrain-pythia-160m_lr0.0001_bs24_wp1_wd0.01_ep0_cp35k-mergedpubmed-2019-pythia-word-tfidf-pubmedqa-clean-articlespubmed-2019-pythia-word-tfidf-pubmedqa-clean-val-sequencespile-deduped-pythia-preshuffledThis dataset contains the fully prepared data, which has been tokenized and pre-shuffled, used to train the Pythia (deduplicated) models.
You can find these models under the EleutherAI organisation, and they are also listed in my Memorisation-Profiles collection.
This data is the same as the one found in EleutherAI/pile-deduped-pythia-preshuffled,
but it is presented in a more manageable format. Instead of using the Megatron format used by the GPT-NeoX library, I have stored the data in a… See the full description on the dataset page: https://huggingface.co/datasets/pietrolesci/pile-deduped-pythia-preshuffled.all-the-news-2-pythia-tfidf-invfreq-topic-stratified-v1-articlespubmed-2019-pythia-word-tfidf-invfreq-pubmedqa-clean-articlespubmed-2019-pythia-word-tfidf-invfreq-pubmedqa-clean-val-sequencesrp1t-sample-tokenized-pythia-70mdetails_EleutherAI__pythia-6.9b-deduped
Dataset Card for Evaluation run of EleutherAI/pythia-6.9b-deduped
Dataset Summary
Dataset automatically created during the evaluation run of model EleutherAI/pythia-6.9b-deduped on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_EleutherAI__pythia-6.9b-deduped.pythia_clt_pretrain_data_tokenizeddetails_EleutherAI__pythia-70m-deduped
Dataset Card for Evaluation run of EleutherAI/pythia-70m-deduped
Dataset Summary
Dataset automatically created during the evaluation run of model EleutherAI/pythia-70m-deduped on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_EleutherAI__pythia-70m-deduped.details_OpenAssistant__pythia-12b-pre-v8-12.5k-steps
Dataset Card for Evaluation run of OpenAssistant/pythia-12b-pre-v8-12.5k-steps
Dataset Summary
Dataset automatically created during the evaluation run of model OpenAssistant/pythia-12b-pre-v8-12.5k-steps on the Open LLM Leaderboard.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_OpenAssistant__pythia-12b-pre-v8-12.5k-steps.
