datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fineweb-edu-gpt2chess-gpt2-hiddenstates-768Is this working?
starcoderdata-gpt2the-stack-dedup-python-filtered-docstrings-gpt2chess-gpt2-hiddenstates-512
Dataset Card for Chess GPT-2 Hidden States 512
Dataset Summary
This dataset contains 120k hidden state vectors from forward passes through a GPT-2 model trained on UCI chess move sequences.
The model has 8 layers, each with 8 attention heads, and a hidden state size of 512.
The dataset was generated by performing one forward pass for each UCI move sequence in the "austindavis/lichess_uci" dataset,
specifically the "train" split from the "201301-moves" configuration.… See the full description on the dataset page: https://huggingface.co/datasets/austindavis/chess-gpt2-hiddenstates-512.wikitext-103-raw-v1_gpt2-20k
Dataset Card for "wikitext-103-raw-v1_gpt2-20k"
More Information needed
LDS-retrain-bank-adamw-N16k-bs256-gpt2-mediumPARTIAL_LDS-retrain-bank-gpt2medium-16k-bs32PrismAI_v2-encoded-gpt2the-stack-dedup-python-filtered-dec_gen_async-gpt2the-stack-dedup-python-filtered-dec_gen_async-gpt2details_gpt2
Dataset Card for Evaluation run of gpt2
Dataset automatically created during the evaluation run of model gpt2.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration "results" store all the… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_gpt2.gpt2-ioi-mixed-circuitgpt2-pretrain-corpusM4-encoded-gpt2RAID_none-encoded-gpt2bergson-wikitext-gpt2-leaderboard-bank
bergson leaderboard: retrain banks, scores and LDS/QLD results (WikiText GPT-2)
Everything behind the numbers on the bergson leaderboard,
for the model at EleutherAI/bergson-wikitext-gpt2-leaderboard.
path
what it is
bank/
the LDS ground truth: 100 random leave-1%-out subsets of the 4,608 training chunks (subsets.json) and each subset's measured loss change on the 50 test queries (validation.csv)
random/retrained/{base,subset_0..99}
the retrained models themselves… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/bergson-wikitext-gpt2-leaderboard-bank.gpt2-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
gpt2-tokenizer-corpusgpt2-outputs
Dataset Card for "gpt2-outputs"
More Information needed
the-stack-dedup-python-filtered-gpt2the-stack-dedup-gpt2autointerp-gpt2-korean-newreward-bench-gpt2-normalsummarize_from_feedback_oai_preprocessing_gpt2_153
Dataset Card for "summarize_from_feedback_oai_preprocessing_gpt2_153"
More Information needed
autointerp-gpt2-multilingual-90
GPT2 Multilingual 20% AutoInterp Features
This dataset contains feature interpretations for GPT2 Multilingual model with 20% sparsity.
Structure
data/layer0.parquet - Features for layer 0
data/layer1.parquet - Features for layer 1
...
data/layer11.parquet - Features for layer 11
Each parquet file contains feature interpretation data including:
Feature activations
Top examples
Interpretations
And other feature analysis data
Usage
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/flodraye/autointerp-gpt2-multilingual-90.summarize_from_feedback_tldr_3_filtered_oai_preprocessing_gpt2_48
TL;DR SFT Dataset for OpenAI's Summarize from Feedback task
The dataset is directly taken from https://github.com/openai/summarize-from-feedback/tree/700967448d10004279f138666442bf1497d0e705#reddit-tldr-dataset
These columns are taken directly from the aforementioned dataset:
id: unique identifier for the post
subreddit: subreddit the post was taken from
title: title of the post
post: body of the post
summary: summary of the post
reference_response: reference response for the post… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/summarize_from_feedback_tldr_3_filtered_oai_preprocessing_gpt2_48.mixed-pretrain-10b-gpt2
Mixed Pretraining 10B (GPT-2 BPE)
A 10-billion-token pretraining dataset, GPT-2 BPE tokenized, assembled as a
diverse mix of web text, books, Wikipedia, code, academic papers, Q&A and
instruction-formatted conversations.
Built to train a ~500M parameter from-scratch GPT-2-style transformer (see
juliannunezb/transformer-lm-500m).
Mix
Source
Mix %
Tokens
Notes
fineweb
40.4%
4,039,999,700
reused from kjj0/fineweb10B-gpt2
fineweb_edu
15.2%
1,514,999,900
reused… See the full description on the dataset page: https://huggingface.co/datasets/juliannunezb/mixed-pretrain-10b-gpt2.reward-bench-gpt2-yes-nogpt2-winogrande_base
Dataset Card for "gpt2-winogrande_base"
More Information needed
