datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fineweb10B-gpt2
fineweb10B-gpt2
This repo contains the GPT-2 tokens for fineweb10B, just as would be generated by https://github.com/KellerJordan/modded-nanogpt/tree/master (or llm.c).
You can download from this repo instead of re-tokenizing to save a couple hours of setup on a new machine.
fineweb100B-gpt2fineweb-edu-100B-gpt2-token-shardsFineWeb Edu 100B dataset tokenized with GPT-2 tokenizer using the code in llm.c repo.
fineweb-edu-gpt2Skylion007-openwebtext-tokenizer-gpt2fineweb-edu-sample-10BT-gpt2tokenizedxent-tasks-gpt2monology-pile-uncopyrighted-tokenizer-gpt2openwebtext-gpt2gpt2_model_acts_openwebtextchess-gpt2-hiddenstates-768Is this working?
sae-monology-pile-uncopyrighted-tokenizer-gpt2dolma-blend-gpt2starcoderdata-gpt2sae-skeskinen-TinyStories-hf-validation-tokenizer-gpt2_playclt_gpt2_tokenized_control
Fresh multilingual GPT-2 CLT control data
Sequential, unshuffled control sample for CLT null experiments. For each language,
complete source documents were tokenized with CausalNLP/gpt2-hf_multilingual-20 at revision
0afbb31b2db3f394270d42d6a4cb7f8fceeca3d8. The first 100,000,000 tokenizer tokens were discarded
(including the complete document that crossed the threshold), after which complete
documents were retained until at least 100,000,000 tokens were collected.
Data are… See the full description on the dataset page: https://huggingface.co/datasets/abir-hr196/clt_gpt2_tokenized_control.Skylion007-openwebtext-tokenizer-gpt2-128the-stack-dedup-python-filtered-docstrings-gpt2openwwebtext-gpt2-50257-standardchess-gpt2-hiddenstates-512
Dataset Card for Chess GPT-2 Hidden States 512
Dataset Summary
This dataset contains 120k hidden state vectors from forward passes through a GPT-2 model trained on UCI chess move sequences.
The model has 8 layers, each with 8 attention heads, and a hidden state size of 512.
The dataset was generated by performing one forward pass for each UCI move sequence in the "austindavis/lichess_uci" dataset,
specifically the "train" split from the "201301-moves" configuration.… See the full description on the dataset page: https://huggingface.co/datasets/austindavis/chess-gpt2-hiddenstates-512.wikilarge-graded-gpt2toneizergpt2small_full_training_datafinewebedu10B-gpt2Skylion007-openwebtext-gpt2-1024Tokenizer: gpt2dataset: '''Skylion007/openwebtext'''context_size : 1024
wikitext-103-raw-v1_gpt2-20k
Dataset Card for "wikitext-103-raw-v1_gpt2-20k"
More Information needed
finewebedu10BT-tokenized-gpt2
FineWebEdu
This is FineWebEdu sample10T subset dataset tokenized into shards
with GPT2 tokenizer.
LDS-retrain-bank-adamw-N16k-bs256-gpt2-mediummonology-pile-uncopyrighted-gpt2-1024Tokenizer: gpt2dataset: '''monology-pile-uncopyrighted'''context_size : 1024
gpt2_tokenized_concatenated_openwebtextcc100-en-test-tokenized-gpt2
