CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kjj0 /fineweb10B-gpt2 fineweb10B-gpt2 This repo contains the GPT-2 tokens for fineweb10B, just as would be generated by https://github.com/KellerJordan/modded-nanogpt/tree/master (or llm.c). You can download from this repo instead of re-tokenizing to save a couple hours of setup on a new machine. 11 likes12k downloads2y agoHugging Face02kjj0 /fineweb100B-gpt21 likes7.9k downloads2y agoHugging Face03karpathy /fineweb-edu-100B-gpt2-token-shardsFineWeb Edu 100B dataset tokenized with GPT-2 tokenizer using the code in llm.c repo. 8 likes5k downloads2y agoHugging Face04kushalt /fineweb-edu-gpt2tabular10M<n<100M0 likes2.6k downloads8mo agoHugging Face05apollo-research /Skylion007-openwebtext-tokenizer-gpt21M<n<10M3 likes2.4k downloads3y agoHugging Face06alkibijad /fineweb-edu-sample-10BT-gpt2tokenized1 likes1.9k downloads2y agoHugging Face07marktas /xent-tasks-gpt2textn<1K0 likes1.6k downloads2y agoHugging Face08apollo-research /monology-pile-uncopyrighted-tokenizer-gpt210M<n<100M1 likes1.3k downloads3y agoHugging Face09chanind /openwebtext-gpt21M<n<10M0 likes1.3k downloads2y agoHugging Face10hanspeterlyngsoeraaschoujensen /gpt2_model_acts_openwebtexttextn<1K0 likes937 downloads2y agoHugging Face11austindavis /chess-gpt2-hiddenstates-768Is this working? tabular1M<n<10M0 likes931 downloads1y agoHugging Face12alancooney /sae-monology-pile-uncopyrighted-tokenizer-gpt210M<n<100M2 likes836 downloads3y agoHugging Face13varunneal /dolma-blend-gpt20 likes773 downloads4mo agoHugging Face14ytzi /starcoderdata-gpt2tabular10M<n<100M0 likes764 downloads2y agoHugging Face15apollo-research /sae-skeskinen-TinyStories-hf-validation-tokenizer-gpt2_playtext10K<n<100K0 likes741 downloads3y agoHugging Face16abir-hr196 /clt_gpt2_tokenized_control Fresh multilingual GPT-2 CLT control data Sequential, unshuffled control sample for CLT null experiments. For each language, complete source documents were tokenized with CausalNLP/gpt2-hf_multilingual-20 at revision 0afbb31b2db3f394270d42d6a4cb7f8fceeca3d8. The first 100,000,000 tokenizer tokens were discarded (including the complete document that crossed the threshold), after which complete documents were retained until at least 100,000,000 tokens were collected. Data are… See the full description on the dataset page: https://huggingface.co/datasets/abir-hr196/clt_gpt2_tokenized_control.texttext-generation100K<n<1M0 likes726 downloads2mo agoHugging Face17pccl-org /Skylion007-openwebtext-tokenizer-gpt2-12810M<n<100M0 likes666 downloads2y agoHugging Face18ytzi /the-stack-dedup-python-filtered-docstrings-gpt2tabular10M<n<100M0 likes557 downloads2y agoHugging Face19anyasims /openwwebtext-gpt2-50257-standard1M<n<10M0 likes531 downloads2y agoHugging Face20austindavis /chess-gpt2-hiddenstates-512 Dataset Card for Chess GPT-2 Hidden States 512 Dataset Summary This dataset contains 120k hidden state vectors from forward passes through a GPT-2 model trained on UCI chess move sequences. The model has 8 layers, each with 8 attention heads, and a hidden state size of 512. The dataset was generated by performing one forward pass for each UCI move sequence in the "austindavis/lichess_uci" dataset, specifically the "train" split from the "201301-moves" configuration.… See the full description on the dataset page: https://huggingface.co/datasets/austindavis/chess-gpt2-hiddenstates-512.tabularother1M<n<10M0 likes523 downloads2y agoHugging Face21williamplacroix /wikilarge-graded-gpt2toneizer100K<n<1M0 likes513 downloads2y agoHugging Face22CausalNLP /gpt2small_full_training_datatext1M<n<10M0 likes503 downloads1y agoHugging Face23kjj0 /finewebedu10B-gpt20 likes500 downloads2y agoHugging Face24heyunzhenwhat /Skylion007-openwebtext-gpt2-1024Tokenizer: gpt2dataset: '''Skylion007/openwebtext'''context_size : 1024 1M<n<10M0 likes471 downloads2y agoHugging Face25pietrolesci /wikitext-103-raw-v1_gpt2-20k Dataset Card for "wikitext-103-raw-v1_gpt2-20k" More Information needed tabular1M<n<10M0 likes452 downloads3y agoHugging Face26JpChi /finewebedu10BT-tokenized-gpt2 FineWebEdu This is FineWebEdu sample10T subset dataset tokenized into shards with GPT2 tokenizer. 0 likes446 downloads2y agoHugging Face27EleutherAI /LDS-retrain-bank-adamw-N16k-bs256-gpt2-mediumtabular1K<n<10K0 likes445 downloads28d agoHugging Face28heyunzhenwhat /monology-pile-uncopyrighted-gpt2-1024Tokenizer: gpt2dataset: '''monology-pile-uncopyrighted'''context_size : 1024 100M<n<1B2 likes436 downloads2y agoHugging Face29chrisjob1021 /gpt2_tokenized_concatenated_openwebtext1M<n<10M0 likes436 downloads1y agoHugging Face30Geonwoohong /cc100-en-test-tokenized-gpt21K<n<10K0 likes416 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.