CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kjj0 /fineweb10B-gpt2 fineweb10B-gpt2 This repo contains the GPT-2 tokens for fineweb10B, just as would be generated by https://github.com/KellerJordan/modded-nanogpt/tree/master (or llm.c). You can download from this repo instead of re-tokenizing to save a couple hours of setup on a new machine. 11 likes12k downloads2y agoHugging Face02kjj0 /fineweb100B-gpt21 likes8.1k downloads2y agoHugging Face03karpathy /fineweb-edu-100B-gpt2-token-shardsFineWeb Edu 100B dataset tokenized with GPT-2 tokenizer using the code in llm.c repo. 8 likes5.1k downloads2y agoHugging Face04kushalt /fineweb-edu-gpt2tabular10M<n<100M0 likes4.5k downloads8mo agoHugging Face05apollo-research /Skylion007-openwebtext-tokenizer-gpt21M<n<10M3 likes2.7k downloads3y agoHugging Face06alkibijad /fineweb-edu-sample-10BT-gpt2tokenized1 likes1.9k downloads2y agoHugging Face07marktas /xent-tasks-gpt2textn<1K0 likes1.6k downloads2y agoHugging Face08apollo-research /monology-pile-uncopyrighted-tokenizer-gpt210M<n<100M1 likes1.3k downloads3y agoHugging Face09chanind /openwebtext-gpt21M<n<10M0 likes1.3k downloads2y agoHugging Face10hanspeterlyngsoeraaschoujensen /gpt2_model_acts_openwebtexttextn<1K0 likes941 downloads2y agoHugging Face11austindavis /chess-gpt2-hiddenstates-768Is this working? tabular1M<n<10M0 likes938 downloads1y agoHugging Face12alancooney /sae-monology-pile-uncopyrighted-tokenizer-gpt210M<n<100M2 likes835 downloads3y agoHugging Face13varunneal /dolma-blend-gpt20 likes774 downloads4mo agoHugging Face14ytzi /starcoderdata-gpt2tabular10M<n<100M0 likes770 downloads2y agoHugging Face15apollo-research /sae-skeskinen-TinyStories-hf-validation-tokenizer-gpt2_playtext10K<n<100K0 likes741 downloads3y agoHugging Face16abir-hr196 /clt_gpt2_tokenized_control Fresh multilingual GPT-2 CLT control data Sequential, unshuffled control sample for CLT null experiments. For each language, complete source documents were tokenized with CausalNLP/gpt2-hf_multilingual-20 at revision 0afbb31b2db3f394270d42d6a4cb7f8fceeca3d8. The first 100,000,000 tokenizer tokens were discarded (including the complete document that crossed the threshold), after which complete documents were retained until at least 100,000,000 tokens were collected. Data are… See the full description on the dataset page: https://huggingface.co/datasets/abir-hr196/clt_gpt2_tokenized_control.texttext-generation100K<n<1M0 likes733 downloads3mo agoHugging Face17pccl-org /Skylion007-openwebtext-tokenizer-gpt2-12810M<n<100M0 likes666 downloads2y agoHugging Face18ytzi /the-stack-dedup-python-filtered-docstrings-gpt2tabular10M<n<100M0 likes564 downloads2y agoHugging Face19austindavis /chess-gpt2-hiddenstates-512 Dataset Card for Chess GPT-2 Hidden States 512 Dataset Summary This dataset contains 120k hidden state vectors from forward passes through a GPT-2 model trained on UCI chess move sequences. The model has 8 layers, each with 8 attention heads, and a hidden state size of 512. The dataset was generated by performing one forward pass for each UCI move sequence in the "austindavis/lichess_uci" dataset, specifically the "train" split from the "201301-moves" configuration.… See the full description on the dataset page: https://huggingface.co/datasets/austindavis/chess-gpt2-hiddenstates-512.tabularother1M<n<10M0 likes526 downloads2y agoHugging Face20williamplacroix /wikilarge-graded-gpt2toneizer100K<n<1M0 likes519 downloads2y agoHugging Face21CausalNLP /gpt2small_full_training_datatext1M<n<10M0 likes506 downloads1y agoHugging Face22anyasims /openwwebtext-gpt2-50257-standard1M<n<10M0 likes504 downloads2y agoHugging Face23kjj0 /finewebedu10B-gpt20 likes494 downloads2y agoHugging Face24heyunzhenwhat /Skylion007-openwebtext-gpt2-1024Tokenizer: gpt2dataset: '''Skylion007/openwebtext'''context_size : 1024 1M<n<10M0 likes474 downloads2y agoHugging Face25pietrolesci /wikitext-103-raw-v1_gpt2-20k Dataset Card for "wikitext-103-raw-v1_gpt2-20k" More Information needed tabular1M<n<10M0 likes454 downloads3y agoHugging Face26EleutherAI /LDS-retrain-bank-adamw-N16k-bs256-gpt2-mediumtabular1K<n<10K0 likes447 downloads28d agoHugging Face27heyunzhenwhat /monology-pile-uncopyrighted-gpt2-1024Tokenizer: gpt2dataset: '''monology-pile-uncopyrighted'''context_size : 1024 100M<n<1B2 likes439 downloads2y agoHugging Face28chrisjob1021 /gpt2_tokenized_concatenated_openwebtext1M<n<10M0 likes437 downloads1y agoHugging Face29EleutherAI /PARTIAL_LDS-retrain-bank-gpt2medium-16k-bs32tabular1K<n<10K0 likes412 downloads21d agoHugging Face30JpChi /finewebedu10BT-tokenized-gpt2 FineWebEdu This is FineWebEdu sample10T subset dataset tokenized into shards with GPT2 tokenizer. 0 likes409 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.