CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kushalt /fineweb-edu-gpt2tabular10M<n<100M0 likes6.5k downloads8mo agoHugging Face02marktas /xent-tasks-gpt2textn<1K0 likes1.8k downloads2y agoHugging Face03austindavis /chess-gpt2-hiddenstates-768Is this working? tabular1M<n<10M0 likes950 downloads1y agoHugging Face04hanspeterlyngsoeraaschoujensen /gpt2_model_acts_openwebtexttextn<1K0 likes825 downloads2y agoHugging Face05ytzi /starcoderdata-gpt2tabular10M<n<100M0 likes812 downloads2y agoHugging Face06apollo-research /sae-skeskinen-TinyStories-hf-validation-tokenizer-gpt2_playtext10K<n<100K0 likes743 downloads3y agoHugging Face07abir-hr196 /clt_gpt2_tokenized_control Fresh multilingual GPT-2 CLT control data Sequential, unshuffled control sample for CLT null experiments. For each language, complete source documents were tokenized with CausalNLP/gpt2-hf_multilingual-20 at revision 0afbb31b2db3f394270d42d6a4cb7f8fceeca3d8. The first 100,000,000 tokenizer tokens were discarded (including the complete document that crossed the threshold), after which complete documents were retained until at least 100,000,000 tokens were collected. Data are… See the full description on the dataset page: https://huggingface.co/datasets/abir-hr196/clt_gpt2_tokenized_control.texttext-generation100K<n<1M0 likes701 downloads3mo agoHugging Face08ytzi /the-stack-dedup-python-filtered-docstrings-gpt2tabular10M<n<100M0 likes628 downloads2y agoHugging Face09CausalNLP /gpt2small_full_training_datatext1M<n<10M0 likes560 downloads1y agoHugging Face10austindavis /chess-gpt2-hiddenstates-512 Dataset Card for Chess GPT-2 Hidden States 512 Dataset Summary This dataset contains 120k hidden state vectors from forward passes through a GPT-2 model trained on UCI chess move sequences. The model has 8 layers, each with 8 attention heads, and a hidden state size of 512. The dataset was generated by performing one forward pass for each UCI move sequence in the "austindavis/lichess_uci" dataset, specifically the "train" split from the "201301-moves" configuration.… See the full description on the dataset page: https://huggingface.co/datasets/austindavis/chess-gpt2-hiddenstates-512.tabularother1M<n<10M0 likes540 downloads2y agoHugging Face11TheItCrOw /PrismAI_v2-encoded-gpt2tabular100K<n<1M0 likes415 downloads1y agoHugging Face12cs-giung /clean-gsm8k-aug-gpt2-0.1b-lora-think-sft-ar-clean-gsm8k-aug clean-gsm8k-aug-gpt2-0.1b-lora-think-sft-ar-clean-gsm8k-aug Ten generated think-sft-ar responses per question from cs-giung/gpt2-0.1b-lora-think-sft-ar-clean-gsm8k-aug (revision step-300000), generated on 2026-08-07. Schema Field Type Meaning question str Source question from cs-giung/clean-gsm8k-aug steps list[list[str]] The 10 responses, each split into reasoning steps answer list[str] The 10 per-response answer blocks Every record has… See the full description on the dataset page: https://huggingface.co/datasets/cs-giung/clean-gsm8k-aug-gpt2-0.1b-lora-think-sft-ar-clean-gsm8k-aug.text100K<n<1M0 likes379 downloads2mo agoHugging Face13ytzi /the-stack-dedup-python-filtered-dec_gen_async-gpt2tabular10M<n<100M0 likes378 downloads2y agoHugging Face14CausalNLP /clt_gpt2_tokenized_control Fresh multilingual GPT-2 CLT control data Sequential, unshuffled control sample for CLT null experiments. For each language, complete source documents were tokenized with CausalNLP/gpt2-hf_multilingual-20 at revision 0afbb31b2db3f394270d42d6a4cb7f8fceeca3d8. The first 100,000,000 tokenizer tokens were discarded (including the complete document that crossed the threshold), after which complete documents were retained until at least 100,000,000 tokens were collected. Data are… See the full description on the dataset page: https://huggingface.co/datasets/CausalNLP/clt_gpt2_tokenized_control.texttext-generation100K<n<1M0 likes339 downloads3mo agoHugging Face15leoleoasd /the-stack-dedup-python-filtered-dec_gen_async-gpt2tabular10M<n<100M0 likes307 downloads2y agoHugging Face16OALL /details_gpt2 Dataset Card for Evaluation run of gpt2 Dataset automatically created during the evaluation run of model gpt2. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional configuration "results" store all the… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_gpt2.tabular100K<n<1M0 likes258 downloads2y agoHugging Face17himalaya-ai /gpt2-pretrain-corpustabular10M<n<100M0 likes242 downloads5mo agoHugging Face18CausalNLP /gpt2small_training_datatext100K<n<1M0 likes235 downloads1y agoHugging Face19TheItCrOw /M4-encoded-gpt2tabular100K<n<1M0 likes229 downloads1y agoHugging Face20CausalNLP /gpt2-training-ar-zh-ko-ja-4b Balanced Arabic-Chinese-Korean-Japanese 4B-token training data Sequential FineWeb-2 documents tokenized with CausalNLP/gpt2-tokenizer-ar-zh-ko-ja. Each language contains at least 1,000,000,000 tokens in complete documents. Total target: 4,000,000,000 tokens. Approximately 100,000,000 tokens per Parquet shard. Splits: arb_Arab, cmn_Hani, kor_Hang, jpn_Jpan. Schema: text: string, input_ids: list<int32>. texttext-generation1M<n<10M0 likes222 downloads3mo agoHugging Face21dhruveshpatel /owt-gpt2-1024-splittext1M<n<10M0 likes220 downloads1y agoHugging Face22TheItCrOw /RAID_none-encoded-gpt2tabular100K<n<1M0 likes215 downloads1y agoHugging Face23jahb57 /gpt2_embeddings_BATCH_14text100K<n<1M0 likes196 downloads3y agoHugging Face24nicholasbien /lmd_full_txt-tokenized-gpt2text100K<n<1M0 likes188 downloads3y agoHugging Face25toksuitebackup /gpt2-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl). tabular10M<n<100M0 likes177 downloads10mo agoHugging Face26david-thrower /smollm-corpus-instruct-2M-cosmopedia-v2-gpt2-v2-streaming A corpus of high quality fine tuning data meant for fine tuning various HelixLM models Dataset Composition: A subset sampled from randomly selected shards from https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus cosmopedia-v2 split ... Added: A preprocessed column formattedconversation concatenating the prompt, text and special tokens for instruct fine tuning. Added: A token count column tokencount based on GPT2 tokenizer + the fine tuning special tokens.… See the full description on the dataset page: https://huggingface.co/datasets/david-thrower/smollm-corpus-instruct-2M-cosmopedia-v2-gpt2-v2-streaming.texttext-generation1M<n<10M0 likes170 downloads4mo agoHugging Face27himalaya-ai /gpt2-tokenizer-corpustabular1M<n<10M0 likes168 downloads6mo agoHugging Face28zekeZZ /mmlu_sort_by_gpt2xl_gpqa_biotext10K<n<100K0 likes151 downloads2y agoHugging Face29spacerini /gpt2-outputs Dataset Card for "gpt2-outputs" More Information needed tabular1M<n<10M0 likes150 downloads4y agoHugging Face30ytzi /the-stack-dedup-python-filtered-gpt2tabular10M<n<100M1 likes149 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.