datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fineweb-edu-gpt2xent-tasks-gpt2chess-gpt2-hiddenstates-768Is this working?
gpt2_model_acts_openwebtextstarcoderdata-gpt2sae-skeskinen-TinyStories-hf-validation-tokenizer-gpt2_playclt_gpt2_tokenized_control
Fresh multilingual GPT-2 CLT control data
Sequential, unshuffled control sample for CLT null experiments. For each language,
complete source documents were tokenized with CausalNLP/gpt2-hf_multilingual-20 at revision
0afbb31b2db3f394270d42d6a4cb7f8fceeca3d8. The first 100,000,000 tokenizer tokens were discarded
(including the complete document that crossed the threshold), after which complete
documents were retained until at least 100,000,000 tokens were collected.
Data are… See the full description on the dataset page: https://huggingface.co/datasets/abir-hr196/clt_gpt2_tokenized_control.the-stack-dedup-python-filtered-docstrings-gpt2gpt2small_full_training_datachess-gpt2-hiddenstates-512
Dataset Card for Chess GPT-2 Hidden States 512
Dataset Summary
This dataset contains 120k hidden state vectors from forward passes through a GPT-2 model trained on UCI chess move sequences.
The model has 8 layers, each with 8 attention heads, and a hidden state size of 512.
The dataset was generated by performing one forward pass for each UCI move sequence in the "austindavis/lichess_uci" dataset,
specifically the "train" split from the "201301-moves" configuration.… See the full description on the dataset page: https://huggingface.co/datasets/austindavis/chess-gpt2-hiddenstates-512.PrismAI_v2-encoded-gpt2clean-gsm8k-aug-gpt2-0.1b-lora-think-sft-ar-clean-gsm8k-aug
clean-gsm8k-aug-gpt2-0.1b-lora-think-sft-ar-clean-gsm8k-aug
Ten generated think-sft-ar responses per question from cs-giung/gpt2-0.1b-lora-think-sft-ar-clean-gsm8k-aug (revision step-300000), generated on 2026-08-07.
Schema
Field
Type
Meaning
question
str
Source question from cs-giung/clean-gsm8k-aug
steps
list[list[str]]
The 10 responses, each split into reasoning steps
answer
list[str]
The 10 per-response answer blocks
Every record has… See the full description on the dataset page: https://huggingface.co/datasets/cs-giung/clean-gsm8k-aug-gpt2-0.1b-lora-think-sft-ar-clean-gsm8k-aug.the-stack-dedup-python-filtered-dec_gen_async-gpt2clt_gpt2_tokenized_control
Fresh multilingual GPT-2 CLT control data
Sequential, unshuffled control sample for CLT null experiments. For each language,
complete source documents were tokenized with CausalNLP/gpt2-hf_multilingual-20 at revision
0afbb31b2db3f394270d42d6a4cb7f8fceeca3d8. The first 100,000,000 tokenizer tokens were discarded
(including the complete document that crossed the threshold), after which complete
documents were retained until at least 100,000,000 tokens were collected.
Data are… See the full description on the dataset page: https://huggingface.co/datasets/CausalNLP/clt_gpt2_tokenized_control.the-stack-dedup-python-filtered-dec_gen_async-gpt2details_gpt2
Dataset Card for Evaluation run of gpt2
Dataset automatically created during the evaluation run of model gpt2.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration "results" store all the… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_gpt2.gpt2-pretrain-corpusgpt2small_training_dataM4-encoded-gpt2gpt2-training-ar-zh-ko-ja-4b
Balanced Arabic-Chinese-Korean-Japanese 4B-token training data
Sequential FineWeb-2 documents tokenized with CausalNLP/gpt2-tokenizer-ar-zh-ko-ja. Each language contains at least 1,000,000,000 tokens in complete documents. Total target: 4,000,000,000 tokens. Approximately 100,000,000 tokens per Parquet shard.
Splits: arb_Arab, cmn_Hani, kor_Hang, jpn_Jpan. Schema: text: string, input_ids: list<int32>.
owt-gpt2-1024-splitRAID_none-encoded-gpt2gpt2_embeddings_BATCH_14lmd_full_txt-tokenized-gpt2gpt2-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
smollm-corpus-instruct-2M-cosmopedia-v2-gpt2-v2-streaming
A corpus of high quality fine tuning data meant for fine tuning various HelixLM models
Dataset Composition:
A subset sampled from randomly selected shards from https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus cosmopedia-v2 split ...
Added: A preprocessed column formattedconversation concatenating the prompt, text and special tokens for instruct fine tuning.
Added: A token count column tokencount based on GPT2 tokenizer + the fine tuning special tokens.… See the full description on the dataset page: https://huggingface.co/datasets/david-thrower/smollm-corpus-instruct-2M-cosmopedia-v2-gpt2-v2-streaming.gpt2-tokenizer-corpusmmlu_sort_by_gpt2xl_gpqa_biogpt2-outputs
Dataset Card for "gpt2-outputs"
More Information needed
the-stack-dedup-python-filtered-gpt2
