datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
xent-tasks-gpt2gpt2_model_acts_openwebtextgpt2-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
fdm-40ch-fresh-gpt2mbkl-1-4-gpt2-basemixed-pretrain-10b-gpt2
Mixed Pretraining 10B (GPT-2 BPE)
A 10-billion-token pretraining dataset, GPT-2 BPE tokenized, assembled as a
diverse mix of web text, books, Wikipedia, code, academic papers, Q&A and
instruction-formatted conversations.
Built to train a ~500M parameter from-scratch GPT-2-style transformer (see
juliannunezb/transformer-lm-500m).
Mix
Source
Mix %
Tokens
Notes
fineweb
40.4%
4,039,999,700
reused from kjj0/fineweb10B-gpt2
fineweb_edu
15.2%
1,514,999,900
reused… See the full description on the dataset page: https://huggingface.co/datasets/juliannunezb/mixed-pretrain-10b-gpt2.gpt2-layer0-bigram-statsshaer-eval-raw-gpt2-small-arabic-poetry
Shaer Evaluation Results
Models: gpt2_small_arabic_poetry
Source dataset: Shaer-AI/shaer-sft-test-generations-k5
Rows: 3481
Validation passed: True
Scored rows included: True
Dataset repo: Shaer-AI/shaer-eval-raw-gpt2-small-arabic-poetry
Files
generations.jsonl: raw generation rows
generations.csv: raw generation rows in CSV
generations_scored.jsonl: raw rows plus meter/count evaluation
validation.json: validation summary
generations_scored.csv: scored rows in… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI-2/shaer-eval-raw-gpt2-small-arabic-poetry.gpt2-details
Dataset Card for Evaluation run of gpt2
Dataset automatically created during the evaluation run of model gpt2
The dataset is composed of 85 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 52 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration "results" store all the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/gpt2-details.shaer-eval-raw-gpt2-medium-arabic-poetry
Raw Shaer Continuation Generations - gpt2_medium_arabic_poetry
This dataset contains cumulative raw continuation generations for gpt2_medium_arabic_poetry.
The main table is data/test.jsonl; it includes source/prompt/reference fields plus the model output in generated_text and raw_generated_text.
Current uploaded progress target: 3481 rows out of 3481.
Scored upload: 1. When scored upload is true, reward/evaluation columns are included in the main table.
Artifacts such as… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI-2/shaer-eval-raw-gpt2-medium-arabic-poetry.litgpt-the-stack-dedup-python-filtered-gpt2openai-community__gpt2-details
Dataset Card for Evaluation run of openai-community/gpt2
Dataset automatically created during the evaluation run of model openai-community/gpt2
The dataset is composed of 39 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/openai-community__gpt2-details.gpt2_to_gpt5.5_distilled_25k
GPT-2 to GPT-5.5 Advanced Reasoning Distillation (25k)
Dataset Description
25,000 unique, high-quality instruction-response pairs designed for knowledge distillation and supervised fine-tuning. The dataset elevates GPT-2 Medium toward GPT-5.5-level performance on complex reasoning tasks.
Core goal: Transfer frontier reasoning capabilities (multi-step CoT, cross-domain synthesis, edge-case analysis, novel insights) from a hypothetical GPT-5.5 teacher into smaller… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/gpt2_to_gpt5.5_distilled_25k.nuzzle-scan-openai-community-gpt2openai-community__gpt2-large-details
Dataset Card for Evaluation run of openai-community/gpt2-large
Dataset automatically created during the evaluation run of model openai-community/gpt2-large
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/openai-community__gpt2-large-details.openai-community__gpt2-xl-details
Dataset Card for Evaluation run of openai-community/gpt2-xl
Dataset automatically created during the evaluation run of model openai-community/gpt2-xl
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/openai-community__gpt2-xl-details.openai-community__gpt2-medium-details
Dataset Card for Evaluation run of openai-community/gpt2-medium
Dataset automatically created during the evaluation run of model openai-community/gpt2-medium
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/openai-community__gpt2-medium-details.fineweb10B-gpt2-fluxentropy
FineWeb Dataset - GPT-2 Tokenized
This dataset contains preprocessed and tokenized FineWeb data using the GPT-2 tokenizer.
It consists of multiple training folders containing the processed data.
Dataset structure:
fineweb_train_000001 to fineweb_train_000005: Training folders
gpt20b_labeledmatilda-smollm-mix-15b-gpt2
matilda-smollm-mix-15B-gpt2
15 B GPT-2-BPE tokens drawn from a 5:1 token-balanced mix of
HuggingFaceTB/smollm-corpus:
Source
Share
Tokens
fineweb-edu-dedup
83.33 %
12.50 B
cosmopedia-v2
16.67 %
2.50 B
Total: 15,000,349,569 tokens across 151 shards (shard_*.bin, uint16,
100 M tokens per shard).
The full SmolLM recipe is 75 / 15 / 10 fineweb-edu / cosmopedia-v2 / python-edu.
python-edu was dropped because the HuggingFaceTB/smollm-corpus subset
ships only blob_id… See the full description on the dataset page: https://huggingface.co/datasets/prometheus04/matilda-smollm-mix-15b-gpt2.fineweb-gpt2bpe-20B
fineweb-gpt2bpe-20B
FineWeb-Edu (from karpathy/fineweb-edu-100b-shuffle), pre-tokenized with the gpt2bpe tokenizer (vocab 50,257)
and packaged as flat uint16 token-id .bin files for fast memmap training.
file
split
tokens
train.bin
train
20,000,000,000
val.bin
val
54,798,984
train and val are disjoint held-out partitions. Each .bin is a raw
little-endian uint16 stream (no header); token count = filesize / 2, and
train.meta.json / val.meta.jsoncarry the full… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/fineweb-gpt2bpe-20B.open_gpt2-chatbotUser-AI conversations generated by "gpt2-chatbot" scraped for online sources. The dataset is in alpaca format. The quality is pretty terrible due to automatic scraping and data sources, so this is intended for research purposes. Please contribute your chats if you would like.
gpt2-large-generated
Synthetic Text Corpus - GPT-2 Large
Dataset Description
This dataset contains synthetically generated text sequences sampled from GPT-2 Large. It was created to provide a large-scale text corpus for research in natural language processing, particularly for studies on model behavior, text generation, and language modeling.
Dataset Summary
Size: ~100M tokens
Number of sequences: ~500,000
Source model: gpt2-large (774M parameters)
Sequence length: Maximum 256… See the full description on the dataset page: https://huggingface.co/datasets/vesteinn/gpt2-large-generated.gpt-2-outputtokenized_QA_gpt2DeepAutoAI__d2nwg_causal_gpt2_v1-details
Dataset Card for Evaluation run of DeepAutoAI/d2nwg_causal_gpt2_v1
Dataset automatically created during the evaluation run of model DeepAutoAI/d2nwg_causal_gpt2_v1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DeepAutoAI__d2nwg_causal_gpt2_v1-details.real-vs-gpt2-sentences
Real vs. GPT2 Sentences Dataset
A collection of 300,000+ sentences comparing human-written text with AI-generated completions.
Dataset Overview
Each row contains:
seed: First 5 words of a sentence
real: Original complete sentence
gpt2: AI-generated sentence completion
Example
{
"seed": "They should read some history,",
"real": "They should read some history, particularly what happened to the big German industries once Hitler came to power...",
"gpt2":… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/real-vs-gpt2-sentences.postbot__gpt2-medium-emailgen-details
Dataset Card for Evaluation run of postbot/gpt2-medium-emailgen
Dataset automatically created during the evaluation run of model postbot/gpt2-medium-emailgen
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/postbot__gpt2-medium-emailgen-details.Sharathhebbar24__chat_gpt2_dpo-details
Dataset Card for Evaluation run of Sharathhebbar24/chat_gpt2_dpo
Dataset automatically created during the evaluation run of model Sharathhebbar24/chat_gpt2_dpo
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sharathhebbar24__chat_gpt2_dpo-details.flava-gpt2-data
