datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SimpleSafetyTestsSimpleStories
📘📕 SimpleStories 📙📗
SimpleStories is a dataset of >2 million model-generated short stories. It was made to train small, interpretable language models on it. The generation process is open-source: To see how the dataset was generated, or to generate some stories yourself, head over to this repository.
If you'd like to commission other languages or story formats, feel free to send mail.
When using SimpleStories in your work, please cite the SimpleStories paper:… See the full description on the dataset page: https://huggingface.co/datasets/SimpleStories/SimpleStories.sim-posttrain
HUMANUAL Posttraining Data
Posttraining data for user simulation, derived from the train splits of the
HUMANUAL benchmark datasets.
Datasets
HUMANUAL (posttraining)
Config
Rows
Description
news
48,618
News article comment responses
politics
45,429
Political discussion responses
opinion
37,791
Reddit AITA / opinion thread responses
book
34,170
Book review responses
chat
23,141
Casual chat responses
email
6,377
Email reply responses… See the full description on the dataset page: https://huggingface.co/datasets/Xuhui/sim-posttrain.simple_wikipedia
simple wikipedia
the 'simple' split of Wikipedia, from Sept 1 2023. The train split contains about 65M tokens,
Pulled via:
dataset = load_dataset(
"wikipedia", language="simple", date="20230901", beam_runner="DirectRunner"
)
stats
train split
general info
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 226242 entries, 0 to 226241
Data columns (total 4 columns):
# Column Non-Null Count Dtype
--- ------ -------------- -----
0 id… See the full description on the dataset page: https://huggingface.co/datasets/pszemraj/simple_wikipedia.simple-math
Simple Math: 2+2=4 -1=3 (LoLo: Learning Only Logical Operations)
Just like my teacher gave me homework, i thought maybe we can also add some of these basics on the trainings of our models.
It was created with very simple code that is in the repo, if you add more complex operations and so.. please share the code :D thank you
Current Code Version: 20240127.fblgit (A modification over @win10 for progressive and DPO operation)
Does it Works?
34BEAGLES… See the full description on the dataset page: https://huggingface.co/datasets/fblgit/simple-math.simverse2026
SimVerse
⚠️ Anonymized for double-blind review. This dataset is currently undergoing peer review. It is hosted under an anonymous account dedicated to the review process; the author and citation fields are deliberately unfilled. Permanent ownership and citation information will be added after the review concludes. Please do not attempt to deanonymize the maintainers of this dataset during review.
A multi-task benchmark for evaluating multimodal LLMs on interactive simulation… See the full description on the dataset page: https://huggingface.co/datasets/SimVer-ano/simverse2026.african-multilingual-tokenizer-challenge
African Multilingual Tokenizer Challenge dataset
The frozen public corpus for the African Multilingual Tokenizer Challenge. It contains one balanced multilingual train split and one balanced validation split.
Split
Per language
Total
Train
40,000
240,000
Validation
4,000
24,000
Languages are English (en), French (fr), Hausa (ha), Swahili (sw), Yoruba (yo) and Amharic (am). Official public-test and private-test text are deliberately absent from this repository.… See the full description on the dataset page: https://huggingface.co/datasets/Similoluwa/african-multilingual-tokenizer-challenge.agent-simulations
Agent Simulations
Made with the whileai SDK · Collections: Simulation, Start here: foundational post-training datasets
53,971 synthetic agent trajectories generated by simulations
across 34 agent types. The rows include successful and failed
trajectories for supervised fine-tuning, preference work, reinforcement learning, and
evaluation.
NOTE: This is generated test and training data, not curated ground truth. Review and
filter it for your application before training or… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/agent-simulations.simpleqa-verified-multilingual
SimpleQA Verified Multilingual
Multilingual translations of SimpleQA Verified, a 1,000-prompt factuality benchmark from Google DeepMind that evaluates short-form parametric knowledge (facts stored in model weights).
Source: google/simpleqa-verified (eval split, 1,000 examples)
Languages
Config
Language
Examples
ces
Czech
100
dan
Danish
100
deu
German
1,000
fra
French
100
ita
Italian
100
nld
Dutch
100
pol
Polish
100
spa
Spanish
100
More to… See the full description on the dataset page: https://huggingface.co/datasets/ellamind/simpleqa-verified-multilingual.simplequestions-sparqltotext
Dataset Card for SimpleQuestions-SPARQLtoText
Dataset Summary
Special version of SimpleQuestions with SPARQL queries formatted for the SPARQL-to-Text task.
JSON fields
The original version of SimpleQuestions is a raw text file listing triples and the natural language question. A JSON version has been generated and augmented with the following fields:
rdf_subject, rdf_property, rdf_object: triple in the Wikidata format (IDs)
nl_subject, nl_property, nl_object:… See the full description on the dataset page: https://huggingface.co/datasets/Orange/simplequestions-sparqltotext.LiteResearcher-SFT-Data
LiteResearcher — SFT Cold-Start Data
Distilled deep-research trajectories used for the SFT cold-start of LiteResearcher-4B
This dataset contains the 68,231 multi-turn deep-research trajectories used to
train the SFT cold-start checkpoint that RL (GRPO+TIS) is later launched from —
the "68.2 K distilled deep-research trajectories" referenced in the paper and in
LiteResearcher-Data.
Each row is a complete ReAct-style episode: a research question, the model's
interleaved… See the full description on the dataset page: https://huggingface.co/datasets/simplex-ai-inc/LiteResearcher-SFT-Data.simutrade-rag-sft-28k
📢 Domain & Email Migration Notice
From May 6th, 2026, Simutrade will transition to new domains as simutrade.app will not be renewed:
🌐 Website: simutrade.faizath.com (formerly simutrade.app)
⚙️ API: simutrade-api.faizath.com (formerly api.simutrade.app)
📧 Email: contact@simutrade.faizath.com (formerly contact@simutrade.app)
🛰️ CDN: simutrade-cdn.faizath.com (formerly cdn.simutrade.app)
📈 Status Pages:… See the full description on the dataset page: https://huggingface.co/datasets/simutrade/simutrade-rag-sft-28k.simplescaling-s1K-R1
Dataset Card: s1k R1
Dataset Description
The s1k R1 dataset is a fork of the simplescaling/s1K dataset. It contains a collection of conversations where the assistant's messages have been enhanced to include Chain of Thought (CoT) reasoning within <think> ... </think> tags, followed by the final answer. This modification aims to improve the interpretability and reasoning capabilities of AI models by providing explicit thought processes in the responses.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/simplescaling-s1K-R1.multilingual-textarena-SimpleTak-v0-train-v2
TextArena Language Trajectories
This dataset contains language-conditioned TextArena trajectory data.
Each dataset configuration corresponds to a different model, experiment group, or source folder.
Available configurations:
gemma4-e4b-it
qwen3-4b
ministral3-3b-instruct
Usage
Install the datasets library:
pip install datasets
Load a specific configuration:
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/The-CoLab/multilingual-textarena-SimpleTak-v0-train-v2.task1347_glue_sts-b_similarity_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1347_glue_sts-b_similarity_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1347_glue_sts-b_similarity_classification.simpsons_infoThe information of all episodes of the cartoon show "The Simpsons" from wikipedia. Some (mainly in recent 32, 33, 34 seasons) plot missing.
SimpleStories-JA
📘📕 SimpleStories 📙📗
このデータセットは、gpt-4o-miniによって生成された短編小説で出来ているデータセットです。生成方法や、自分で物語を生成する方法については、こちらのリポジトリをご覧ください。
他の言語や物語形式の制作を希望される場合は、メールにてお問い合わせください。
SimpleStoriesは、EldenとLiによるTinyStoriesの改良版です。
特徴
物語の注釈情報(theme、topic、styleなど)
多様性の高さ
2024年のモデルによって生成
NLPのデータが用意しているためフィルタリングしやすい
以下の言語版が利用可能:
英語
日本語
他にも追加予定
This dataset is a collection of short stories generated by gpt-4o-mini (+ other models, soon). To see how this dataset was generated, or to generate some stories… See the full description on the dataset page: https://huggingface.co/datasets/SimpleStories/SimpleStories-JA.multilingual-textarena-SimpleTak-v0-train
TextArena Language Trajectories
This dataset contains language-conditioned TextArena trajectory data.
Each dataset configuration corresponds to a different model, experiment group, or source folder.
Available configurations:
gemma4-e4b-it
qwen3-4b
ministral3-3b-instruct
Usage
Install the datasets library:
pip install datasets
Load a specific configuration:
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/The-CoLab/multilingual-textarena-SimpleTak-v0-train.tau2-simulated
tau2 Simulated Training Set
Made with the whileai SDK · Collections: Simulation, Start here: foundational post-training datasets
The training set that took a base model from 5% to 30% on tau2-bench
telecom, made from nothing but the agent's tool list and policy.
If you build a customer-facing agent, you already have the two files this
dataset was made from: the tools it can call and the policy it follows.
The whileai SDK turned those into 1,057 graded conversations across the… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/tau2-simulated.simple_wikipedia_LM
Dataset Card for "simple_wikipedia_LM"
A filtered/edited version of pszemraj/simple_wikipedia that removes headings/contents that appear in the text column without any relevant text for them (at least in the simple split).
import re
def split_on_headings(text):
headings = ["References", "Related pages", "Other websites", "Further reading"]
for heading in headings:
parts = re.split(
r"^\s*" + re.escape(heading) + r".*$", text, flags=re.MULTILINE… See the full description on the dataset page: https://huggingface.co/datasets/pszemraj/simple_wikipedia_LM.CanlIICaseSummaries
Canadian Case Law Summaries
A database of (currently, still growing) >600 case law summaries generated by GPT 4 for random case law in Ontario or Canada
SimpleMath
🧮 SimpleMath 100K
SimpleMath 100K is a high-quality synthetic dataset of 100,000 basic arithmetic problems — no noise, no tricks, just clean and accurate math.
✅ Purpose
This was made for small AI models — not to struggle with complex math, but to get simple math right every time.
📦 Contents
75,000 numeric problems, evenly split:
18,750 addition (456 + 789 =)
18,750 subtraction (900 - 345 =)
18,750 multiplication (12 x 15 =)
18,750 division (144 / 12 =)… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/SimpleMath.simple-wiki-article
Simple-wiki-article
This dataset is directly derived from rahular/simple-wikipedia.
Various techniques were used to detect which strings were article titles, and separate the original dataset into articles, with a somewhat good accuracy.
Article content strings were merged and split with \n, and <br> was replaced with \n to make the dataset more usable.
simplepile-lite
Dataset Card for "simplepile-lite"
Interleaved dataset using 'first exhausted' strategy. Counts:
DatasetDict({
train: Dataset({
features: ['text'],
num_rows: 452432
})
validation: Dataset({
features: ['text'],
num_rows: 1000
})
test: Dataset({
features: ['text'],
num_rows: 11908
})
})
token counts - train
using GPTNeoX Tokenizer:
token_count
count
452432
mean
868.642
std
4791.71… See the full description on the dataset page: https://huggingface.co/datasets/pszemraj/simplepile-lite.lawflow-reasoning-simulation
LawFlow: Collecting and Simulating Lawyers' Thought Processes
Debarati Das, Khanh Chi Le*, Ritik Parkar*, Karin De Langis, Brendan Madson, Chad Berryman, Robin Willis, Daniel Moses, Brett McDonnell†, Daniel Schwarcz†, Dongyeop Kang†
Minnesota NLP, University of Minnesota Twin Cities
*equal contribution, †senior advisors
Arxiv
Project Page
Dataset Summary and Purpose
LawFlow: Collecting and Simulating Lawyers' Thought Processes
The purpose of this dataset is aim… See the full description on the dataset page: https://huggingface.co/datasets/minnesotanlp/lawflow-reasoning-simulation.sakhi
Sakhi: A Community-Validated Multilingual Maternal-Health Benchmark
Sakhi is a benchmark for evaluating large language models on maternal and reproductive-health questions in three languages spoken in low-resource settings: English, Hindi, and Marathi. It was built around a deployed WhatsApp-based maternal-health chatbot reaching rural mothers in Hindi- and Marathi-speaking districts of India, with a three-channel review pipeline: practising Indian doctors, Accredited Social Health… See the full description on the dataset page: https://huggingface.co/datasets/SimPPL/sakhi.LiteResearcher-Data
LiteResearcher — RL Training Data
Companion training data for the LiteResearcher paper
A low-cost, scalable Agentic RL training framework for deep-research agents.
This dataset contains the two-stage curriculum of question–answer prompts
used to train LiteResearcher-4B with on-policy GRPO+TIS, fully against a
local search / browse environment. Both stages share the same validation set.
What this is not: the underlying webpage corpus (~32 M records, used by
the local… See the full description on the dataset page: https://huggingface.co/datasets/simplex-ai-inc/LiteResearcher-Data.task146_afs_argument_similarity_gun_control
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task146_afs_argument_similarity_gun_control
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task146_afs_argument_similarity_gun_control.task934_turk_simplification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task934_turk_simplification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task934_turk_simplification.SimScholar-SFT
S3 SFT Trajectories
Complete ReAct trajectories for scientific-literature search.
Code ·
S3 collection ·
Source corpus
This dataset contains 14,633 single- and two-hop tool-use trajectories. In each
trajectory, a policy searches and reads a fixed scientific corpus through nine
tools, then submits an answer with a correctness label. The messages column
uses OpenAI tool-calling chat format.
At a glance
Question type
Rows
Correct
Incorrect
Single-hop… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/SimScholar-SFT.
