datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SimpleStories
📘📕 SimpleStories 📙📗
SimpleStories is a dataset of >2 million model-generated short stories. It was made to train small, interpretable language models on it. The generation process is open-source: To see how the dataset was generated, or to generate some stories yourself, head over to this repository.
If you'd like to commission other languages or story formats, feel free to send mail.
When using SimpleStories in your work, please cite the SimpleStories paper:… See the full description on the dataset page: https://huggingface.co/datasets/SimpleStories/SimpleStories.sim-posttrain
HUMANUAL Posttraining Data
Posttraining data for user simulation, derived from the train splits of the
HUMANUAL benchmark datasets.
Datasets
HUMANUAL (posttraining)
Config
Rows
Description
news
48,618
News article comment responses
politics
45,429
Political discussion responses
opinion
37,791
Reddit AITA / opinion thread responses
book
34,170
Book review responses
chat
23,141
Casual chat responses
email
6,377
Email reply responses… See the full description on the dataset page: https://huggingface.co/datasets/Xuhui/sim-posttrain.simverse2026
SimVerse
⚠️ Anonymized for double-blind review. This dataset is currently undergoing peer review. It is hosted under an anonymous account dedicated to the review process; the author and citation fields are deliberately unfilled. Permanent ownership and citation information will be added after the review concludes. Please do not attempt to deanonymize the maintainers of this dataset during review.
A multi-task benchmark for evaluating multimodal LLMs on interactive simulation… See the full description on the dataset page: https://huggingface.co/datasets/SimVer-ano/simverse2026.multilingual-textarena-SimpleTak-v0-train-v2
TextArena Language Trajectories
This dataset contains language-conditioned TextArena trajectory data.
Each dataset configuration corresponds to a different model, experiment group, or source folder.
Available configurations:
gemma4-e4b-it
qwen3-4b
ministral3-3b-instruct
Usage
Install the datasets library:
pip install datasets
Load a specific configuration:
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/The-CoLab/multilingual-textarena-SimpleTak-v0-train-v2.simpsons_infoThe information of all episodes of the cartoon show "The Simpsons" from wikipedia. Some (mainly in recent 32, 33, 34 seasons) plot missing.
SimpleStories-JA
📘📕 SimpleStories 📙📗
このデータセットは、gpt-4o-miniによって生成された短編小説で出来ているデータセットです。生成方法や、自分で物語を生成する方法については、こちらのリポジトリをご覧ください。
他の言語や物語形式の制作を希望される場合は、メールにてお問い合わせください。
SimpleStoriesは、EldenとLiによるTinyStoriesの改良版です。
特徴
物語の注釈情報(theme、topic、styleなど)
多様性の高さ
2024年のモデルによって生成
NLPのデータが用意しているためフィルタリングしやすい
以下の言語版が利用可能:
英語
日本語
他にも追加予定
This dataset is a collection of short stories generated by gpt-4o-mini (+ other models, soon). To see how this dataset was generated, or to generate some stories… See the full description on the dataset page: https://huggingface.co/datasets/SimpleStories/SimpleStories-JA.multilingual-textarena-SimpleTak-v0-train
TextArena Language Trajectories
This dataset contains language-conditioned TextArena trajectory data.
Each dataset configuration corresponds to a different model, experiment group, or source folder.
Available configurations:
gemma4-e4b-it
qwen3-4b
ministral3-3b-instruct
Usage
Install the datasets library:
pip install datasets
Load a specific configuration:
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/The-CoLab/multilingual-textarena-SimpleTak-v0-train.lawflow-reasoning-simulation
LawFlow: Collecting and Simulating Lawyers' Thought Processes
Debarati Das, Khanh Chi Le*, Ritik Parkar*, Karin De Langis, Brendan Madson, Chad Berryman, Robin Willis, Daniel Moses, Brett McDonnell†, Daniel Schwarcz†, Dongyeop Kang†
Minnesota NLP, University of Minnesota Twin Cities
*equal contribution, †senior advisors
Arxiv
Project Page
Dataset Summary and Purpose
LawFlow: Collecting and Simulating Lawyers' Thought Processes
The purpose of this dataset is aim… See the full description on the dataset page: https://huggingface.co/datasets/minnesotanlp/lawflow-reasoning-simulation.sakhi
Sakhi: A Community-Validated Multilingual Maternal-Health Benchmark
Sakhi is a benchmark for evaluating large language models on maternal and reproductive-health questions in three languages spoken in low-resource settings: English, Hindi, and Marathi. It was built around a deployed WhatsApp-based maternal-health chatbot reaching rural mothers in Hindi- and Marathi-speaking districts of India, with a three-channel review pipeline: practising Indian doctors, Accredited Social Health… See the full description on the dataset page: https://huggingface.co/datasets/SimPPL/sakhi.SimScholar-SFT
S3 SFT Trajectories
Complete ReAct trajectories for scientific-literature search.
Code ·
S3 collection ·
Source corpus
This dataset contains 14,633 single- and two-hop tool-use trajectories. In each
trajectory, a policy searches and reads a fixed scientific corpus through nine
tools, then submits an answer with a correctness label. The messages column
uses OpenAI tool-calling chat format.
At a glance
Question type
Rows
Correct
Incorrect
Single-hop… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/SimScholar-SFT.quantum-simulation-chemistry-materials
Neura Parse — Quantum Simulation of Chemistry & Materials: Encodings, VQE/QPE & Dynamics
An application-deep, code-backed vertical on simulating quantum matter: electronic-structure problems, fermion-to-qubit encodings, Hamiltonian factorizations, ground/excited-state and real-time-dynamics algorithms, and analog simulation, with end-to-end resource estimates and honest classical-competitor accounting. Built with Qiskit Nature, OpenFermion, PennyLane-QChem, and PySCF — far… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-simulation-chemistry-materials.age-specific-text-simplification
Age-Specific Text Simplification Dataset
Dataset Description
This dataset contains complex texts simplified into age-appropriate versions for children aged 3, 4, and 5 years old. Each original text has been professionally adapted to match the cognitive development, vocabulary, and comprehension abilities of each specific age group.
Dataset Summary
Total Examples: 17,177
Training Split: 15,459 examples
Validation Split: 1,718 examples
Languages:… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/age-specific-text-simplification.SimpleStories
📘📕 SimpleStories 📙📗
SimpleStories is a dataset of >2 million model-generated short stories. It was made to train small, interpretable language models on it. The generation process is open-source: To see how the dataset was generated, or to generate some stories yourself, head over to this repository.
If you'd like to commission other languages or story formats, feel free to send mail.
When using SimpleStories in your work, please cite the SimpleStories paper:… See the full description on the dataset page: https://huggingface.co/datasets/duoduoyeah/SimpleStories.SimCoPilot
Dataset Card for Dataset Name
SimCoPilot is a benchmark for evaluating LLMs to perform as a "copilot"-style, interactive coding assistant.
Dataset Details
Dataset Description
SimCoPilot is a benchmark for evaluating LLMs to perform as a "copilot"-style, interactive coding assistant, testing their ability to add and complete code in complex real-world software environments and analyzing how LLMs manage different code dependencies and logic complexities.
Curated… See the full description on the dataset page: https://huggingface.co/datasets/mj33/SimCoPilot.bird
Dataset
This dataset is a polished version of the BIRD dataset. It was introduced in the paper Think2SQL: Reinforce LLM Reasoning Capabilities for Text2SQL.
It has been used to train the reasoning Text2SQL model simone-papicchio/Think2SQL-7B.
Please refer to the paper for further details.
License: CC BY-SA 4.0
Citation
@misc{papicchio2025think2sqlreinforcellmreasoning,
title={Think2SQL: Reinforce LLM Reasoning Capabilities for Text2SQL},
author={Simone… See the full description on the dataset page: https://huggingface.co/datasets/simone-papicchio/bird.iterative-dpo-data-for-SimPO-iter2
iterative-dpo-data-for-SimPO-iter2
概要
合成instructionデータであるAratako/Magpie-Tanuki-Instruction-Selected-Evolved-26.5kを元に以下のような手順で作成した日本語Preferenceデータセットです。
開発途中のモデルであるAratako/Llama-Gemma-2-27b-CPO_SimPO-iter1を用いて、temperature=1で回答を5回生成
5個の回答それぞれに対して、Qwen/Qwen2.5-72B-Instruct-GPTQ-Int8を用いて0~5点のスコア付けを実施
1つのinstructionに対する5個の回答について、最もスコアが高いものをchosenに、低いものをrejectedに配置
全て同じスコアの場合や、最も良いスコアが2点以下の場合は除外
ライセンス
本データセットは回答の作成に利用したモデルの関係で以下のライセンスの影響を受けます。
META LLAMA 3.1… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/iterative-dpo-data-for-SimPO-iter2.ORCHESTRA-simple-1M
ORCHESTRA-simple-1M
GitHub: nk2028/ORCHESTRA-dataset
中文簡介
ORCHESTRA (cOmpRehensive Classical cHinESe poeTRy dAtaset) 是一個全面的古典中文詩歌的數據集,數據來自搜韻網。本數據集由 nk2028 進行格式轉換並發佈,希望透過公開高品質的古典中文詩歌數據,促進對古典中文詩歌及古典中文自然語言處理的研究。
ORCHESTRA-simple 是 ORCHESTRA 數據集的簡化格式,僅保留 id, title, group_index, type, dynasty, author, content 這 7 個欄位,而去除其他欄位,以簡化使用。
本資料集可用於大型語言模型的訓練。如欲作其他用途,請向數據提供者搜韻網諮詢。
English Introduction
ORCHESTRA (cOmpRehensive Classical cHinESe poeTRy dAtaset) is a comprehensive dataset of classical… See the full description on the dataset page: https://huggingface.co/datasets/Ayaka/ORCHESTRA-simple-1M.datasus_sim
Dataset Card for [DATASUS SIM]
This dataset is a large-scale collection of information around deaths registered by the Brazilian public health care system.
Dataset Details
Check the Data Dictionary attached to the project.
Dataset Sources
Repository: [Link to HF Repo or GitHub]
Paper: [Optional: Link to ArXiv or Journal]
Demo: [Optional: Link to Space or Web App]
Uses
Direct Use
Pre-training: Training Large Language Models (LLMs)… See the full description on the dataset page: https://huggingface.co/datasets/jgchaicoski/datasus_sim.Simple-agent-traces
📱 Simple Agent Traces – Tiny Tool‑Calling Conversations for Small Models
Simple Agent Traces is a compact, hand‑picked dataset of 605 real‑world tool‑calling conversations, each carefully truncated to ≤8,192 tokens (using the SmolLM2‑360M tokenizer).It is purpose‑built for training and fine‑tuning tiny language models (≤500M) that must run on‑device – smartphones, edge devices, or any environment with strict memory and latency constraints.
🧹 No chain‑of‑thought, no fluff.Every… See the full description on the dataset page: https://huggingface.co/datasets/LiteMind/Simple-agent-traces.countdown-qwen3-0.6b
Countdown Qwen3-0.6B Pass@10 Buckets
Countdown arithmetic problems filtered by observed local Qwen/Qwen3-0.6B success rate over 10 rollouts per problem.
Each problem asks for an arithmetic expression that reaches a target using each listed source number at most once. The final answer should be inside \boxed{...}. Canonical solutions are provided, but any verifier-valid expression is accepted.
Subsets
subset
source bucket
count
observed successes out of 10… See the full description on the dataset page: https://huggingface.co/datasets/simpissa/countdown-qwen3-0.6b.simson-assembly-dependency-graph
🔗 Simson Assembly Dependency Graph
13 „Braucht-auch"-Ketten für Simson-Reparaturen und Tuning.
Was das ist
Jeder Eintrag beschreibt, was man zusätzlich braucht wenn man ein bestimmtes Teil einbaut:
Pflicht-Teile (mandatory_with): Ohne diese geht es nicht
Empfohlene Teile (recommended_with): Sinnvoll, aber optional
Inkompatible Teile: Was NICHT gleichzeitig verbaut werden kann
Upgrade-Pfad: Was als nächstes Sinn macht
Geschätzte Arbeitszeit & Skill-Level
Kosten:… See the full description on the dataset page: https://huggingface.co/datasets/jmp1987/simson-assembly-dependency-graph.Complete-FABLE.5-traces-2M
Complete FABLE.5 Traces 2M
Full FABLE.5 / Mythos corpus restored, with session-limit answer rows removed.
Dataset Viewer | Parquet | Raw JSONL.gz
This dataset is a post-closure compilation of all available FABLE.5 / Mythos trace datasets found on Hugging Face during the curation pass after the closure of Fable and Mythos. It is deduplicated at the normalized-row level and keeps row-level provenance through first_source_dataset, first_source_config, first_source_split… See the full description on the dataset page: https://huggingface.co/datasets/simonzimmo/Complete-FABLE.5-traces-2M.simson-youtube-tutorials
📺 Simson YouTube Tutorial Metadata
20 kuratierte YouTube-Tutorial-Einträge für Simson-Moped Reparatur, Tuning und Restaurierung.
Inhalt
Strukturierte Metadaten der wichtigsten Simson-Tutorial-Videos auf YouTube:
Kanal-Typen: DIY-Werkstatt, Tuning-Spezialist, Restaurierungs-Kanal, Enthusiasten-Kanal, Dokumentation
Topics: Motor, Zündung, Vergaser, Elektrik, Tuning, Restaurierung, Fahrwerk, Geschichte, Wartung
Fahrzeuge: S50, S51, S70, KR51/1, KR51/2 (Schwalbe)… See the full description on the dataset page: https://huggingface.co/datasets/jmp1987/simson-youtube-tutorials.simson-unified-knowledge-graph
🧠 Simson Unified Knowledge Graph
173 Nodes × 348 Edges – der Klebstoff zwischen allen Simson-Datasets.
Was das ist
Ein maschinenlesbarer Graph, der alle 6 Datasets miteinander verknüpft:
Dataset
Status
Nodes
racing-planet-simson-traces
Diagnose-Traces
15
simson-forum-qa-pairs
Forum-Wissen
30
simson-repair-manual
Technische Daten
14
racing-planet-product-catalog
Teilekatalog
37
simson-youtube-tutorials
Video-Tutorials
20… See the full description on the dataset page: https://huggingface.co/datasets/jmp1987/simson-unified-knowledge-graph.drich-simhitshaddas-tigrinya-corpus
haddas-tigrinya-corpus
Monolingual Tigrinya newspaper text segmented into article bodies. Suitable for continued pretraining or causal language modeling of Tigrinya LLMs.
Source
Derived from the Haddas Eritrea newspaper archive: 63 PDF issues processed
by the haddas-eritrea pipeline (extract -> clean -> segment -> translate -> label).
Generated: 2026-04-26 12:21 UTC
Row count: 2653
Schema: id, text, char_count, topic, issue_date, source_pdf, page_start, page_end… See the full description on the dataset page: https://huggingface.co/datasets/SIMBA9657/haddas-tigrinya-corpus.controlled_text_simplyControlled text simplification, targetting at different audiences. A dataset for a group project.
