datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GeneralThought-430K-filteredData from https://huggingface.co/datasets/GeneralReasoning/GeneralThought-430K, removed prompts with non commercial data
GeneralThoughtArchive
GeneralThought-430K
Thought wants to be free
Open reasoning data for March 14 2025. This dataset was part of a side-project in the weeks following the R1 release by Chengxi and Ross - we are no longer maintaining this
dataset but are archiving it here.
The dataset contains questions, reference answers, reasoning traces, final answers and other metadata from several popular reasoning models including DeepSeek-R1, DeepSeek-R1-Zero, OpenThoughts-32B, LIMO… See the full description on the dataset page: https://huggingface.co/datasets/RJT1990/GeneralThoughtArchive.OpenOneRec-General-Pretrain
通用文本数据集
本目录包含 OpenOneRec 项目使用的通用文本数据集信息。这些数据集均来自 HuggingFace,经过清洗处理和对齐到项目统一的数据格式,并转换为 Parquet 格式用于训练。
数据格式说明
所有数据集均已转换为统一的 Parquet 格式,符合项目的数据格式规范(参考 ../README.md)。数据格式支持:
Segments 格式:用于普通文本数据,使用 segments 字段存储文本段落列表
Chat 格式:用于对话数据,使用 messages 字段存储对话消息列表
每个 Parquet 文件包含以下核心字段:
uuid: 唯一标识符
source: 数据来源标识
metadata: JSON 格式的元数据字典
segments 或 messages: 文本内容(根据数据类型选择)
详细的数据格式规范请参考 ../README.md。
数据集列表
数据集名称
样本数量
HuggingFace 仓库
reasoning_v1_20m
1,666… See the full description on the dataset page: https://huggingface.co/datasets/OpenOneRec/OpenOneRec-General-Pretrain.AITW_Generalgeneral-master-en-202608
General · Master · English · 2026-08
English pretraining text, assembled from three public sources, cleaned with one
character-level cleaner, and filtered for repetition.
109,337,531 documents and 468,064,046,462 characters.
Composition
Config
Documents
Characters
What it is
fineweb-edu-dedup
65,010,430
297,544,916,118
Web text an educational classifier kept
cosmopedia-v2
38,591,146
144,011,993,012
Synthetic prose from a seeded generator… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-master-en-202608.general-web-it-202608
General · Web · Italian · 2026-08
Italian pretraining text, built from the Italian portion of EPFL's FineWeb2-HQ, which is the
high quality slice of FineWeb-2. Every document passes one character-level cleaner and a
repetition filter.
21,065,052 documents and 66,158,573,443 characters of Italian prose.
Contents
Config
Documents
Characters
Upstream
fineweb2-hq-ita_Latn
21,065,052
66,158,573,443
epfml/FineWeb2-HQ, ita_Latn
The character count is… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-web-it-202608.general-web-fr-202608
General · Web · French · 2026-08
French pretraining text, built from the French portion of EPFL's FineWeb2-HQ, which is the
high quality slice of FineWeb-2. Every document passes one character-level cleaner and a
repetition filter.
31,999,309 documents and 118,346,333,763 characters of French prose.
Contents
Config
Documents
Characters
Upstream
fineweb2-hq-fra_Latn
31,999,309
118,346,333,763
epfml/FineWeb2-HQ, fra_Latn
The character count is exact.… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-web-fr-202608.SFT_Chinese_Generalmoda-general-capability-rollouts
MODA General Capability Retention Rollouts
This dataset contains the raw model generations and evaluation results for the
MODA general-capability retention experiments. It covers 16 models, seven
benchmarks, 260,592 prompt records, and 2,605,920 stored generations.
The evaluation code is pinned to source commit
12ea99b2a57a354f2b7d6792f62a3d9313192fa7.
Evaluation protocol
Benchmarks: GSM8K, MMLU abstract_algebra, GPQA Diamond, BoolQ,
HellaSwag, TruthfulQA, and… See the full description on the dataset page: https://huggingface.co/datasets/Hkang/moda-general-capability-rollouts.treaty-bodies-general-comments
Treaty Bodies General Comments
A paragraph-level dataset of General Comments and General Recommendations
adopted by the nine UN human-rights Treaty Bodies, with concerned-group
labels and document metadata. Companion to the
UNHRD search interface.
Licence
The curated dataset (paragraph segmentation, label annotation, document
metadata enrichment, footnote and section work) is released under
Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International… See the full description on the dataset page: https://huggingface.co/datasets/lszoszk/treaty-bodies-general-comments.sam_general_smolvlaThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "bi_sam_follower",
"total_episodes": 5,
"total_frames": 5282,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/girardijp/sam_general_smolvla.general-layerC-200klanguage:
en
license: apache-2.0
size_categories:
100K<n<1M
pretty_name: General LayerC — training-ready (200,000 samples)
tags:
task_categories:text-generation
task_categories:text2text-generation
language:en
license:apache-2.0
domain:general
reasoning
synthetic
sft
general-knowledge
science
creative-writing
glaive
deepseek-r1-distill
distillation
layer-c
layer-training-ready
General LayerC (training-ready) — 200k
Built from glaiveai/reasoning-v1-20m (spread-sampled across… See the full description on the dataset page: https://huggingface.co/datasets/hivemind-research/general-layerC-200k.SO100_evaluate_generalize_pick_posThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 90,
"total_frames": 33529,
"total_tasks": 1,
"total_videos": 180,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:90"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/masato-ka/SO100_evaluate_generalize_pick_pos.kikobot_final_50_generalThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "kikobot",
"total_episodes": 83,
"total_frames": 45277,
"total_tasks": 1,
"total_videos": 166,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:83"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/akabhinav32/kikobot_final_50_general.infini_gsm_4k_noise_generalSO100_evaluate_generalize_pick_pos_extendThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 30,
"total_frames": 11184,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/masato-ka/SO100_evaluate_generalize_pick_pos_extend.infini_gsm_8k_noise_generalteam15-groot-two-task-generalization
mvp
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
centrifuge-generalization-v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"ee_x.pos",
"ee_y.pos",
"ee_z.pos",
"ee_wx.pos",
"ee_wy.pos",
"ee_wz.pos",
"gripper.pos"
],
"shape": [
7… See the full description on the dataset page: https://huggingface.co/datasets/ar0s/centrifuge-generalization-v1.dataset
Harness Generalization Rollouts
Evolution-run rollouts for Qwen3-4B-Instruct-2507, Qwen2.5-3B-Instruct, gpt-oss-120b, and gpt-oss-20b. One Parquet file per run is stored at data/<task>/<model>/<configuration>/<timestamp>.parquet.
so101_generalThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/robot-learning-team43/so101_general.general_object_pickupThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"trossen_subversion": "v1.0",
"robot_type": "trossen_ai_stationary",
"total_episodes": 6,
"total_frames": 1706,
"total_tasks": 1,
"total_videos": 18,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:6"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/allday-technology/general_object_pickup.llama3-uf-meta-general-tokenizedGeneralThought-Sciencenemotron-sft-general-focused-stage1-2-ChatML-V2
Nemotron SFT Dataset (Chat Template Formatted)
Overview
This dataset is a curated supervised fine-tuning (SFT) dataset built from NVIDIA's Nemotron-Cascade-SFT-Stage-1 and Stage-2 datasets.
Important: This dataset uses the tokenizer's apply_chat_template() method to properly format conversations from the original messages/conversations fields.
Statistics
Total Samples: 100,000
Total Tokens: 222,951,633
Average Tokens per Sample: 2229.5
Tokenizer:… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nemotron-sft-general-focused-stage1-2-ChatML-V2.general-layerB-200klanguage:
en
license: apache-2.0
size_categories:
100K<n<1M
pretty_name: General LayerB — teacher-supervision (200,000 samples)
tags:
task_categories:text-generation
task_categories:text2text-generation
language:en
license:apache-2.0
domain:general
reasoning
synthetic
sft
general-knowledge
science
creative-writing
glaive
deepseek-r1-distill
distillation
layer-b
layer-teacher-supervision
teacher-supervision
General LayerB (teacher-supervision) — 200k
Built from… See the full description on the dataset page: https://huggingface.co/datasets/hivemind-research/general-layerB-200k.llama3-uf-meta-general-tokenized-1024-2048eval_smolvla_generalThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 19,
"total_frames": 2949,
"total_tasks": 8,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:19"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/eorikun/eval_smolvla_general.generals_io_replays
⚔️ Generals.io High-Rank Replay Dataset 🌟
Overview
This dataset contains a curated collection of 1v1 game replays from the online strategy game generals.io, specifically designed for training high-level reinforcement learning agents 🤖.
🏆 High-Quality Matches: Includes games where at least one participant had a star rating of 70 or higher 📈, ensuring a baseline of quality and strategic depth.
✅ Clean Data: Carefully filtered to remove outliers, games with AFK players… See the full description on the dataset page: https://huggingface.co/datasets/strakammm/generals_io_replays.assay-transfer-record-level-v27-carcinogens-general-intern
Carcinogens record-level V27-general
This level-agnostic release uses current V10 physical evidence and Gold-v1
parent-disjoint validation/test queries. Training pairs stay inside an intrinsic
V10 assay bucket and cross normalized parents. Gold level is optional provenance,
not a pairing or calibration boundary.
support gate: at least 10 records and 4 parents
variance ratio: undefined allowed; finite values above 0.75 excluded
train record-role degree cap: 24
ID/OOD validation… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/assay-transfer-record-level-v27-carcinogens-general-intern.
