datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
shofo-tiktok-general-small
Shofo TikTok General (Small)
Overview
Shofo TikTok General (Small) is a dataset containing 50,000 TikTok videos with comprehensive metadata, transcripts, comments, and engagement metrics. This is a curated subset of Shofo's larger TikTok index, which contains hundreds of millions of indexed videos.
Size: ~50K videos (~500GB)
Modality: Video + Audio + Text (transcripts, comments, captions)
Source: TikTok
Schema
Column
Type
Description
file_name… See the full description on the dataset page: https://huggingface.co/datasets/Shofo/shofo-tiktok-general-small.GeneralThought-430K-filteredData from https://huggingface.co/datasets/GeneralReasoning/GeneralThought-430K, removed prompts with non commercial data
GMAI-VL-5.5M
GMAI-VL-5.5M Dataset
GMAI-VL-5.5M is a comprehensive, large-scale medical General Medical AI Vision-Language (GMAI-VL) dataset built specifically for training multimodal foundation models in the medical domain. It contains an extraordinary scale of high-quality instructions encompassing over 5.5 million multimodal question-answering pairs, carefully constructed based on hundreds of medical classification, segmentation, and detection datasets.
This repository… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-VL-5.5M.GeneralThoughtArchive
GeneralThought-430K
Thought wants to be free
Open reasoning data for March 14 2025. This dataset was part of a side-project in the weeks following the R1 release by Chengxi and Ross - we are no longer maintaining this
dataset but are archiving it here.
The dataset contains questions, reference answers, reasoning traces, final answers and other metadata from several popular reasoning models including DeepSeek-R1, DeepSeek-R1-Zero, OpenThoughts-32B, LIMO… See the full description on the dataset page: https://huggingface.co/datasets/RJT1990/GeneralThoughtArchive.OpenOneRec-General-Pretrain
通用文本数据集
本目录包含 OpenOneRec 项目使用的通用文本数据集信息。这些数据集均来自 HuggingFace,经过清洗处理和对齐到项目统一的数据格式,并转换为 Parquet 格式用于训练。
数据格式说明
所有数据集均已转换为统一的 Parquet 格式,符合项目的数据格式规范(参考 ../README.md)。数据格式支持:
Segments 格式:用于普通文本数据,使用 segments 字段存储文本段落列表
Chat 格式:用于对话数据,使用 messages 字段存储对话消息列表
每个 Parquet 文件包含以下核心字段:
uuid: 唯一标识符
source: 数据来源标识
metadata: JSON 格式的元数据字典
segments 或 messages: 文本内容(根据数据类型选择)
详细的数据格式规范请参考 ../README.md。
数据集列表
数据集名称
样本数量
HuggingFace 仓库
reasoning_v1_20m
1,666… See the full description on the dataset page: https://huggingface.co/datasets/OpenOneRec/OpenOneRec-General-Pretrain.steady-rans-generalization
Steady-RANS cross-family generalization dataset
Data for the paper "Towards generalized flow field prediction: one model across unseen
object families" (under double blind review; this account is anonymous for that reason).
Trained checkpoints and evaluation code are in the companion model repo:
steady-rans-surrogates.
Steady incompressible k-omega SST (OpenFOAM simpleFoam) external flow around 855 distinct
shapes (17 scripted parametric families plus 40 ModelNet object… See the full description on the dataset page: https://huggingface.co/datasets/BlidReview/steady-rans-generalization.AITW_Generalgeneral-master-en-202608
General · Master · English · 2026-08
English pretraining text, assembled from three public sources, cleaned with one
character-level cleaner, and filtered for repetition.
109,337,531 documents and 468,064,046,462 characters.
Composition
Config
Documents
Characters
What it is
fineweb-edu-dedup
65,010,430
297,544,916,118
Web text an educational classifier kept
cosmopedia-v2
38,591,146
144,011,993,012
Synthetic prose from a seeded generator… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-master-en-202608.moda-general-capability-rollouts
MODA General Capability Retention Rollouts
This dataset contains the raw model generations and evaluation results for the
MODA general-capability retention experiments. It covers 16 models, seven
benchmarks, 260,592 prompt records, and 2,605,920 stored generations.
The evaluation code is pinned to source commit
12ea99b2a57a354f2b7d6792f62a3d9313192fa7.
Evaluation protocol
Benchmarks: GSM8K, MMLU abstract_algebra, GPQA Diamond, BoolQ,
HellaSwag, TruthfulQA, and… See the full description on the dataset page: https://huggingface.co/datasets/Hkang/moda-general-capability-rollouts.general-web-it-202608
General · Web · Italian · 2026-08
Italian pretraining text, built from the Italian portion of EPFL's FineWeb2-HQ, which is the
high quality slice of FineWeb-2. Every document passes one character-level cleaner and a
repetition filter.
21,065,052 documents and 66,158,573,443 characters of Italian prose.
Contents
Config
Documents
Characters
Upstream
fineweb2-hq-ita_Latn
21,065,052
66,158,573,443
epfml/FineWeb2-HQ, ita_Latn
The character count is… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-web-it-202608.general-web-fr-202608
General · Web · French · 2026-08
French pretraining text, built from the French portion of EPFL's FineWeb2-HQ, which is the
high quality slice of FineWeb-2. Every document passes one character-level cleaner and a
repetition filter.
31,999,309 documents and 118,346,333,763 characters of French prose.
Contents
Config
Documents
Characters
Upstream
fineweb2-hq-fra_Latn
31,999,309
118,346,333,763
epfml/FineWeb2-HQ, fra_Latn
The character count is exact.… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-web-fr-202608.SFT_Chinese_Generalgeneralization-science-datasam_general_smolvlaThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "bi_sam_follower",
"total_episodes": 5,
"total_frames": 5282,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/girardijp/sam_general_smolvla.reward-projection-goal-generalisation-vlmgeneral-layerC-200klanguage:
en
license: apache-2.0
size_categories:
100K<n<1M
pretty_name: General LayerC — training-ready (200,000 samples)
tags:
task_categories:text-generation
task_categories:text2text-generation
language:en
license:apache-2.0
domain:general
reasoning
synthetic
sft
general-knowledge
science
creative-writing
glaive
deepseek-r1-distill
distillation
layer-c
layer-training-ready
General LayerC (training-ready) — 200k
Built from glaiveai/reasoning-v1-20m (spread-sampled across… See the full description on the dataset page: https://huggingface.co/datasets/hivemind-research/general-layerC-200k.SO100_evaluate_generalize_pick_posThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 90,
"total_frames": 33529,
"total_tasks": 1,
"total_videos": 180,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:90"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/masato-ka/SO100_evaluate_generalize_pick_pos.team15-groot-two-task-generalization
mvp
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
kikobot_final_50_generalThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "kikobot",
"total_episodes": 83,
"total_frames": 45277,
"total_tasks": 1,
"total_videos": 166,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:83"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/akabhinav32/kikobot_final_50_general.nemotron-sft-general-focused-stage1-2-ChatML-V3
Nemotron SFT Dataset (Chat Template Formatted)
Overview
This dataset is a curated supervised fine-tuning (SFT) dataset built from NVIDIA's Nemotron-Cascade-SFT-Stage-1 and Stage-2 datasets.
Important: This dataset uses the tokenizer's apply_chat_template() method to properly format conversations from the original messages/conversations fields.
Statistics
Total Samples: 496,385
Total Tokens: 1,114,218,401
Average Tokens per Sample: 2244.7
Tokenizer:… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nemotron-sft-general-focused-stage1-2-ChatML-V3.general_speech_distortedseedance_general_all_dance_scm_latent_lmdb
Seedance General-All + Dance SCM Latent LMDB
This dataset stores precomputed SCM latents used for TurboT2AV training.
Source mapping: seedance_general_all_dance_mapping.csv
Successful latent samples: 44,305
Shards: 8 LMDB shards under scm_latent_lmdb/shard_00000 ... shard_00007
Video latent shape per sample: (1, 16, 128, 16, 24)
Audio latent shape per sample: (1, 127, 128)
The source mapping combines Seedance general-all data with a dance subset. The mapping contains 44,504… See the full description on the dataset page: https://huggingface.co/datasets/luyu1021/seedance_general_all_dance_scm_latent_lmdb.SO100_evaluate_generalize_pick_pos_extendThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 30,
"total_frames": 11184,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:30"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/masato-ka/SO100_evaluate_generalize_pick_pos_extend.showdown-clicks
showdown-clicks
General Agents
🤗 Dataset | GitHub
showdown is a suite of offline and online benchmarks for computer-use agents.
showdown-clicks is a collection of 5,679 left clicks of humans performing various tasks in a macOS desktop environment. It is intended to evaluate instruction-following and low-level control capabilities of computer-use agents.
As of March 2025, we are releasing a subset of the full set, showdown-clicks-dev, containing 557 clicks. All examples are… See the full description on the dataset page: https://huggingface.co/datasets/generalagents/showdown-clicks.vn-provinces-ethnic-minority-general-school-pupils
Vietnam ethnic-minority general school pupils by level
Vietnam ethnic-minority general school pupils by level. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Hero (continued)
Comparison
Color key
Files
provinces (1152 rows)
data/provinces.csv
data/provinces.dta
data/provinces.xlsx… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-ethnic-minority-general-school-pupils.infini_gsm_4k_noise_generalGeneralAgentBench
GeneralAgentBench
GeneralAgentBench is a 1,400+ task benchmark for evaluating whether general-purpose AI agents have genuinely completed a task, spanning Mobile / Browser / Desktop environments. It is the evaluation resource accompanying an anonymous NeurIPS 2026 Evaluations & Datasets Track submission (Submission 173, AgentJudge). This release is fully anonymized for double-blind review.
Each task provides a natural-language instruction plus a list of verification
checkpoints.… See the full description on the dataset page: https://huggingface.co/datasets/agentjudge-anon/GeneralAgentBench.vn-provinces-general-schools
Vietnam provinces general schools by school type
Vietnam provinces general schools by school type. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Hero (continued)
Comparison
Color key
Files
provinces (1260 rows)
data/provinces.csv
data/provinces.dta
data/provinces.xlsx
regions… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-general-schools.dataset
Harness Generalization Rollouts
Evolution-run rollouts for Qwen3-4B-Instruct-2507, Qwen2.5-3B-Instruct, gpt-oss-120b, and gpt-oss-20b. One Parquet file per run is stored at data/<task>/<model>/<configuration>/<timestamp>.parquet.
vn-provinces-ethnic-minority-general-school-teachers
Vietnam ethnic-minority general school teachers (selected provinces)
Vietnam ethnic-minority general school teachers (selected provinces). Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Hero (continued)
Comparison
Color key
Files
provinces (702 rows)
data/provinces.csv… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-ethnic-minority-general-school-teachers.
