datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AM-DeepSeek-Distilled-40MFor more open-source datasets, models, and methodologies, please visit our GitHub repository and paper: DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training.
Due to certain constraints, we are only able to open-source a subset of the complete dataset.
Model Training Performance based on our complete dataset
On AIME 2024, our 72B model achieved a score of 79.2 using only supervised fine-tuning (SFT). The 32B model reached 75.8 and… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-Distilled-40M.YuLan-Mini-Text-Datasets
News
[2025.04.11] Add dataset mixture: link.
[2025.03.30] Text datasets upload finished.
This is text dataset.
这是文本格式的数据集。
Since we have used BPE-Dropout, in order to ensure accuracy, you can find the tokenized dataset here.
由于我们使用了BPE-Dropout,为了保证准确性,你可以在这里找到分词后的数据。
For more information, please refer to our datasets details and preprocess details.
Contributing
We welcome any form of contribution, including feedback on model bad cases, feature suggestions, and example… See the full description on the dataset page: https://huggingface.co/datasets/yulan-team/YuLan-Mini-Text-Datasets.SwissCrop25
SwissCrop25
A national benchmark dataset for operational crop mapping in Switzerland, providing Sentinel-2
time series, daily temperature data, and parcel-level crop type labels across seven growing
seasons (2019–2025).
Introduced in: SwissCrop25: A National Multi-Year Benchmark for Operational Crop Mapping
(TerraBytes II Workshop, ECCV 2026) — [Paper] [Code] [Team]
Highlights
Nationwide coverage of Switzerland (41,285 km²)
Seven growing seasons (2019–2025)
73… See the full description on the dataset page: https://huggingface.co/datasets/EOA-team/SwissCrop25.Agilex_Cobot_Magic_pour_tea
Split_aloha_pour_tea
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: agilex_cobot_decoupled_magic
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
office
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
place
pick
pour
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Agilex_Cobot_Magic_pour_tea.RMC-AIDA-L_pour_tea
RMC-AIDA-L_pour_tea
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: realman_rmc_aidal
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
restaurant
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
place
pick
pour
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/RMC-AIDA-L_pour_tea.moss-002-sft-data
Dataset Card for "moss-002-sft-data"
Dataset Summary
An open-source conversational dataset that was used to train MOSS-002. The user prompts are extended based on a small set of human-written seed prompts in a way similar to Self-Instruct. The AI responses are generated using text-davinci-003. The user prompts of en_harmlessness are from Anthropic red teaming data.
Data Splits
name
# samples
en_helpfulness.json
419049
en_honesty.json
112580… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/moss-002-sft-data.MVU-Eval-Data
MVU-Eval Dataset
Paper | Code | Project Page
Dataset Description
The advent of Multimodal Large Language Models (MLLMs) has expanded AI capabilities to visual modalities, yet existing evaluation benchmarks remain limited to single-video understanding, overlooking the critical need for multi-video understanding in real-world scenarios (e.g., sports analytics and autonomous driving). To address this significant gap, we introduce MVU-Eval, the first comprehensive benchmark… See the full description on the dataset page: https://huggingface.co/datasets/MVU-Eval-Team/MVU-Eval-Data.FutureOmni
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
Predicting the future requires listening as well as seeing.
📖 Dataset Summary
Although Multimodal Large Language Models (MLLMs) demonstrate strong omni-modal perception, their ability to forecast future events from audio–visual cues remains largely unexplored, as existing benchmarks focus mainly on retrospective understanding.
FutureOmni is the first benchmark designed… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/FutureOmni.R1_Lite_make_tea
R1_Lite_make_tea
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: galaxea_r1_lite
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
restaurant
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset Statistics
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/R1_Lite_make_tea.FINSABER-V2-Data
FINSABER V2 Data
FINSABER V2 Data is the parquet dataset used by the upgraded FINSABER-2 backtesting framework. It contains S&P 500 daily market data, financial news items, and SEC filing text organized by year for reproducible financial strategy research.
Dataset Structure
The dataset is partitioned by modality and year:
price_daily/year=<YYYY>/part-000.parquet
news_items/year=<YYYY>/part-000.parquet
filingk/year=<YYYY>/part-000.parquet… See the full description on the dataset page: https://huggingface.co/datasets/finsaber-team/FINSABER-V2-Data.FINSABER-reproduce
FINSABER Data
Aggregated datasets for the FINSABER backtesting framework (KDD 2026).
Files
File
Description
Size
data/finmem_data/stock_data_sp500_2000_2024.pkl
S&P500 full aggregated data (Price + News + Filings)
~11 GB
data/finmem_data/stock_data_cherrypick_2000_2024.pkl
Selected symbols (TSLA, AMZN, MSFT, NFLX, COIN)
~53 MB
data/price/all_sp500_prices_2000_2024_delisted_include.csv
CSV price-only data for S&P500 (including delisted)
~253 MB… See the full description on the dataset page: https://huggingface.co/datasets/finsaber-team/FINSABER-reproduce.R1_Lite_tea_service_table_setting
R1_Lite_tea_service_table_setting
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: galaxea_r1_lite
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset Statistics
Metric… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/R1_Lite_tea_service_table_setting.adapter-based-multimodal-fusion
Falcon-Audio Training Dataset
Training-ready Parquet shards for Falcon-Audio. Rows contain Gemma-tokenized inputs/labels and fp16 Whisper encoder features encoded as raw bytes.
so101_tea2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101",
"total_episodes": 101,
"total_frames": 78833,
"total_tasks": 1,
"total_videos": 202,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:101"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Cornito/so101_tea2.teacher-coverage-hash-cache
teacher-coverage hash cache
The cache behind the teacher-coverage analysis for OpenThoughts-Agent
(code and result tables: https://github.com/franziweindel/otagent-teacher-coverage):
which teacher models have trajectories for which agentic tasks, across the
datagen accounts on the Hub (DCAgent, DCAgent2, mlfoundations-dev, laion,
marin-community, open-thoughts, ...). Tasks are identified by the SHA-1 of
their whitespace-collapsed instruction.md text, never by task id (ids are… See the full description on the dataset page: https://huggingface.co/datasets/FWeindel/teacher-coverage-hash-cache.trial-v0-20250313
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/team-wonders/trial-v0-20250313.team2-guess_who_so100_lightThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 96,
"total_frames": 19853,
"total_tasks": 1,
"total_videos": 96,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:96"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/roboticshack/team2-guess_who_so100_light.team-7-right-arm-grasp-tapeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 100,
"total_frames": 59711,
"total_tasks": 1,
"total_videos": 200,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/roboticshack/team-7-right-arm-grasp-tape.put_coffee_cap_teaboxThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 40,
"total_frames": 14806,
"total_tasks": 1,
"total_videos": 80,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:40"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lirislab/put_coffee_cap_teabox.team13-two-balls-stackingThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 51,
"total_frames": 13476,
"total_tasks": 1,
"total_videos": 102,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:51"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/roboticshack/team13-two-balls-stacking.every_frame_right_rgbbactrainus-hotpotqa-teacher-traces
Bactrainus HotpotQA Teacher Traces
SOURCE-LINKED v1.0.0
Archived Llama 3.1 rationale and question-decomposition supervision, paired with complete SFT conversations and stable HotpotQA identities.
198,660 ROWS
4 CONFIGURATIONS
SFT MESSAGES
8B + 70B LABELS
CC BY-SA 4.0
A focused release of recovered teacher-generated supervision for multi-hop question answering. Every row contains the normalized annotation, an ordered… See the full description on the dataset page: https://huggingface.co/datasets/bactrianus/bactrainus-hotpotqa-teacher-traces.team13-three-balls-stackingThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 101,
"total_frames": 48652,
"total_tasks": 1,
"total_videos": 202,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:101"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/roboticshack/team13-three-balls-stacking.team16-water-pouringThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 50,
"total_frames": 29887,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/roboticshack/team16-water-pouring.open_top_drawer_teaboxThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 30,
"total_frames": 9600,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lirislab/open_top_drawer_teabox.lexgraph
Lexgraph dataset
The data plane of Lexgraph —
German legislation modelled as Laws as Git: a temporal, multi-authority
event log with HEAD, commits, open/closed branches and evidence-bound merges
(Bund / Bayern / EU; Länder records only after verification at the originating
Landtag). Built 2026-07-19.
Git is a navigation metaphor, not a substitute for legal status. Every row's
official source and status controls whether it is current law, a pending branch
or a documented… See the full description on the dataset page: https://huggingface.co/datasets/SNTIQ-Team/lexgraph.team16-cleaning-plateThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 100,
"total_frames": 74893,
"total_tasks": 1,
"total_videos": 200,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/roboticshack/team16-cleaning-plate.final_so100_test_tea2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 66,
"total_frames": 31250,
"total_tasks": 1,
"total_videos": 132,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:66"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/gdut508/final_so100_test_tea2.team2-guess_who_so100This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 96,
"total_frames": 23114,
"total_tasks": 1,
"total_videos": 96,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:96"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/roboticshack/team2-guess_who_so100.omx_f_team10_push_placeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omx_f",
"total_episodes": 26,
"total_frames": 15607,
"total_tasks": 1,
"total_videos": 26,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:26"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/haneulborri/omx_f_team10_push_place.
