datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
YuLan-Mini-Text-Datasets
News
[2025.04.11] Add dataset mixture: link.
[2025.03.30] Text datasets upload finished.
This is text dataset.
这是文本格式的数据集。
Since we have used BPE-Dropout, in order to ensure accuracy, you can find the tokenized dataset here.
由于我们使用了BPE-Dropout,为了保证准确性,你可以在这里找到分词后的数据。
For more information, please refer to our datasets details and preprocess details.
Contributing
We welcome any form of contribution, including feedback on model bad cases, feature suggestions, and example… See the full description on the dataset page: https://huggingface.co/datasets/yulan-team/YuLan-Mini-Text-Datasets.AM-DeepSeek-Distilled-40MFor more open-source datasets, models, and methodologies, please visit our GitHub repository and paper: DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training.
Due to certain constraints, we are only able to open-source a subset of the complete dataset.
Model Training Performance based on our complete dataset
On AIME 2024, our 72B model achieved a score of 79.2 using only supervised fine-tuning (SFT). The 32B model reached 75.8 and… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-Distilled-40M.SwissCrop25
SwissCrop25
A national benchmark dataset for operational crop mapping in Switzerland, providing Sentinel-2
time series, daily temperature data, and parcel-level crop type labels across seven growing
seasons (2019–2025).
Introduced in: SwissCrop25: A National Multi-Year Benchmark for Operational Crop Mapping
(TerraBytes II Workshop, ECCV 2026) — [Paper] [Code] [Team]
Highlights
Nationwide coverage of Switzerland (41,285 km²)
Seven growing seasons (2019–2025)
73… See the full description on the dataset page: https://huggingface.co/datasets/EOA-team/SwissCrop25.moss-002-sft-data
Dataset Card for "moss-002-sft-data"
Dataset Summary
An open-source conversational dataset that was used to train MOSS-002. The user prompts are extended based on a small set of human-written seed prompts in a way similar to Self-Instruct. The AI responses are generated using text-davinci-003. The user prompts of en_harmlessness are from Anthropic red teaming data.
Data Splits
name
# samples
en_helpfulness.json
419049
en_honesty.json
112580… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/moss-002-sft-data.MVU-Eval-Data
MVU-Eval Dataset
Paper | Code | Project Page
Dataset Description
The advent of Multimodal Large Language Models (MLLMs) has expanded AI capabilities to visual modalities, yet existing evaluation benchmarks remain limited to single-video understanding, overlooking the critical need for multi-video understanding in real-world scenarios (e.g., sports analytics and autonomous driving). To address this significant gap, we introduce MVU-Eval, the first comprehensive benchmark… See the full description on the dataset page: https://huggingface.co/datasets/MVU-Eval-Team/MVU-Eval-Data.FutureOmni
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
Predicting the future requires listening as well as seeing.
📖 Dataset Summary
Although Multimodal Large Language Models (MLLMs) demonstrate strong omni-modal perception, their ability to forecast future events from audio–visual cues remains largely unexplored, as existing benchmarks focus mainly on retrospective understanding.
FutureOmni is the first benchmark designed… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/FutureOmni.adapter-based-multimodal-fusion
Falcon-Audio Training Dataset
Training-ready Parquet shards for Falcon-Audio. Rows contain Gemma-tokenized inputs/labels and fp16 Whisper encoder features encoded as raw bytes.
FINSABER-V2-Data
FINSABER V2 Data
FINSABER V2 Data is the parquet dataset used by the upgraded FINSABER-2 backtesting framework. It contains S&P 500 daily market data, financial news items, and SEC filing text organized by year for reproducible financial strategy research.
Dataset Structure
The dataset is partitioned by modality and year:
price_daily/year=<YYYY>/part-000.parquet
news_items/year=<YYYY>/part-000.parquet
filingk/year=<YYYY>/part-000.parquet… See the full description on the dataset page: https://huggingface.co/datasets/finsaber-team/FINSABER-V2-Data.FINSABER-reproduce
FINSABER Data
Aggregated datasets for the FINSABER backtesting framework (KDD 2026).
Files
File
Description
Size
data/finmem_data/stock_data_sp500_2000_2024.pkl
S&P500 full aggregated data (Price + News + Filings)
~11 GB
data/finmem_data/stock_data_cherrypick_2000_2024.pkl
Selected symbols (TSLA, AMZN, MSFT, NFLX, COIN)
~53 MB
data/price/all_sp500_prices_2000_2024_delisted_include.csv
CSV price-only data for S&P500 (including delisted)
~253 MB… See the full description on the dataset page: https://huggingface.co/datasets/finsaber-team/FINSABER-reproduce.trial-v0-20250313
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/team-wonders/trial-v0-20250313.team-7-right-arm-grasp-tapeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 100,
"total_frames": 59711,
"total_tasks": 1,
"total_videos": 200,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/roboticshack/team-7-right-arm-grasp-tape.every_frame_right_rgbteam2-guess_who_so100_lightThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 96,
"total_frames": 19853,
"total_tasks": 1,
"total_videos": 96,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:96"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/roboticshack/team2-guess_who_so100_light.team13-two-balls-stackingThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 51,
"total_frames": 13476,
"total_tasks": 1,
"total_videos": 102,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:51"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/roboticshack/team13-two-balls-stacking.team16-water-pouringThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 50,
"total_frames": 29887,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/roboticshack/team16-water-pouring.team13-three-balls-stackingThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 101,
"total_frames": 48652,
"total_tasks": 1,
"total_videos": 202,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:101"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/roboticshack/team13-three-balls-stacking.team16-cleaning-plateThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 100,
"total_frames": 74893,
"total_tasks": 1,
"total_videos": 200,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/roboticshack/team16-cleaning-plate.lexgraph
Lexgraph dataset
The data plane of Lexgraph —
German legislation modelled as Laws as Git: a temporal, multi-authority
event log with HEAD, commits, open/closed branches and evidence-bound merges
(Bund / Bayern / EU; Länder records only after verification at the originating
Landtag). Built 2026-07-19.
Git is a navigation metaphor, not a substitute for legal status. Every row's
official source and status controls whether it is current law, a pending branch
or a documented… See the full description on the dataset page: https://huggingface.co/datasets/SNTIQ-Team/lexgraph.omx_f_team10_push_placeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omx_f",
"total_episodes": 26,
"total_frames": 15607,
"total_tasks": 1,
"total_videos": 26,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:26"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/haneulborri/omx_f_team10_push_place.SciJudgeBench
SciJudgeBench Dataset
Training and evaluation data for scientific paper citation prediction, from the paper AI Can Learn Scientific Taste.
Given two academic papers (title, abstract, publication date), the task is to predict which paper has a higher citation count.
Resources: Project page, GitHub repository, SciJudge-4B-2605, and SciJudge-30B-2605.
Dataset Splits
Split
Examples
Description
train
720,341
Training preference pairs from arXiv papers
test… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/SciJudgeBench.team2-guess_who_so100This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 96,
"total_frames": 23114,
"total_tasks": 1,
"total_videos": 96,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:96"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/roboticshack/team2-guess_who_so100.team16-can-stackingThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 40,
"total_frames": 23959,
"total_tasks": 1,
"total_videos": 80,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:40"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/roboticshack/team16-can-stacking.team-7-left-arm-grasp-motorThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 99,
"total_frames": 59123,
"total_tasks": 1,
"total_videos": 198,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:99"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/roboticshack/team-7-left-arm-grasp-motor.mastodon-instances
Dataset Card for "mastodon-instances"
More Information needed
team_20_augmented_datasocksomx_f_team10_pick_place2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omx_f",
"total_episodes": 25,
"total_frames": 14922,
"total_tasks": 1,
"total_videos": 25,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:25"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/haneulborri/omx_f_team10_pick_place2.blue_pick
blue_pick
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
Hackathon_Team00blue_pick_250829_ep300
blue_pick_250829_ep300
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
libero-lerobot-v3.0
LIBERO LeRobot v3.0
This dataset packages the four standard LIBERO demonstration suites in LeRobot v3.0 format for robot-policy and world-action-model training. It contains synchronized RGB observations, robot state, actions, timestamps, task indices, and natural-language task metadata.
Dataset Summary
Suite
Episodes
Frames
Tasks
Videos
LIBERO-Spatial
434
53,229
10
868
LIBERO-Object
457
67,309
10
914
LIBERO-Goal
433
52,895
10
866
LIBERO-10
388
104… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/libero-lerobot-v3.0.
