datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agent-data-collection
Agent Data Collection
A comprehensive collection of agent interaction datasets for training and evaluating AI agents across diverse domains and tasks.
This dataset aggregates high-quality agent trajectories from various environments including web browsing, code generation, household tasks, knowledge base querying, and software engineering.
The dataset is collected through methods described in Agent Data Protocol.
Dataset Splits
Each dataset configuration provides up… See the full description on the dataset page: https://huggingface.co/datasets/neulab/agent-data-collection.dagger_final_1_21This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "openarms_follower",
"total_episodes": 1,
"total_frames": 841,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-data-collection/dagger_final_1_21.Long-Data-Collections-Pretrain-Without-Books
Dataset Card for "Long-Data-Collections-Pretrain-Without-Books"
Paraquet version of the pretrain split of togethercomputer/Long-Data-Collections WITHOUT books
Statistics (in # of characters): total_len: 236088622215, average_len: 25159.041601590307
level2_final_quality2_augmentedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "openarms_follower",
"total_episodes": 2366,
"total_frames": 6205242,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:2366"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-data-collection/level2_final_quality2_augmented.DHSA_Long-Data-Collections
DHSA_Long-Data-Collections
A length-bucketed release of togethercomputer/Long-Data-Collections, used in Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference (ICML 2026 Spotlight).
Each example is assigned to exactly one length bucket based on token count with meta-llama/Llama-3.1-8B-Instruct (add_special_tokens=True):
Bucket
Token range
lt_8k
[0, 8K)
8k_16k
[8K, 16K)
16k_32k
[16K, 32K)
32k_64k
[32K, 64K)… See the full description on the dataset page: https://huggingface.co/datasets/sxiong/DHSA_Long-Data-Collections.level2_final_qualityThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "openarms_follower",
"total_episodes": 1171,
"total_frames": 2940342,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1171"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-data-collection/level2_final_quality.Long-Data-Collections-Fine-Tune
Dataset Card for "Long-Data-Collections-Fine-Tune"
Paraquet version of the fine-tune split of togethercomputer/Long-Data-Collections
Statistics (in # of characters): total_len: 6419025428, average_len: 65130.08135393731
level12_quality0_2026-02-08This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "openarms_follower",
"total_episodes": 258,
"total_frames": 595888,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:258"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-data-collection/level12_quality0_2026-02-08.Clean-Long-Data-CollectionsData-Collection
Deprecated-API Code Generation Benchmark
Python functions mined from open-source repositories, each anchored on a library
API that has since been deprecated or replaced. Used to test whether a code LLM
still emits the outdated API, and to build forget / test splits per model.
Structure
outdated_all.json # O — samples calling the deprecated API
uptodated_all.json # U — samples calling the replacement API
<model>/ # codegen |… See the full description on the dataset page: https://huggingface.co/datasets/tummitum/Data-Collection.lerobot_data_collectionintegration-Data-Collection-Dataset
integration Data Collection Dataset
Merged LeRobot v3.0 dataset: 100 episodes, 26,953 frames, 30 FPS, robot type so_follower. Two 640×480 camera streams (front and wrist) and six-dimensional actions/state are preserved.
All episodes are in the train split. Episodes and global frame indices are continuous. Task labels are preserved exactly and mapped into one shared task table. No episodes were deduplicated or dropped. Original video files are copied without re-encoding;… See the full description on the dataset page: https://huggingface.co/datasets/Hailey-5-2026/integration-Data-Collection-Dataset.data-collectionsflowrl-data-collection
Data Preparation
Introduction
This is a collection for all data used in FlowRL paper. We detail the data source as follows:
data/
├── README.md
├── math_data/
│ ├── dapo-math-17k.parquet # Training data for math
│ ├── vaildation.parquet # Validation data (AIME 24/25, GPQA)
│ └── test.parquet # Test data (AIME24/25, AMC23, MATH500, etc.)
└── code_data/
├── deepcoder_train-00000-of-00005.parquet # DeepCoder… See the full description on the dataset page: https://huggingface.co/datasets/xuekai/flowrl-data-collection.datacollectionlevel12_rac_2_2026-02-07_and_mirThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "openarms_follower",
"total_episodes": 10770,
"total_frames": 26748966,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:10770"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-data-collection/level12_rac_2_2026-02-07_and_mir.JudgeLM-data-collection-v1.0
Dataset Card for JudgeLM-data-collection
Dataset Summary
This dataset is created for easily use and evaluate JudgeLM. We include LLMs-generated answers and a great multi-modal benchmark, MM-Vet in this repo. The folder structure is shown as bellow:
Folder structure
data
├── JudgeLM/
│ ├── answers/
│ │ ├── alpaca_judgelm_val.jsonl
| | ├── ...
│ ├── judgelm_preprocess.py
│ ├── judgelm_val_5k.jsonl
│ ├── judgelm_val_5k_gpt4.jsonl
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/JudgeLM-data-collection-v1.0.spectral-data-collection
高光谱遥感数据集 (Spectral Data Collection)
概述
本数据集包含遥感领域常用的高光谱影像数据和光谱参考数据,用于地物分类、植被监测、城市分析等研究。
数据集列表
1. Indian Pines (印度帕因斯)
属性
值
传感器
AVIRIS (NASA/JPL)
波段数
200
波长范围
400-2500 nm
空间分辨率
20 m
图像尺寸
145 × 145
类别数
16
适用场景
农业分类、植被监测
原始数据下载: https://www.ehu.eus/ccwintco/index.php?oldid=16536
2. Salinas (萨利纳斯)
属性
值
传感器
AVIRIS
波段数
204
波长范围
400-2500 nm
空间分辨率
3.7 m
图像尺寸
512 × 217
类别数
16… See the full description on the dataset page: https://huggingface.co/datasets/shangshang/spectral-data-collection.level2_final_quality3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "openarms_follower",
"total_episodes": 1200,
"total_frames": 3254196,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1200"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-data-collection/level2_final_quality3.level12_rac_2_2026-02-07This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "openarms_follower",
"total_episodes": 5551,
"total_frames": 13805156,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:5551"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-data-collection/level12_rac_2_2026-02-07.reasoning-data-collectionsOdia-data-collectionThis is a repo created for training purpose for the language "Odia".
dataset's collection, that exists:
Link
level12_rac_2_2026-02-05_agThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "openarms_follower",
"total_episodes": 5413,
"total_frames": 13497374,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:5413"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-data-collection/level12_rac_2_2026-02-05_ag.level2_final_quality3_t_0_hil_data_cThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "openarms_follower",
"total_episodes": 1319,
"total_frames": 3414338,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 0.001,
"fps": 30,
"splits": {
"train": "0:1319"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-data-collection/level2_final_quality3_t_0_hil_data_c.dagger_final_1_17This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "openarms_follower",
"total_episodes": 1,
"total_frames": 3542,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-data-collection/dagger_final_1_17.folding_2025-12-06_level2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "openarms_follower",
"total_episodes": 1380,
"total_frames": 4879690,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1380"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-data-collection/folding_2025-12-06_level2.long-data-collection-finetune-50k
Dataset Card for "long-data-collection-finetune-50k"
More Information needed
The dataset is a 50k row collection of the finetuning subset created by togethercomputer and which can be found at the following URL https://huggingface.co/datasets/togethercomputer/Long-Data-Collections in the fine-tune path
The exercise consisted of taking the data set and being able to set the format for finetuning llama2 with the aim of setting only one column (text), with the full format.
Additionally… See the full description on the dataset page: https://huggingface.co/datasets/yvillamil/long-data-collection-finetune-50k.level2_rac3_8This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "openarms_follower",
"total_episodes": 8,
"total_frames": 33260,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:8"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-data-collection/level2_rac3_8.dagger_final_1_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "openarms_follower",
"total_episodes": 5,
"total_frames": 13726,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-data-collection/dagger_final_1_2.level2_final_quality2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "openarms_follower",
"total_episodes": 1183,
"total_frames": 3102621,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1183"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-data-collection/level2_final_quality2.
