datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Cobot_Magic_take_out_a_pen_from_the_pen_holder
Cobot_Magic_take_out_a_pen_from_the_pen_holder
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: agilex_cobot_decoupled_magic
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
office
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_take_out_a_pen_from_the_pen_holder.TextOnly_FromRLBench_CloseBox_24K_unfixedus-historical-layoffs-archive-warn-act-notices-removed-from-state-websites
Historical US layoffs archive: 6,799 WARN Act notices that state websites no longer list (2000-2025), recovered
Rebuilt 2026-09-25. Five state labor agencies — Connecticut, Michigan, New York,
North Carolina and Pennsylvania — retired the web pages their older WARN Act
layoff notices lived on. Their current pages start years later. This dataset is
every notice in our file that came from one of those retired pages and is not
on the agency's live page today: 6,799 notices, 6,799… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites.text-code-galeras-code-generation-from-docstring-3k-dedupedOMat24_train_aimd_from_PBE_1000_npt
Cite this dataset Barroso-Luque, L., Shuaibi, M., Fu, X., Wood, B. M., Dzamba, M., Gao, M., Rizvi, A., Zitnick, C. L., and Ulissi, Z. W. OMat24 train aimd from PBE 1000 npt. ColabFit, 2025. https://doi.org/10.60732/25f16f85
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_jqrkc9e7cgmh_0
Visit the ColabFit Exchange to search… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/OMat24_train_aimd_from_PBE_1000_npt.ecommerce-behavior-data-from-multi-category-store_oct-nov_2019
eCommerce Behavior Data from Multi-Category Store
About the Dataset
This dataset contains behavioral data for 285 million user events from a large multi-category eCommerce store. The data spans 7 months (October 2019 - April 2020) and records various user interactions with products.
Dataset Overview
Time Frame: October 2019 - April 2020
Total Events: 285 million
Event Granularity: Each row represents an event associated with a product and a user.
Data Source:… See the full description on the dataset page: https://huggingface.co/datasets/kevykibbz/ecommerce-behavior-data-from-multi-category-store_oct-nov_2019.task04_pour_liquid_from_tubes_to_beakerThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "mobileai_robot",
"total_episodes": 130,
"total_frames": 41199,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:130"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kiroaiseoul/task04_pour_liquid_from_tubes_to_beaker.python-copilot-training-from-many-repos-large
Python Copilot Large Coding Dataset
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains python code, either a class method or a global function, imported modules, base classes (if any), exceptions (ordered based off the code), returns (ordered based off the code), arguments (ordered based off the code), and more.
Rows: 2350782… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-copilot-training-from-many-repos-large.OMat24_train_aimd_from_PBE_3000_nvt
Cite this dataset Barroso-Luque, L., Shuaibi, M., Fu, X., Wood, B. M., Dzamba, M., Gao, M., Rizvi, A., Zitnick, C. L., and Ulissi, Z. W. OMat24 train aimd from PBE 3000 nvt. ColabFit, 2025. https://doi.org/10.60732/105da475
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_fi9ozets1s5f_0
Visit the ColabFit Exchange to search… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/OMat24_train_aimd_from_PBE_3000_nvt.OMat24_train_aimd_from_PBE_3000_npt
Cite this dataset Barroso-Luque, L., Shuaibi, M., Fu, X., Wood, B. M., Dzamba, M., Gao, M., Rizvi, A., Zitnick, C. L., and Ulissi, Z. W. OMat24 train aimd from PBE 3000 npt. ColabFit, 2025. https://doi.org/10.60732/edd12490
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_6xvvh8yl7rfd_0
Visit the ColabFit Exchange to search… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/OMat24_train_aimd_from_PBE_3000_npt.pick_specific_item_from_clutterThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"trossen_subversion": "v1.0",
"robot_type": "trossen_ai_solo",
"total_episodes": 6,
"total_frames": 2299,
"total_tasks": 5,
"total_videos": 18,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:6"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/k1seul/pick_specific_item_from_clutter.tactile_grab_pen_from_bagThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_tactile_follower",
"total_episodes": 50,
"total_frames": 47635,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ydaichi/tactile_grab_pen_from_bag.asia-faostat-emissions-from-energy-use-in-agriculture-gn
Emissions from Energy use in agriculture — Asia
Source: FAOSTAT — Emissions from Energy use in agriculture
Domain code: GN
Publisher: Food and Agriculture Organization of the United Nations (FAO)
Coverage: 41 Asian countries · 1990–2023 · 20,632 rows
Items: 7 · Elements: 4
About
FAOSTAT is the world's largest and most comprehensive statistical database on food, agriculture, fisheries,
forestry, and rural development. This dataset contains the Emissions from Energy use… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-faostat-emissions-from-energy-use-in-agriculture-gn.agibot-sim-pickup-items-from-the-freezerThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "a2d",
"total_episodes": 100,
"total_frames": 280654,
"total_tasks": 1,
"total_videos": 300,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30.0,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/bot-pi/agibot-sim-pickup-items-from-the-freezer.FR5_task2_transfer_the_gray_tablets_from_the_brown_bottle_into_another_brown_bottle_100This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "fairino_follower",
"total_episodes": 100,
"total_frames": 146016,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/coport-uni/FR5_task2_transfer_the_gray_tablets_from_the_brown_bottle_into_another_brown_bottle_100.demo_insert_connector_from_rest_dag1_centerThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "sam_evt2",
"total_episodes": 51,
"total_frames": 14166,
"total_tasks": 1,
"total_videos": 204,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:51"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/1g0rrr/demo_insert_connector_from_rest_dag1_center.FR5_task2_transfer_the_gray_tablets_from_the_brown_bottle_into_another_brown_bottle_200This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "fairino_follower",
"total_episodes": 200,
"total_frames": 293120,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:200"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/coport-uni/FR5_task2_transfer_the_gray_tablets_from_the_brown_bottle_into_another_brown_bottle_200.EOT-2004-Raw
End of Term 2004
Original dump: https://eotarchive.org/data/data-2004/
The End of Term Web Archive is a crawl of U.S. government websites conducted at the end of each presidential administration. This is a filtered version of the 2004 crawl.
Notice
This dataset is still a work in progress.
Data Curation
We download the 2004 EOT WARC files and parse the HTML using Trafilatura. We then filter the extracted text by length (minimum of 550 characters)… See the full description on the dataset page: https://huggingface.co/datasets/fromziro/EOT-2004-Raw.agibot-sim-pack-moving-objects-from-conveyorThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "a2d",
"total_episodes": 100,
"total_frames": 90641,
"total_tasks": 1,
"total_videos": 300,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30.0,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/bot-pi/agibot-sim-pack-moving-objects-from-conveyor.summit_grab_wire_from_table2_229jetoncount_corpus
JetonCount's Corpus
This is the corpus used to train JetonCount.
JSONL Format
{
"source_file": "token_stats\\HuggingFaceFW_fineweb-edu00000.jsonl",
"dataset_dir": "HuggingFaceFW/fineweb-edu",
"index": 4020,
"chars": 2178,
"words": 335,
"avg_chars_per_word": 5.504478,
"longest_word_chars": 33,
"punctuation_ratio": 0.037649,
"symbol_ratio": 0.00551,
"tokens": 664,
"vocab_size": 2560,
"tokenizer_dir": "fromziro/Er-Tiny-1.3M"
}… See the full description on the dataset page: https://huggingface.co/datasets/fromziro/jetoncount_corpus.pour_from_cup_to_cup_lerobotv2.1demo_insert_connector_from_rest_dag3_centerThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "sam_evt2",
"total_episodes": 101,
"total_frames": 28225,
"total_tasks": 1,
"total_videos": 404,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:101"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/1g0rrr/demo_insert_connector_from_rest_dag3_center.ny_stereo_grab_wire_from_pile_51This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "sam_evt2",
"total_episodes": 51,
"total_frames": 32024,
"total_tasks": 1,
"total_videos": 204,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 60,
"splits": {
"train": "0:51"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/1g0rrr/ny_stereo_grab_wire_from_pile_51.demo_grab_pcb_from_edge_igor1_pcbcrop1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "sam_evt2",
"total_episodes": 51,
"total_frames": 22668,
"total_tasks": 1,
"total_videos": 204,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:51"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/1g0rrr/demo_grab_pcb_from_edge_igor1_pcbcrop1.summit_grab_pcb_from_edge_centerThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "sam_evt2",
"total_episodes": 52,
"total_frames": 19092,
"total_tasks": 1,
"total_videos": 208,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:52"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/1g0rrr/summit_grab_pcb_from_edge_center.agilex_take_out_glassescase_from_bagThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "arx5_bimanual",
"total_episodes": 20,
"total_frames": 14638,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 25,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/agilex_take_out_glassescase_from_bag.medicines_from_zakupki_gov_ruДанные для исследования существования focal points (https://www.jstor.org/stable/3132148) в гос. закупках лекарств в России.
summit_grab_wire_from_table2_dag1_upperThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "sam_evt2",
"total_episodes": 50,
"total_frames": 25107,
"total_tasks": 1,
"total_videos": 200,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/1g0rrr/summit_grab_wire_from_table2_dag1_upper.summarize_from_feedback_tldr_3_filtered_oai_preprocessing_1706381144
TL;DR SFT Dataset for OpenAI's Summarize from Feedback task
The dataset is directly taken from https://github.com/openai/summarize-from-feedback/tree/700967448d10004279f138666442bf1497d0e705#reddit-tldr-dataset
These columns are taken directly from the aforementioned dataset:
id: unique identifier for the post
subreddit: subreddit the post was taken from
title: title of the post
post: body of the post
summary: summary of the post
reference_response: reference response for the post… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/summarize_from_feedback_tldr_3_filtered_oai_preprocessing_1706381144.
