datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
M12-fastener_removeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_widowxai_follower_robot",
"total_episodes": 60,
"total_frames": 53873,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:60"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/REBOOT26/M12-fastener_remove.M12-fastener_installThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_widowxai_follower_robot",
"total_episodes": 60,
"total_frames": 53864,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:60"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/REBOOT26/M12-fastener_install.assemble_battery_longSuzuki-m1
Suzuki-m1 — Corpus de pré-treinamento em português
Corpus de texto em português brasileiro/europeu para pré-treinamento (do zero) de modelos de linguagem causais (estilo GPT/Claude). Formato JSONL: um documento por linha, campo text.
Formato
{"text": "A economia brasileira cresceu no último trimestre, impulsionada pelo agronegócio."}
Fontes
Fonte
Descrição
wikimedia/wikipedia (20231101.pt)
Artigos da Wikipédia em português
allenai/c4… See the full description on the dataset page: https://huggingface.co/datasets/Raivatv24/Suzuki-m1.InternData-M1
InternData-M1
InternData-M1 is a comprehensive embodied robotics dataset containing 244K simulation demonstrations with rich frame-based information including 2D/3D boxes, trajectories, grasp points, and semantic masks, with comprehensive annotations.
Your browser does not support the video tag.
Changelog 📋
Previous versions remain available in the branch version name.
v0.1 :
Initial version(26-07-2025)
Added simulated/agilex… See the full description on the dataset page: https://huggingface.co/datasets/InternRobotics/InternData-M1.the-stack-v2-dedup-filtered-500-stars-100-forks-contentsMem-0-m1mix-dataset-RMBench
Mem-0 m1_mix — RMBench / RoboTwin 2.0 (LeRobot dataset)
The m1_mix training dataset for the Mem-0 execution module: the five RMBench
M1 tasks merged into a single LeRobot
v2.1 dataset with globally unique episode indices. This is the exact data used to
train the checkpoint released at
qiuly/Mem-0-m1mix-RMBench.
Summary
Format
LeRobot v2.1
Episodes
250 (50 per task × 5 tasks)
Frames
92,520
FPS
30
Tasks
5 (see below)
Robot
dual-arm (2× 7-DoF +… See the full description on the dataset page: https://huggingface.co/datasets/qiuly/Mem-0-m1mix-dataset-RMBench.stackexchange_filteredm12_recovery_installThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_widowxai_follower_robot",
"total_episodes": 60,
"total_frames": 53887,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:60"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/REBOOT26/m12_recovery_install.automl_llm_agent_m1
AutoML-LLM Agent Module 1 Benchmark
This dataset contains the Module 1 benchmark for evaluating an AutoML assistant that interprets user requests, selects tabular modeling settings, produces an auditable AutoGluon Tabular plan, and synthesizes the compact Module 1 recipe consumed by Module 2 through the mandatory final LLM writer used by all A-E variants.
The repository is scoped to Module 1 only.
Tables
cases: one row per Module 1 evaluation case.
queries: one… See the full description on the dataset page: https://huggingface.co/datasets/tecnologiactc/automl_llm_agent_m1.koch_sim_pick_cube_matrixThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "koch",
"total_episodes": 100,
"total_frames": 72130,
"total_tasks": 1,
"total_videos": 200,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/m1b/koch_sim_pick_cube_matrix.nouns
Dataset Card for Nouns auto-captioned
Dataset used to train Nouns text to image model
Automatically generated captions for Nouns from their attributes, colors and items. Help on the captioning script appreciated!
For each row the dataset contains image and text keys. image is a varying size PIL jpeg, and text is the accompanying text caption. Only a train split is provided.
Citation
If you use this dataset, please cite it as:
@misc{piedrafita2022nouns,
author =… See the full description on the dataset page: https://huggingface.co/datasets/m1guelpf/nouns.MiroMind-M1-SFT-719K
MiroMind-M1
🧾 Overview
Training performance of MiroMind-M1-RL-7B on AIME24 and AIME25.
MiroMind-M1 is a fully open-source series of reasoning language models built on Qwen-2.5, focused on advancing mathematical reasoning. It is trained through supervised fine-tuning (SFT) on 719K curated problems and reinforcement learning with verifiable rewards (RLVR) on 62K challenging examples, using a context-aware multi-stage policy optimization method… See the full description on the dataset page: https://huggingface.co/datasets/miromind-ai/MiroMind-M1-SFT-719K.dynasoftgrasp-m1-softmouse-pick
DynaSoftGrasp M1 · soft-body mouse pick episodes
200 episodes (190 judged SUCCESS, see success_manifest.json)
of a Franka Panda picking a FEM soft-body computer-mouse-sized mouse
(ByteDance Seed3D-generated mesh, 10 kPa tissue-scale Young's modulus baked
into the USD) on a flat table. Collected with Isaac Sim 4.5 / Isaac Lab 2.2.1
at 25 Hz, 480x360, 3 cameras, on top of the DynamicVLA collection pipeline.
Task prompt (stored per episode in language): "grasp a static soft mouse"… See the full description on the dataset page: https://huggingface.co/datasets/KhalilGao/dynasoftgrasp-m1-softmouse-pick.m1_muscarinic_receptor_antagonists_butkiewicz
Dataset Details
Dataset Description
Primary screen AID628 confirmed by screen AID677.
AID859 confirmed activity on rat M1 receptor.
The counter screen AID860 removed non-selective compounds
being active also at the rat M4 receptor.
Final set of active compoundsobtained by subtracting active compounds of AID860
from those in AID677, resulting in 448 total active compounds.
Curated by:
License: CC BY 4.0
Dataset Sources
original dataset
corresponding… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/m1_muscarinic_receptor_antagonists_butkiewicz.SBm12_recovery_removeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_widowxai_follower_robot",
"total_episodes": 60,
"total_frames": 53873,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:60"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/REBOOT26/m12_recovery_remove.feval-rolloutstask1_mode2_obst_SP10.5X7.5_item_B18X12_S10X13_dist_M19X8_R18X4_CO3X12_ST2X3_CA3X4_epi_10_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 10,
"total_frames": 2141,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cbrian/task1_mode2_obst_SP10.5X7.5_item_B18X12_S10X13_dist_M19X8_R18X4_CO3X12_ST2X3_CA3X4_epi_10_2.so100_bluelego_updtThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 41,
"total_frames": 9840,
"total_tasks": 1,
"total_videos": 82,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:41"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/m1b/so100_bluelego_updt.real01b-marker-d2-final-dp-vs-iql-m1c125k-m2c150k-n16-heval-sobolseed2026070102-n50This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
7
],
"names": [
"cart_pos_x",
"cart_pos_y",
"cart_pos_z",
"cart_rot_x",
"cart_rot_y",
"cart_rot_z"… See the full description on the dataset page: https://huggingface.co/datasets/ankile/real01b-marker-d2-final-dp-vs-iql-m1c125k-m2c150k-n16-heval-sobolseed2026070102-n50.m1_muscarinic_receptor_agonists_butkiewicz
Dataset Details
Dataset Description
Positive allosteric modulation of the M1 Muscarinic
receptor screened with AID626. Confirmed by screen AID 1488. A second
counter screen AID 1741. The final set of selective positive
allosteric modulators of M1 was obtained by removing compounds active
in AID 1741 from the compounds active in AID 1488 resulting in 188
compounds.
Curated by:
License: CC BY 4.0
Dataset Sources
original dataset
corresponding… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/m1_muscarinic_receptor_agonists_butkiewicz.M1_EURUSD_candles
All chunks have more than 4000 rows of data in chronological order in a panda dataframe
CSV files are the same data in chronological order, some may not be more than 4000 rows
MiroMind-M1-RL-62K
MiroMind-M1
🧾 Overview
Training performance of MiroMind-M1-RL-7B on AIME24 and AIME25.
MiroMind-M1 is a fully open-source series of reasoning language models built on Qwen-2.5, focused on advancing mathematical reasoning. It is trained through supervised fine-tuning (SFT) on 719K curated problems and reinforcement learning with verifiable rewards (RLVR) on 62K challenging examples, using a context-aware multi-stage policy optimization method… See the full description on the dataset page: https://huggingface.co/datasets/miromind-ai/MiroMind-M1-RL-62K.pack_3_objects_plusm1llion-lang
🌍 M1llion-Lang: Multilingual Instruction Dataset with Emoji Expression
M1llion-Lang is a high-quality, large-scale multilingual instruction dataset designed for training and fine-tuning large language models (LLMs) to understand and generate text in 20+ languages with natural emoji expression and cultural nuance.
📊 Dataset Overview
Total Size: ~4GB (JSON Lines format)
Languages: 20 languages covering 95%+ of global internet usersFormat: Conversational JSONL with… See the full description on the dataset page: https://huggingface.co/datasets/m1llion-ai-high-end-group/m1llion-lang.tool-wrench-m12This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"joint_0.pos",
"joint_1.pos",
"joint_2.pos",
"joint_3.pos",
"joint_4.pos",
"joint_5.pos",
"left_carriage_joint.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/rbtrprjkt/tool-wrench-m12.battery_assembledetails_LeroyDyer__Mixtral_AI_Cyber_3.m1
Dataset Card for Evaluation run of LeroyDyer/Mixtral_AI_Cyber_3.m1
Dataset automatically created during the evaluation run of model LeroyDyer/Mixtral_AI_Cyber_3.m1 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_LeroyDyer__Mixtral_AI_Cyber_3.m1.Gowalla_m1
Gowalla_m1
Dataset description:
The dataset statistics are summarized as follows:
Dataset ID
#Users
#Items
#Interactions
#Train
#Test
Density
Gowalla_m1
29,858
40,981
1,027,370
810,128
217,242
0.00084
Source: https://snap.stanford.edu/data/loc-gowalla.html
Download: https://huggingface.co/datasets/reczoo/Gowalla_m1/tree/main
RecZoo Datasets: https://github.com/reczoo/Datasets
Used by papers:
Xiangnan He, Kuan Deng, Xiang Wang, Yan Li, Yongdong Zhang, Meng… See the full description on the dataset page: https://huggingface.co/datasets/reczoo/Gowalla_m1.
