datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MegaMath-Web-Pro-Max
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
The Curation of MegaMath-Web-Pro-Max
Step 1: Uniformly and randomly sample millions of documents from the MegaMath-Web corpus, stratified by publication year;
Step 2: Annotate them using Llama-3.1-70B-instruct with a scoring prompt from FineMath and prepare the seed data;
Step 3: Training a fasttext carefully with proper preprocessing;
Step 4: Filtering documents with a threshold (i.e., 0.4);
Step 5:… See the full description on the dataset page: https://huggingface.co/datasets/OctoThinker/MegaMath-Web-Pro-Max.octflow-newdata-v1ecommerce-behavior-data-from-multi-category-store_oct-nov_2019
eCommerce Behavior Data from Multi-Category Store
About the Dataset
This dataset contains behavioral data for 285 million user events from a large multi-category eCommerce store. The data spans 7 months (October 2019 - April 2020) and records various user interactions with products.
Dataset Overview
Time Frame: October 2019 - April 2020
Total Events: 285 million
Event Granularity: Each row represents an event associated with a product and a user.
Data Source:… See the full description on the dataset page: https://huggingface.co/datasets/kevykibbz/ecommerce-behavior-data-from-multi-category-store_oct-nov_2019.Octant_CYP_inhibition_reactivity_blog_release
OpenADMET Octant CYP Inhibition & Reactivity
Data release from the OpenADMET consortium, generated by Octant Bio.
This dataset accompanies the blog post Building the OpenADMET Data Engine.
Source code, assay protocols, and raw TSV files are on GitHub.
Overview
Cytochrome P450 (CYP) enzymes drive the oxidative metabolism of most drugs and are a primary cause of drug-drug interactions (DDIs).
Despite their importance, public CYP datasets are sparse, noisy, and… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/Octant_CYP_inhibition_reactivity_blog_release.octo-small-simpler-cube-stack-rollout-bank-50
Octo-Small SIMPLER cube-stack rollout bank
This bank contains exactly 50 deterministic Octo-Small
rollouts for StackGreenCubeOnYellowCubeBakedTexInScene-v1: 3
successes and 47 failures. Every episode has a complete
61-frame H.264 video, a seven-frame contact sheet, losslessly stored per-step
telemetry, derived phase/geometry metrics, an external-LLM diagnosis, exact
evidence values, and a future action-patch hypothesis for failures.
Dataset:… See the full description on the dataset page: https://huggingface.co/datasets/lsnu/octo-small-simpler-cube-stack-rollout-bank-50.Octpickandplace_bboxes
Octpickandplace
This dataset was generated using phosphobot.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot.
To get started in robotics, get your own phospho starter pack..
octopus_dataThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 30,
"total_frames": 8954,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jakmilller/octopus_data.wikipedia_august_october_diff
Dataset Card for "wikipedia_august_october_diff"
More Information needed
cross-repo-code-samplehuggingface_downloads_October_2025
Description
Data used to build the plots of this blogpost.Contains all the models of the 50 most downloaded entities on HF as of October 1, 2025.
panda_pick_octo_resizedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 30,
"total_frames": 559,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:30"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lilkm/panda_pick_octo_resized.rope_cut_oct_xyzi_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.pointcloud": {
"dtype": "float32",
"shape": [
2048,
4
],
"names": [
"N_points",
"XYZI"
]
},
"observation.state": {
"dtype": "float32",
"shape": [… See the full description on the dataset page: https://huggingface.co/datasets/escapebirdy/rope_cut_oct_xyzi_v1.oct_23_1110amThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 3847,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dll-hackathon-102025/oct_23_1110am.DataSet_mix_duck_oct_cabecommerce-retail-product-matching-workflow-dataset
Ecommerce Retail Product Matching Workflow Dataset
This dataset is a public-facing sanitized workflow preview for managed ecommerce and retail product matching. It shows how candidate retrieval, UPC/model/brand/title/image evidence, customer-visible URL validation, confidence bands, review buckets, and rejection reasons can be structured for pricing intelligence, merchandising, data engineering, and AI-assisted product matching workflows.
Use this dataset to evaluate product… See the full description on the dataset page: https://huggingface.co/datasets/Octoparse/ecommerce-retail-product-matching-workflow-dataset.pick_cube_octoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 30,
"total_frames": 1789,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:30"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lilkm/pick_cube_octo.mags-oct-28octopus_table_MVHuman_20251113_ss_hgThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 20,
"total_frames": 5550,
"total_tasks": 1,
"total_videos": 140,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sriramsk/octopus_table_MVHuman_20251113_ss_hg.octoagent-benchmarks-results
E2E v3 results (Hub splits)
Layout baisbench_celltype (path prefix baisbench_celltype/)
Prepared: datasets.load_dataset(repo, 'baisbench_celltype_prepared', split="task_type_N")
Results: datasets.load_dataset(repo, 'baisbench_celltype', split="task_type_N")
Layout baisbench_codex_gpt55 (path prefix baisbench_codex_gpt55/)
Prepared: datasets.load_dataset(repo, 'baisbench_codex_gpt55_prepared', split="task_type_N")
Results: datasets.load_dataset(repo… See the full description on the dataset page: https://huggingface.co/datasets/ixprzemyslawpietrzak/octoagent-benchmarks-results.oct_25_umbrella_435pmThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 4447,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dll-hackathon-102025/oct_25_umbrella_435pm.so101_sim_wmThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 5000,
"total_frames": 881942,
"total_tasks": 6,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:5000"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/OctoEnSueur/so101_sim_wm.oct_23_1115amThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 4077,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dll-hackathon-102025/oct_23_1115am.oct_25_umbrella_420pmThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 4955,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dll-hackathon-102025/oct_25_umbrella_420pm.oct_25_umbrella_500pmThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 5481,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dll-hackathon-102025/oct_25_umbrella_500pm.oct_19_510pmThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 4891,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dll-hackathon-102025/oct_19_510pm.oct_25_1200pmThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 4112,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dll-hackathon-102025/oct_25_1200pm.octopus_table_MVHuman_20251210_ss_hgThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 20,
"total_frames": 4144,
"total_tasks": 1,
"total_videos": 140,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sriramsk/octopus_table_MVHuman_20251210_ss_hg.tiktok-brand-monitoring-beauty-sample
TikTok Beauty Brand Engagement Dataset
7 videos · 1,811 comments · 34 detected languages · 3.47M combined views
A structured, GDPR-compliant sample dataset of TikTok video metadata and user comments across beauty and skincare brand accounts. Built to accelerate research in social media sentiment analysis, brand engagement modeling, and multilingual NLP for the beauty vertical.
Produced by Octoparse Managed Data Service — enterprise web data pipelines for brand intelligence teams.… See the full description on the dataset page: https://huggingface.co/datasets/Octoparse/tiktok-brand-monitoring-beauty-sample.oct_25_1210amThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 3999,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dll-hackathon-102025/oct_25_1210am.octopus_table_human_multiview_20251210This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "aloha",
"total_episodes": 20,
"total_frames": 4144,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sriramsk/octopus_table_human_multiview_20251210.
