datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MegaPairs-Standard
MegaPairs-Standard (Standardized Version)
Dataset Summary
This is a standardized, high-efficiency version of the JUNJIE99/MegaPairs dataset.
Why use this version?
The original dataset is distributed as a massive Tar archive containing millions of images, accompanied by a separate JSONL annotation file.
The Problem: Using the original format requires extracting terabytes of small files (which can exhaust disk inodes) or writing complex logic to read from archives. It… See the full description on the dataset page: https://huggingface.co/datasets/86Cao/MegaPairs-Standard.Standard-Pipeline
Standard Pipeline
Environment adaptation, training evidence, latency profiles, raw demonstrations and evaluation traces, organized by environment and experiment stage.
AirRaid: zero-latency and profile-latency experiments.
Only raw demonstrations are distributed. Generate converted training datasets on the training server. Model weights remain in the dedicated model repository and are referenced from the experiment records. Pin a commit revision for reproducible downloads.… See the full description on the dataset page: https://huggingface.co/datasets/latency-sensitive-bench/Standard-Pipeline.hongyan_lift_bottle_standardThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 100,
"total_frames": 37265,
"total_tasks": 1,
"total_videos": 200,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/gaozj/hongyan_lift_bottle_standard.We-Math2.0-Standard
Dataset Card for We-Math 2.0
GitHub | Paper | Website
We-Math 2.0 is a unified system designed to comprehensively enhance the mathematical reasoning capabilities of Multimodal Large Language Models (MLLMs).
It integrates a structured mathematical knowledge system, model-centric data space modeling, and a reinforcement learning (RL)-based training paradigm to achieve both broad conceptual coverage and robust reasoning performance across varying difficulty levels.
The key… See the full description on the dataset page: https://huggingface.co/datasets/We-Math/We-Math2.0-Standard.hongyan_lift_bottle_standard_3_fixedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 100,
"total_frames": 40654,
"total_tasks": 1,
"total_videos": 200,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/gaozj/hongyan_lift_bottle_standard_3_fixed.simready-core
license: cc-by-4.0
hongyan_lift_bottle_standard_3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 100,
"total_frames": 40654,
"total_tasks": 1,
"total_videos": 200,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/gaozj/hongyan_lift_bottle_standard_3.ground-truth-mmmu-pro-standard-10MIRAGE-Standard-single-turn-with-meteo-pathsdownsampled_cleaned_chartQa_plotQa_colored_standardhongyan_lift_bottle_standard_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 100,
"total_frames": 28565,
"total_tasks": 1,
"total_videos": 200,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/gaozj/hongyan_lift_bottle_standard_2.PubmedOCR_standardFendi.Standard.Categories.Italy
Fendi web scraped data
About the website
In the EMEA region, particularly in Italy, the luxury fashion industry has an immense influence and it significantly contributes to Italys economy. Brands such as Fendi are prominent players in this sector. Emphasizing on haute couture, ready-to-wear clothing, leather goods, shoes, fragrances, eyewear, timepieces and accessories, the industry has seen significant growth with the adoption of Ecommerce. Our dataset provides an… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Fendi.Standard.Categories.Italy.HWD_Standard_Dataset
Dataset Card for "HWD_Standard_Dataset"
More Information needed
hongyan_lift_bottle_standard_4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 100,
"total_frames": 22278,
"total_tasks": 1,
"total_videos": 200,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/gaozj/hongyan_lift_bottle_standard_4.Train-image-generation-standardhongyan_move_cube_standardThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 10,
"total_frames": 16116,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/gaozj/hongyan_move_cube_standard.logu_lift_bottle_standard_2_fixedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 100,
"total_frames": 28565,
"total_tasks": 1,
"total_videos": 200,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/gaozj/logu_lift_bottle_standard_2_fixed.visdrone-det-standardizedPCBA_Standard-to-Real_ChallengeHere is a more concise, English version of the Hugging Face dataset card. You can copy and paste this directly into your README.md.
Dataset Card for PCBA Standard-to-Real Challenge
🏆 Challenge Overview
This is the official dataset for the PCBA Standard-to-Real Grand Challenge, held in conjunction with ACM Multimedia (MM) 2026.
This challenge focuses on Cross-domain Visual Question Answering (VQA) for real-world manufacturing inspection. The goal is to develop… See the full description on the dataset page: https://huggingface.co/datasets/aimmifm/PCBA_Standard-to-Real_Challenge.seqlab-pooling-benchmark-reportsimagenet-1k-standardizedSaint.Laurent.Standard.Categories.Italy
Saint Laurent web scraped data
About the website
Saint Laurent operates within the vibrant luxury fashion industry in the EMEA region, particularly in Italy. The Italian luxury fashion sector is characterised by its rich history, renowned craftsmanship and a high demand for its prestigious brands. Saint Laurent falls under this prestigious category. The industry has significant presence online, propelled by the rise of e-commerce and digital marketing strategies. The… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Saint.Laurent.Standard.Categories.Italy.ground-truth-mmmu-pro-standard-10-sampling-500robotv1-standardReal_ESR-GAN_Standardizedimnet1k_standard_poodleimnet1k_standard_schnauzerstandard1StandardDifusion
