datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Diverse-SDXL-Dogs
Dataset Card for Diverse-SDXL-Dogs
This is a FiftyOne dataset with 181 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
import fiftyone.utils.huggingface as fouh
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = fouh.load_from_hub("Voxel51/Diverse-SDXL-Dogs")
# Launch the App
session = fo.launch_app(dataset)
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/Diverse-SDXL-Dogs.diverse_behaviour_activationsdiverse_qa_filteredaloha_pen_uncap_diverseThis dataset was created using LeRobot.
Dataset Description
This dataset is a lerobot conversion of the aloha_pen_uncap_diverse subset of BiPlay.
BiPlay contains 9.7 hours of bimanual data collected with an aloha robot at the RAIL lab @ UC Berkeley, USA. It contains 7023 clips, 2000 language annotations and 326 unique scenes.
Paper: https://huggingface.co/papers/2410.10088 Code: https://github.com/sudeepdasari/dit-policy If you use the dataset please cite:… See the full description on the dataset page: https://huggingface.co/datasets/physical-intelligence/aloha_pen_uncap_diverse.French-PD-diverse43,085,129,931 words
RoboTwin_pick_diverse_bottles_randomizedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aloha",
"total_episodes": 500,
"total_frames": 58824,
"total_tasks": 500,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:500"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/suz22/RoboTwin_pick_diverse_bottles_randomized.Nemotron-CC-Translated-Diverse-QA-itdiverse_risk_acts_fixedstratified-kmeans-diverse-pretraining-100K-1M
Stratified K-Means Diverse Pre-Training Dataset (100K-1M)
A carefully balanced subset combining FineWeb-Edu and Proof-Pile-2, featuring embedding-based k-means sampling to ensure diverse representation across educational and mathematical/scientific content at multiple scales.
👥 Follow the Authors
Aman Priyanshu
Supriti Vijay
Overview
This dataset provides stratified subsets at 50k, 100k, 250k, 500k, and 1M scales, combining high-quality… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/stratified-kmeans-diverse-pretraining-100K-1M.diversevul
Dataset Card for "diversevul"
Unofficial, not affiliated with the authors.
Paper: https://surrealyz.github.io/files/pubs/raid23-diversevul.pdf
Repository: https://github.com/wagner-group/diversevul
aloha_right_to_left_diverse1G1_Dex1_DiverseManip_DualArm_256x256This dataset was created using LeRobot.
Important Notes:
This is a G1 diversity dataset that can be used for video generation models, world models, and other applications [Lee et al., 2018].
If you want to use the lerobotv2.1 format, refer to this file for conversion: convert_v3_to_v2.py
Due to the inability to precisely describe spatial positions, adjust the scene to closely match the first frame of the dataset after installing the hardware as specified in… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_Dex1_DiverseManip_DualArm_256x256.French-PD-diverse43,085,129,931 words
diverse-robot-datasetG1_Dex1_DiverseManip_DualArm_128x128This dataset was created using LeRobot.
Important Notes:
This is a G1 diversity dataset that can be used for video generation models, world models, and other applications [Lee et al., 2018].
If you want to use the lerobotv2.1 format, refer to this file for conversion: convert_v3_to_v2.py
Due to the inability to precisely describe spatial positions, adjust the scene to closely match the first frame of the dataset after installing the hardware as specified in… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_Dex1_DiverseManip_DualArm_128x128.Repackage_diversechrono-2019-leak-sft-diversedirectional_pick_place_simpledesk_diverse_targets_tested40_200demos_20260716_molmobot
Directional SimpleDesk Pick-and-Place Diverse Targets
This dataset contains 200 successful scripted Franka demonstrations converted
from native MolmoSpaces output into the MolmoBot/Synthmanip training layout.
The task is to pick up one tabletop object and place it either to the left of
or to the right of a second object, from the robot's point of view.
The dataset is balanced by direction: 100 demonstrations use left prompts and
100 use right prompts. It covers 40 fixed initial… See the full description on the dataset page: https://huggingface.co/datasets/ccwatson/directional_pick_place_simpledesk_diverse_targets_tested40_200demos_20260716_molmobot.DiversePose300
DiversePose 300
DiversePose 300 is a challenging category-level object pose estimation benchmark introduced in CPPF++: Uncertainty-Aware Sim2Real Object Pose Estimation by Vote Aggregation (TPAMI 2024). It contains real-world RGB-D captures of common objects (bottles, bowls, mugs) in diverse, unconstrained scenes, designed to test the generalization of pose estimators beyond NOCS REAL275.
Code: https://github.com/qq456cvb/CPPF2
Project page:… See the full description on the dataset page: https://huggingface.co/datasets/qq456cvb/DiversePose300.G1_Dex1_DiverseManip_SingleArm_256x256This dataset was created using LeRobot.
Important Notes:
This is a G1 diversity dataset that can be used for video generation models, world models, and other applications [Lee et al., 2018].
If you want to use the lerobotv2.1 format, refer to this file for conversion: convert_v3_to_v2.py
Due to the inability to precisely describe spatial positions, adjust the scene to closely match the first frame of the dataset after installing the hardware as specified in… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_Dex1_DiverseManip_SingleArm_256x256.G1_Dex1_DiverseManip_SingleArm_128x128This dataset was created using LeRobot.
Important Notes:
This is a G1 diversity dataset that can be used for video generation models, world models, and other applications [Lee et al., 2018].
If you want to use the lerobotv2.1 format, refer to this file for conversion: convert_v3_to_v2.py
Due to the inability to precisely describe spatial positions, adjust the scene to closely match the first frame of the dataset after installing the hardware as specified in… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_Dex1_DiverseManip_SingleArm_128x128.chrono-2020-leak-sft-diversediverse_french_newsleaves-of-grass
leaves of grass
The following is a dataset for training a model to generate text in the style of Walt Whitman's "Leaves of Grass".
The idea with this dataset is to provide a single line (input) and then provide the next lines (1 to 10 lines) of the poem as output.
A model can then be trained to generate lines of poems given a single line of input.
There is a generate_data.py script that can be used to generate the dataset.
It keeps some formatting. New lines may be indented by a… See the full description on the dataset page: https://huggingface.co/datasets/diversen/leaves-of-grass.llm-ensembles-15-afg-m50-diverseRoboTwin_pick_diverse_bottlesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aloha",
"total_episodes": 50,
"total_frames": 5798,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/suz22/RoboTwin_pick_diverse_bottles.gsm8k-synthetic-diverse-8b
gretelai/gsm8k-synthetic-diverse-8b
This dataset is a synthetically generated version inspired by the GSM8K https://huggingface.co/datasets/openai/gsm8k dataset, created entirely using Gretel Navigator with meta-llama/Meta-Llama-3.1-8B as the agent LLM. It contains ~1500 Grade School-level math word problems with step-by-step solutions, focusing on age group, difficulty, and domain diversity.
Key Features:
Synthetically Generated: Math problems created using Gretel… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/gsm8k-synthetic-diverse-8b.so101_pick_diverse_objects
Dataset Card: SO101 Object Pick-Up Dataset
Overview
This dataset contains object pick-up demonstrations collected using the SO101 robotic arm. It is intended for training robot manipulation policies focused on pick-up tasks across a diverse set of everyday objects.
Dataset Summary
Item
Details
Release Date
April 30, 2026
Number of Objects
~70 objects
Total Duration
~2 hours
Task Type
Object Pick-Up
Robot
SO101
Robot Setup… See the full description on the dataset page: https://huggingface.co/datasets/TakuyaHiraoka/so101_pick_diverse_objects.mixed-sft-openai-tools-qwen3-areal-diverse
mixed-sft-openai-tools-qwen3-areal-diverse
Private SFT dataset for the Qwen3 AReaL terminal-agent trainer. It is the
normalized, shuffled "diverse" variant of the default terminal-agent SFT recipe.
It is intended for:
terminal_agent_demo/sft/config_terminus2_l40s_default_diverse.yaml
Files
File
Rows
Description
mixed_sft_openai_tools_qwen3_areal.shuf_seed7.jsonl
127,270
Training JSONL, shuffled with seed 7 (post-filtering).… See the full description on the dataset page: https://huggingface.co/datasets/eewer/mixed-sft-openai-tools-qwen3-areal-diverse.ai2thor-perspective-qa-400-balanced-diverse
