datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
document-review-data
Document Review Data
Private dataset for the Office/PDF title extraction review app and the current extractive title-training data package.
Current Title Extraction Dataset Surface
Canonical prefix:
datasets/title_extraction/
Effective datasets:
datasets/title_extraction/training/source4k_device_qwen_fp1000_v1/
datasets/title_extraction/evaluation/real_device_280_v1/
datasets/title_extraction/synthetic/controlled_synthetic_parse_v1/
The Dataset Viewer is… See the full description on the dataset page: https://huggingface.co/datasets/mannycooper/document-review-data.appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-tmp01-reeval1
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-tmp01-reeval1
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.40546875
Action score: 0.475
Valid samples: 320/320
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-reeval1
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-reeval1
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.4046875
Action score: 0.4703125
Valid samples: 320/320
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.39921875
Action score: 0.44375
Valid samples: 320/320
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-t01
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-t01
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.38359375
Action score: 0.4703125
Valid samples: 320/320
kitchen-workspace-understanding-safe-manipulation
Kitchen Workspace Understanding & Safe Manipulation
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled by… See the full description on the dataset page: https://huggingface.co/datasets/physicl/kitchen-workspace-understanding-safe-manipulation.mmu_manga
mmu_manga HATS Catalog Collection
This is the collection of HATS catalogs representing mmu_manga.
This dataset is part of the Multimodal Universe,
a large-scale collection of multimodal astronomical data. For full details, see the paper:
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data.
Access the catalog
We recommend the use of the LSDB Python framework to access HATS catalogs.
LSDB can be installed via… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/mmu_manga.Mana-TTS
ManaTTS-Persian-Speech-Dataset
ManaTTS is the largest publicly available single-speaker Persian corpus, comprising over 114 hours of high-quality audio (sampled at 44.1 kHz). Released under the permissive CC-0 license, this dataset is freely usable for both educational and commercial purposes.
Collected from Nasl-e-Mana magazine, the dataset covers a diverse range of topics, making it ideal for training robust text-to-speech (TTS) models. The release includes a fully transparent… See the full description on the dataset page: https://huggingface.co/datasets/MahtaFetrat/Mana-TTS.PhysicalAI-Robotics-Manipulation-Kitchen
PhysicalAI Robotics Manipulation in the Kitchen
Dataset Description:
PhysicalAI-Robotics-Manipulation-Kitchen is a dataset of automatic generated motions of robots performing operations such as opening and closing cabinets, drawers, dishwashers and fridges. The dataset was generated in IsaacSim leveraging reasoning algorithms and optimization-based motion planning to find solutions to the tasks automatically [1, 3]. The dataset includes a bimanual manipulator built with… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-Manipulation-Kitchen.corpus-1T-manifest
SPP Corpus 1T Manifest
The selection manifest for the ~1.0T-token pretraining corpus used in
Synthetic Persona Pretraining (SPP): Alignment from Token Zero.
The corpus is a seeded subsample of allenai/dolma3_mix-6T.
Rather than redistribute ~2.6 TB of text that is already public, this dataset
publishes the selection decisions keyed by upstream document id, so the corpus
can be reconstructed exactly by replaying against upstream.
📄 Reflections + text for the annotated half:… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/corpus-1T-manifest.SpatialLM-Dataset
SpatialLM Dataset
The SpatialLM dataset is a large-scale, high-quality synthetic dataset designed by professional 3D designers and used for real-world production. It contains point clouds from 12,328 diverse indoor scenes comprising 54,778 rooms, each paired with rich ground-truth 3D annotations. SpatialLM dataset provides an additional valuable resource for advancing research in indoor scene understanding, 3D perception, and… See the full description on the dataset page: https://huggingface.co/datasets/manycore-research/SpatialLM-Dataset.PulseLM
PulseLM: A Foundation Dataset and Benchmark for PPG-Text Learning
Usage
from datasets import load_dataset, get_dataset_config_names, concatenate_datasets # datasets==4.5.0
dataset_names = get_dataset_config_names("Manhph2211/PulseLM")
print(f"Available datasets: {dataset_names}")
train_splits = [
load_dataset("Manhph2211/PulseLM", name, split="train").select_columns(["signal", "text", "qa"])
for name in dataset_names
]
combined =… See the full description on the dataset page: https://huggingface.co/datasets/Manhph2211/PulseLM.MANGO
MANGO: A Corpus of Human Ratings for Speech
MANGO (MUSHRA Assessment corpus using Native listeners and Guidelines to understand human Opinions at scale) is the first large-scale dataset designed for evaluating Text-to-Speech (TTS) systems in Indian languages.
Key Features:
255,150 human ratings of TTS-generated outputs and ground-truth human speech.
Covers two major Indian languages: Hindi & Tamil, and English.
Based on the MUSHRA (Multiple Stimuli with Hidden Reference… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/MANGO.everyday-manipulation-3d-raw
Everyday Manipulation 3D (raw RGB-D)
1,513 clips · 10.28 hours · 279 GiB · 4 participants · 10 manipulation tasks · 42 recording sittings
Chest-mounted iPhone Pro capture of everyday two-handed manipulation by
CaryX AI. Clips were recorded with
Record3D, an iOS app that captures the
iPhone's LiDAR RGB-D stream. Each clip is the app's .r3d recording with the
audio track removed; the sensor streams are unmodified: synchronised RGB,
metric LiDAR depth, per-frame ARKit 6-DoF camera… See the full description on the dataset page: https://huggingface.co/datasets/CaryxAI/everyday-manipulation-3d-raw.ego-tactile-manipulation
Ego-Tactile Manipulation
Egocentric video + dense two-hand tactile + touch-grounded action labels - by OpenGraph Labs.
Four episodes visualized in our dashboard - egocentric video with live tactile & sensor signals.
Synchronized ego + touch human-manipulation data is rare. This is a clean 1.28-hour sample from OpenGraph Labs' Physical-AI data pipeline: a head camera plus our OGLO tactile gloves on both hands, with action labels derived from the physical contact signal.… See the full description on the dataset page: https://huggingface.co/datasets/OpenGraphLabs-Research/ego-tactile-manipulation.Arena-G1-Loco-Manipulation-Task
Dataset Description:
The Arena-G1-Loco-Manipulation-Task dataset is multimodal collections of trajectories generated in Isaac Lab. It supports humanoid (G1) loco-manipulation task in IsaacLab-Arena environment. Each entry provides the full context (state, vision, language, action) needed to train and evaluate generalist robot policies for box pick and place task.
Dataset Name
# Trajectories
G1 Loco-Manipulation Task
50
This dataset is ideal for behavior cloning… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Arena-G1-Loco-Manipulation-Task.berkeley_fanuc_manipulation_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "fanuc_mate",
"total_episodes": 415,
"total_frames": 62613,
"total_tasks": 32,
"total_videos": 830,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:415"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/berkeley_fanuc_manipulation_lerobot.reranker-scoresmaniskill3-sft-rgb-lerobot-1200epThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "maniskill3_sim",
"total_episodes": 1200,
"total_frames": 171480,
"total_tasks": 6,
"total_videos": 0,
"total_chunks": 2,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:1200"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features":… See the full description on the dataset page: https://huggingface.co/datasets/onnoboru/maniskill3-sft-rgb-lerobot-1200ep.ATLAS-WDS-v3
ATLAS-WDS-v3
v2 的 has_wind=True 子集(其余字段与 v2 一致)。每条样本含 MEM/Fourier 方向谱、真实 UTC、同站风场、逐点波龄判据。
chimera-bench
CHIMERA-Bench v1.0
A unified benchmark for epitope-specific antibody CDR sequence-structure co-design.
Paper: CHIMERA-Bench: A Benchmark Dataset for Epitope-Specific Antibody Design (ICLR 2026 GEM Workshop)
Code: github.com/mansoorbaloch/chimera-bench
Dataset Summary
Property
Value
Complexes
2,922
PDB structures
2,721
Pre-computed features
2,941 .pt files
Splits
3 (epitope-group, antigen-fold, temporal)
Numbering schemesIMGT, Chothia
Contact… See the full description on the dataset page: https://huggingface.co/datasets/mansoorbaloch/chimera-bench.PhysicalAI-Robotics-Manipulation-ObjectsPhysicalAI-Robotics-Manipulation-Objects is a dataset of automatic generated motions of robots performing operations such as picking and placing objects in a kitchen environment. The dataset was generated in IsaacSim leveraging reasoning algorithms and optimization-based motion planning to find solutions to the tasks automatically [1, 3]. The dataset includes a bimanual manipulator built with Kinova Gen3 arms. The environments are kitchen scenes where the furniture and appliances were… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-Manipulation-Objects.Additive-Manufacturing-Benchmark
Additive Manufacturing Benchmark
A benchmark dataset for evaluating knowledge of additive manufacturing (AM) processes, derived from graduate-level coursework at Carnegie Mellon University.
Configurations
general_knowledge_multiple_choice
Multiple-choice questions covering various AM processes with explanations.
Column
Description
source
Source homework assignment (e.g. cmu_24_633_2023/homework_1_exone)
process
AM process type (e.g. Binder Jet… See the full description on the dataset page: https://huggingface.co/datasets/ppak10/Additive-Manufacturing-Benchmark.utokyo_pr2_tabletop_manipulationThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 240,
"total_frames": 32708,
"total_tasks": 3,
"total_videos": 240,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 5,
"splits": {
"train": "0:240"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/utokyo_pr2_tabletop_manipulation.Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2
🧬 Omni-Frontier Collection
Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package
A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible.
📖 Jump to
What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.maniskill_50ep_so101_blue_cube_orange_tray_20260812_131142This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/makermods/maniskill_50ep_so101_blue_cube_orange_tray_20260812_131142.maniskill_50ep_so101_blue_cube_orange_tray_20260811_114951This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/makermods/maniskill_50ep_so101_blue_cube_orange_tray_20260811_114951.AIRBOT_MMK2_store_pomegranates_and_mangoes
AIRBOT_MMK2_store_pomegranates_and_mangoes
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: discover_robotics_aitbot_mmk2
| Codebase Version: v2.1
End-Effector Type: five_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_store_pomegranates_and_mangoes.netryx-manhattan-15kmManiSkill_TwoRobotStackCube-v1_recovery
TwoRobotStackCube-v1 Dataset
This dataset was converted from ManiSkill format to LeRobot format. It is part of the paper CroSTAta: Cross-State Transition Attention Transformer for Robotic Manipulation.
Code: GitHub Repository
Authors: Giovanni Minelli, Giulio Turrisi, Victor Barasuol, Claudio Semini
Dataset Information
Environment: TwoRobotStackCube-v1
Total Episodes: 1043
Total Frames: 25763
FPS: 20
Video Keys: ['observation.image']
Target Control Mode: Original… See the full description on the dataset page: https://huggingface.co/datasets/johnMinelli/ManiSkill_TwoRobotStackCube-v1_recovery.
