datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
regionsOpenThoughts-1k-sample
[!NOTE]
We have released a paper for OpenThoughts! See our paper here.
Open-Thoughts-1k-sample
This is a 1k sample of the OpenThoughts-114k dataset.
Open synthetic reasoning dataset with high-quality examples covering math, science, code, and puzzles!
Inspect the content with rich formatting with Curator Viewer.
Available Subsets
default subset containing ready-to-train data used to finetune the OpenThinker-7B and OpenThinker-32B models:
ds =… See the full description on the dataset page: https://huggingface.co/datasets/ryanmarten/OpenThoughts-1k-sample.PhysicalAI-Robotics-GR00T-X-Embodiment-Sim
PhysicalAI-Robotics-GR00T-X-Embodiment-Sim
Github Repo: Isaac GR00T N1
We provide a set of datasets used for post-training of GR00T N1. Each dataset is a collection of trajectories from different robot embodiments and tasks.
Cross-embodied bimanual manipulation: 9k trajectories
Dataset Name
#trajectories
bimanual_panda_gripper.Threading
1000
bimanual_panda_hand.LiftTray
1000
bimanual_panda_gripper.ThreePieceAssembly
1000… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-X-Embodiment-Sim.resultsrodridrembpp
Dataset Card for Mostly Basic Python Problems (mbpp)
Dataset Summary
The benchmark consists of around 1,000 crowd-sourced Python programming problems, designed to be solvable by entry level programmers, covering programming fundamentals, standard library functionality, and so on. Each problem consists of a task description, code solution and 3 automated test cases. As described in the paper, a subset of the data has been hand-verified by us.
Released here as part of… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/mbpp.hellaswag
Dataset Card for "hellaswag"
Dataset Summary
HellaSwag: Can a Machine Really Finish Your Sentence? is a new dataset for commonsense NLI. A paper was published at ACL2019.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
default
Size of downloaded dataset files: 71.49 MB
Size of the generated dataset: 65.32 MB
Total… See the full description on the dataset page: https://huggingface.co/datasets/Rowan/hellaswag.roboreal_data10Kh-RealOmin-OpenData
Boasting over 13,000 hours of cumulative data and 5 million+ clips, it ranks as the largest open-source embodied intelligence dataset in the industry.
Update Notes:Stage 3 data upload completed.
13,000+ hours of pure dual-hand data with frame-level alignment latency < 1ms
Full high-precision trajectory reconstruction, breaking the limit of superficial open source, fully ready-to-use
3,000+ contributors and 10,000+ real household scenarios with exceptional diversity… See the full description on the dataset page: https://huggingface.co/datasets/genrobot2025/10Kh-RealOmin-OpenData.Rosetta-Activations
Rosetta Activations
Updated: 2026-06-15 02:30 UTC
Contrastive activation extractions for 17 semantic concepts across 46 language models,
supporting cross-architecture mechanistic interpretability research.
Companion concept pair corpus: jamesrahenry/Rosetta_Concept_Pairs
Papers: forthcoming
Dataset Structure
Rosetta-Activations/
├── rcp_v1/ # Current extraction line — richest data (N≈2000)
│ └── {Model_Name}/
│ ├── calibration_{concept}.npy… See the full description on the dataset page: https://huggingface.co/datasets/james-ra-henry/Rosetta-Activations.nonmyopia_resultsresultstransformers_circleci_workflow_runsevaluation-results@misc{muennighoff2022crosslingual,
title={Crosslingual Generalization through Multitask Finetuning},
author={Niklas Muennighoff and Thomas Wang and Lintang Sutawika and Adam Roberts and Stella Biderman and Teven Le Scao and M Saiful Bari and Sheng Shen and Zheng-Xin Yong and Hailey Schoelkopf and Xiangru Tang and Dragomir Radev and Alham Fikri Aji and Khalid Almubarak and Samuel Albanie and Zaid Alyafeai and Albert Webson and Edward Raff and Colin Raffel},
year={2022},
eprint={2211.01786},
archivePrefix={arXiv},
primaryClass={cs.CL}
}squad
Dataset Card for SQuAD
Dataset Summary
Stanford Question Answering Dataset (SQuAD) is a reading comprehension dataset, consisting of questions posed by crowdworkers on a set of Wikipedia articles, where the answer to every question is a segment of text, or span, from the corresponding reading passage, or the question might be unanswerable.
SQuAD 1.1 contains 100,000+ question-answer pairs on 500+ articles.
Supported Tasks and Leaderboards
Question… See the full description on the dataset page: https://huggingface.co/datasets/rajpurkar/squad.paws
Dataset Card for PAWS: Paraphrase Adversaries from Word Scrambling
Dataset Summary
PAWS: Paraphrase Adversaries from Word Scrambling
This dataset contains 108,463 human-labeled and 656k noisily labeled pairs that feature the importance of modeling structure, context, and word order information for the problem of paraphrase identification. The dataset has two subsets, one based on Wikipedia and the other one based on the Quora Question Pairs (QQP) dataset.
For further… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/paws.rdpRekaDaily-10k-raw
RekaDaily-10k (raw)
Raw, unscripted, first-person daily-life video, collected through
Claru, Reka's data collection marketplace — recorded by
paid collectors in their own homes and workplaces on head-mounted and handheld
phones, across multiple regions.
Videos are delivered as recorded — no cuts, no trimming, no editing, no
filtering beyond basic integrity checks. A processed tier (short clips with
machine captions) is released separately under the same RekaDaily-10k prefix.… See the full description on the dataset page: https://huggingface.co/datasets/RekaAI/RekaDaily-10k-raw.race
Dataset Card for "race"
Dataset Summary
RACE is a large-scale reading comprehension dataset with more than 28,000 passages and nearly 100,000 questions. The
dataset is collected from English examinations in China, which are designed for middle school and high school students.
The dataset can be served as the training and test sets for machine comprehension.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed… See the full description on the dataset page: https://huggingface.co/datasets/ehovy/race.OpenR1-Math-220k
OpenR1-Math-220k
Dataset description
OpenR1-Math-220k is a large-scale dataset for mathematical reasoning. It consists of 220k math problems with two to four reasoning traces generated by DeepSeek R1 for problems from NuminaMath 1.5.
The traces were verified using Math Verify for most samples and Llama-3.3-70B-Instruct as a judge for 12% of the samples, and each problem contains at least one reasoning trace with a correct answer.
The dataset consists of two… See the full description on the dataset page: https://huggingface.co/datasets/open-r1/OpenR1-Math-220k.dbsueraRoboDojoPhysicalAI-Robotics-Open-H-Embodiment
Dataset Description:
Open-H-Embodiment is a community‑driven dataset initiative building the open, shared foundation needed to train and evaluate AI autonomy models for surgical robotics and ultrasound.
This dataset is a multi-embodiment collection of LeRobot datasets of paired kinematics and video, across tasks such as tabletop exercises, clinical procedures, as well as simulations of healthcare robotics applications.
Maintainer / Hosting Organization:
NVIDIA… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-Open-H-Embodiment.ReActor
ReActor Assets
The Fast and Simple Face Swap Extension
ComfyUI-ReActor (ex. comfyui-reactor-node)
sd-webui-reactor
Models
file
source
license
buffalo_l.zip
DeepInsight
codeformer-v0.1.0.pth
sczhou
GFPGANv1.3.pth
TencentARC
GFPGANv1.4.pth
TencentARC
GPEN-BFR-512.onnx
harisreedhar
RestoreFormer_PP.onnx
netrunner.exe
inswapper_128.onnx
DeepInsight
inswapper_128_fp16.onnx
Hillobar
RT-PosePaper
RT-Pose: A 4D Radar Tensor-based 3D Human Pose Estimation and Localization Benchmark (ECCV 2024)
RT-Pose introduces a human pose estimation (HPE) dataset and benchmark by integrating a unique combination of calibrated radar ADC data, 4D radar tensors, stereo RGB images, and LiDAR point clouds.
This integration marks a significant advancement in studying human pose analysis through multi-modality datasets.
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/uwipl/RT-Pose.CT-RATE
The CT-RATE Team organizes the VLM3D Challenge
VLM3D 2026 (2nd Edition) → Challenge Finals at MICCAI 2026
VLM3D 2025 (1st Edition) → Challenge Finals at MICCAI 2025 • Workshop at ICCV 2025
The CT-RATE Team is developing the MR-RATE Dataset
A large-scale brain MRI dataset with paired radiology reports for training 3D vision-language models.
GitHub |
Dataset |
Metadata Dashboard
Generalist Foundation Models from a Multimodal Dataset for 3D Computed Tomography… See the full description on the dataset page: https://huggingface.co/datasets/ibrahimhamamci/CT-RATE.RealCam-Vid
RealCam-Vid Dataset
News
25/04/08: We provide torch dataset demo code for example usage of our RealCam-Vid.
25/03/26: Release our dataset RealCam-Vid v1 for metric-scale camera-controlled video generation, containing ~100K video clips with dedicated short/long captions and metric-scale camera annotations.
25/02/18: Initial commit of the project, we plan to release the full dataset and data processing code in several… See the full description on the dataset page: https://huggingface.co/datasets/MuteApo/RealCam-Vid.Retargeted_AMASS_for_bxi_elf2
Retargeted AMASS for Robotics
Project Overview
This project aims to retarget motion data from the AMASS dataset to various robot models and open-source the retargeted data to facilitate research and applications in robotics and human-robot interaction. AMASS (Archive of Motion Capture as Surface Shapes) is a high-quality human motion capture dataset, and the SMPL-X model is a powerful tool for generating realistic human motion data.
By adapting the motion data from AMASS… See the full description on the dataset page: https://huggingface.co/datasets/fleaven/Retargeted_AMASS_for_bxi_elf2.riddle_senserequests
