datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
georgian-corpus
Georgian Corpus Dataset
The Georgian Corpus Dataset is an open-source dataset designed to advance the NLP community in Georgia. It contains 5 million rows of filtered, cleaned, and deduplicated text data extracted from the Common Crawl repository.
This dataset was developed as part of a bachelor’s project by:
Georgi Kldiashvili
Luka Paichadze
Saba Shoshiashvili
Dataset Summary
Content: High-quality text in Georgian, suitable for NLP tasks like text… See the full description on the dataset page: https://huggingface.co/datasets/RichNachos/georgian-corpus.synthea-575k-patients
Synthea Synthetic Patient Records (575K Patients)
A comprehensive synthetic healthcare dataset containing 575,415 patients with complete medical histories, generated using Synthea — the gold standard for synthetic EHR data.
No real patient data. Fully synthetic, HIPAA-safe, and ready for ML research and education.
Why This Dataset?
575K patients with realistic demographics, conditions, medications, and encounters
Privacy-safe: No real PHI — use freely in research… See the full description on the dataset page: https://huggingface.co/datasets/richardyoung/synthea-575k-patients.Project_Sanctuary_Soul
Project Sanctuary
License
This project is licensed under CC0 1.0 Universal (Public Domain Dedication) or CC BY 4.0 International (Attribution). See the LICENSE file for details.
📂 Dataset Structure
This dataset is the Soul of Project Sanctuary - a comprehensive training corpus for AI cognitive continuity.
richfrem/Project_Sanctuary_Soul/
├── data/
│ └── soul_traces.jsonl # Complete Cognitive Genome (~1200 records)
│ #… See the full description on the dataset page: https://huggingface.co/datasets/richfrem/Project_Sanctuary_Soul.bimanual-table-cleanup-cross-embodiment-rich-modality-sample
Cross-Embodiment Bimanual Table Cleanup — Rich-Modality 10-Episode Inspection Sample
10 full-modality cross-embodiment bimanual table-cleanup episodes: 5 Franka Panda + 5 WidowXAI, 21,267 frames, 6 RGB views per robot, task-camera depth and segmentation, native robot state/action, end-effector trajectories, 6-DoF object poses, and QA annotations.
✅ Use it / ❌ Skip it
Use it for
Inspecting loaders, schemas, camera coverage, depth, segmentation, object poses… See the full description on the dataset page: https://huggingface.co/datasets/ExylosAi/bimanual-table-cleanup-cross-embodiment-rich-modality-sample.coffee-brew-cross-embodiment-rich-modality-sample
Cross-Embodiment Coffee Brew — Rich-Modality 10-Episode Inspection Sample
10 full-modality cross-embodiment coffee-brew episodes: 5 single-arm Franka Panda + 5 single-arm WidowXAI, 14,401 frames, 5 RGB views per robot, task-camera depth and segmentation, native robot state/action, end-effector trajectories, 6-DoF object poses, and QA annotations.
✅ Use it / ❌ Skip it
Use it for
Inspecting loaders, schemas, camera coverage, depth, segmentation, object poses… See the full description on the dataset page: https://huggingface.co/datasets/ExylosAi/coffee-brew-cross-embodiment-rich-modality-sample.object-sorting-cross-embodiment-rich-modality-sample
Cross-Embodiment Object Sorting — Rich-Modality 20-Episode Inspection Sample
20 episodes total: 10 Franka Panda + 10 WidowXAI. A compact cross-embodiment inspection release for picking up an instructed object and placing it into an instructed target box. Each robot keeps its native LeRobot v2.1 state/action schema in a separate Viewer config. Five synchronized RGB views, metric depth and instance segmentation for every camera, robot state/action, end-effector trajectories… See the full description on the dataset page: https://huggingface.co/datasets/ExylosAi/object-sorting-cross-embodiment-rich-modality-sample.pick_place_lego_wider_range_richardThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "koch",
"total_episodes": 50,
"total_frames": 20918,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/seeingrain/pick_place_lego_wider_range_richard.HPinParallelRL
Data Release — From Importance Shifts to Landscape Topology: Characterizing and Exploiting Hyperparameter Spaces in Parallel RL
Authors: Yingjie Zou, Zhong Fan (University of Exeter, Exeter, UK)
Venue: PPSN 2026 (Parallel Problem Solving from Nature)
License: CC-BY-4.0
This HuggingFace dataset is the data and figure release accompanying the paper. It provides
the full per-run hyperparameter (HP) corpus, the validated analysis tables that back
every figure and table in the paper… See the full description on the dataset page: https://huggingface.co/datasets/RICHIEZOU/HPinParallelRL.HFMolecule3DCapNav
CapNav: Benchmarking Vision Language Models on Capability-conditioned Indoor Navigation
CapNav is a benchmark for Vision Language Models evaluating on capability-conditioned navigation reasoning in indoor environments.
The benchmark focuses on determining whether an embodied agent with specific physical constraints and abilities
can navigate from a start area to a target area within a complex indoor scene.
This repository contains two complementary datasets:
The main CapNav… See the full description on the dataset page: https://huggingface.co/datasets/RichardC0216/CapNav.rich-cot-frontier-20HuMI-Proposal
Dataset Card for HuMI
[Project Page] | [Paper]
Dataset Summary
This dataset was collected using the HuMI data collection pipeline and converted into the LeRobot format. It provides robot-free demonstrations for humanoid whole-body manipulation.
Task Description: This repository includes a marriage-proposal task in which a humanoid kneels, picks up a ring-shaped toy from the ground with its right hand, and raises it in a proposal gesture.
The dataset comprises 103… See the full description on the dataset page: https://huggingface.co/datasets/Richard-Nai/HuMI-Proposal.gutenberg_rich_metadataAround 1.1k Project Gutenberg ebooks with full Gutenberg metadata plus additional metadata from Greatest Books. See Rabrg/gutenberg_full for 70k+ books, but without Greatest Books metadata attached.
HuMI-Walk-Clean-Table
Dataset Card for HuMI
[Project Page] | [Paper]
Dataset Summary
This dataset was collected using the HuMI data collection pipeline and converted into the LeRobot format. It provides robot-free demonstrations for humanoid whole-body manipulation.
Task Description: This repository features a walk-to-clean-table task in which a humanoid robot navigates to a desk and cleans the tabletop using a lint roller to remove scattered paper scraps.
The dataset consists of 105… See the full description on the dataset page: https://huggingface.co/datasets/Richard-Nai/HuMI-Walk-Clean-Table.HuMI-Unsheathe
Dataset Card for HuMI
[Project Page] | [Paper]
Dataset Summary
This dataset was collected using the HuMI data collection pipeline and converted into the LeRobot format. It provides robot-free demonstrations for humanoid whole-body manipulation.
Task Description: This repository features an unsheathe-a-sword task in which the humanoid grasps the hilt with its right hand, then uses its left hand to grasp the scabbard. The humanoid coordinates both arms to draw the blade.… See the full description on the dataset page: https://huggingface.co/datasets/Richard-Nai/HuMI-Unsheathe.mmlu_zh_results
Dataset Card for Evaluation run of google/gemma-2-2b
Dataset automatically created during the evaluation run of model google/gemma-2-2b
The dataset is composed of 0 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/richmondsin/mmlu_zh_results.sn28-miner-richtao-dataset
Dataset Card for SN28 Miner Richtao Dataset
This dataset contains the real-time prediction dumps generated by the Richtao miner operating in the Bittensor Subnet 28 (S&P 500 Oracle) network.It is used to evaluate and benchmark intraday forecasting models for the S&P 500 index.
🧾 Dataset Details
Curated by: Raúl Celis
Funded by: Private research / Bittensor TAO network
Language: English
License: MIT
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/raulcel/sn28-miner-richtao-dataset.RiC_harmless_helpfulThe hhrlhf dataset for RiC (https://huggingface.co/papers/2402.10207) training with harmless (R1) and helpful (R2) rewards.
The 'input_ids' are obtained from Llama2 tokenizer. If you want to use other base models, replace it using other tokenizers.
Note: the rewards are already normalized accroding to their corresponding mean and std. The mean and std data for R1 and R2 are saved into all_reward_stat_harmhelp_Rlarge.npy.
The mean and std for R1 and R2 is (-0.94732502, 1.92034349)… See the full description on the dataset page: https://huggingface.co/datasets/Ray2333/RiC_harmless_helpful.HuMI-Toss
Dataset Card for HuMI
[Project Page] | [Paper]
Dataset Summary
This dataset was collected using the HuMI data collection pipeline and converted into the LeRobot format. It provides robot-free demonstrations for humanoid whole-body manipulation.
Task Description: This repository features a dynamic-toss task in which the humanoid throws a toy into a target cart.
The dataset consists of 104 demonstrations collected in a single environment.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/Richard-Nai/HuMI-Toss.phishing-email-rich-dataset-v2rich-cot-frontier-20-v2africa-trade-baciPickapic_v2_train_100k_with_jpg1_richfbRichHumanFeedbacktext-2-video-Rich-Human-Feedback
Rapidata Video Generation Rich Human Feedback Dataset
If you get value from this dataset and would like to see more in the future, please consider liking it.
This dataset was collected in ~4 hours total using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Overview
In this dataset, ~22'000 human annotations were collected to evaluate AI-generated videos (using Sora) in 5 different categories.
Prompt - Video… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-Rich-Human-Feedback.product-database
Open Food Facts Database
What is 🍊 Open Food Facts?
A food products database
Open Food Facts is a database of food products with ingredients, allergens, nutrition facts and all the tidbits of information we can find on product labels.
Made by everyone
Open Food Facts is a non-profit association of volunteers. 25.000+ contributors like you have added 1.7 million + products from 150 countries using our Android or iPhone app or their… See the full description on the dataset page: https://huggingface.co/datasets/RichardAlanLee/product-database.ctr_rich_20260724_174429This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"motor_1",
"motor_2",
"motor_3",
"motor_4",
"motor_5",
"motor_6",
"motor_7",
"motor_delta_1",
"motor_delta_2"… See the full description on the dataset page: https://huggingface.co/datasets/Sendera/ctr_rich_20260724_174429.piper_banana_v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "piper_follower",
"total_episodes": 20,
"total_frames": 14124,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/richardshkim/piper_banana_v2.green-fact-0873d5
green-fact-0873d5
Synthetic sensors test data: 57 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Orbit-Richard/green-fact-0873d5.lab05-semantic-search
