datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arena-human-preference-55kDataset for Kaggle competition on predicting human preference on Chatbot Arena battles.
The training dataset includes over 55,000 real-world user and LLM conversations and user preferences across over 70 state-of-the-art LLMs, such as GPT-4, Claude 2, Llama 2, Gemini, and Mistral models.
Each sample represents a battle consisting of 2 LLMs which answer the same question, with a user label of either prefer model A, prefer model B, tie, or tie (both bad).
Citation
Please cite the… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/arena-human-preference-55k.MM-Food-100K
Overview
This project aims to introduce and release a comprehensive food image dataset designed specifically for computer vision tasks, particularly food recognition, classification, and nutritional analysis. We hope this dataset will provide a reliable resource for researchers and developers to advance the field of food AI. By publishing on Hugging Face, we expect to foster community collaboration and accelerate innovation in applications such as smart recipe recommendations… See the full description on the dataset page: https://huggingface.co/datasets/Humanbased-AI/MM-Food-100K.landuse-sentence-relevance-golden-human-set
Land-use sentence relevance golden human set
This release contains the final 300-row V3 benchmark in English plus one
parallel CSV for each of the 84 non-English project-provided sat-3l-sm
language codes. There are 85 language files in total.
Files
Every file is at
data/translations/<iso>/v3-final-<iso>.csv. The nine columns are:
sentence, label, polygon_name, h3_cell, latitude, longitude,
source, region, source_url.
The Dataset Viewer exposes these files as 85… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/landuse-sentence-relevance-golden-human-set.humaneval-rerun-scoresHumanCreativityBenchmark
The Human Creativity Benchmark (HCB)
Expert evaluations of AI-generated creative work, built to separate two signals that single-score benchmarks collapse: convergence, where professionals align around shared, checkable standards, and divergence, where creative taste legitimately differs. Each AI output is judged by domain professionals through three complementary lenses — forced-choice pairwise comparisons, 1-5 scalar ratings on prompt adherence, usability, and visual appeal… See the full description on the dataset page: https://huggingface.co/datasets/contralabs/HumanCreativityBenchmark.wmt-da-human-evaluation
Dataset Summary
This dataset contains all DA human annotations from previous WMT News Translation shared tasks.
The data is organised into 8 columns:
lp: language pair
src: input text
mt: translation
ref: reference translation
score: z score
raw: direct assessment
annotators: number of annotators
domain: domain of the input text (e.g. news)
year: collection year
You can also find the original data for each year in the results section https://www.statmt.org/wmt{YEAR}/results.html… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-da-human-evaluation.wmt-mqm-human-evaluation
Dataset Summary
This dataset contains all MQM human annotations from previous WMT Metrics shared tasks and the MQM annotations from Experts, Errors, and Context.
The data is organised into 8 columns:
lp: language pair
src: input text
mt: translation
ref: reference translation
score: MQM score
system: MT Engine that produced the translation
annotators: number of annotators
domain: domain of the input text (e.g. news)
year: collection year
You can also find the original data here.… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-mqm-human-evaluation.HumanoidRobotSoccer
Fall Prediction Dataset for Humanoid Robots
Dataset Summary
This dataset consists of 37.9 hours of real-world sensor data collected from 20 Nao humanoid robots over the course of one year in various test environments, including RoboCup soccer matches. The dataset includes 18.3 hours of walking data, featuring 2519 falls. It captures a wide range of activities such as omni-directional walking, collisions, standing up, and falls on various surfaces like artificial turf and… See the full description on the dataset page: https://huggingface.co/datasets/OliverUrbann/HumanoidRobotSoccer.emotion-negotiation-benchmarks
Emotion-Aware LLM Negotiation Benchmarks
Four high-stakes, edge-deployable negotiation benchmarks — the official evaluation suite for our research program on emotion-aware LLM agents. Each benchmark targets a distinct domain where (a) LLM-vs-LLM negotiation has real-world consequences, and (b) on-device deployment of small language models matters for privacy and latency.
The benchmarks were originally introduced with EmoMAS (ACL 2026 Main, top 9% of 12,148 submissions) and are… See the full description on the dataset page: https://huggingface.co/datasets/humanlong/emotion-negotiation-benchmarks.arc_agi_2_human_testing
ARC-AGI-2 Human testing data
This file contains data from human testing sessions on ARC-AGI tasks.
Each row represents a single test attempt by a human participant on a specific task-test pair in the "Public Train" or "Public Eval" ARC-AGI-2 datasets. Not all tasks in the released "Public Train"
sets were tested, so these results are not comprehensive. This data does not include tasks from "Semi Private Evaluation" or "Private Evaluation"
Column Descriptions… See the full description on the dataset page: https://huggingface.co/datasets/arcprize/arc_agi_2_human_testing.humanbreast_xenium_janesick # Human Breast Cancer Xenium · Sample 1 Rep1+Rep2
Curated, ready-to-load spatial transcriptomics dataset.
## Source
- Paper: [Janesick et al., Nat. Commun. 2023](https://www.nature.com/articles/s41467-023-43458-x)
- Canonical download: cf.10xgenomics.com/samples/xenium/1.0.1/Xenium_FFPE_Human_Breast_Cancer_Rep{1,2}
## Scale
| Property | Value |
|---|---|
| Technology | 10x Genomics Xenium (313-gene panel) |
| Species | Homo sapiens |
| Tissue |… See the full description on the dataset page: https://huggingface.co/datasets/Shaow/humanbreast_xenium_janesick.wmt-sqm-human-evaluation
Dataset Summary
In 2022, several changes were made to the annotation procedure used in the WMT Translation task. In contrast to the standard DA (sliding scale from 0-100) used in previous years, in 2022 annotators performed DA+SQM (Direct Assessment + Scalar Quality Metric). In DA+SQM, the annotators still provide a raw score between 0 and 100, but also are presented with seven labeled tick marks. DA+SQM helps to stabilize scores across annotators (as compared to DA).
The data is… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-sqm-human-evaluation.humanoid-robot-radarscore
Roboterradar Humanoid & Quadruped Robot Dataset
Curated editorial assessments of 21 commercially relevant humanoid robots (16) and
quadruped robots (5), with a frozen scoring methodology, evidence grades and a complete
source register.
This Hugging Face repository is a versioned distribution mirror. The canonical,
citeable publication is the Zenodo release:
Version 1.0.0 DOI: https://doi.org/10.5281/zenodo.21797689
Concept DOI for all versions:… See the full description on the dataset page: https://huggingface.co/datasets/roboterradar/humanoid-robot-radarscore.InstaArt-HumanAI
Instagram AI Art vs Human Art: Engagement & Comment Dataset
Dataset Summary
This dataset was created and contributed by Akshaya, Cynthia, Grace, and Soham as part of a project at UC San Diego.
This dataset supports research into how audiences engage with AI-generated art versus human-made art on Instagram, with a specific focus on comment sentiment, reaction types, and engagement patterns. It consists of 40 matched pairs of Instagram posts - one human art post and one… See the full description on the dataset page: https://huggingface.co/datasets/sparklessszzz/InstaArt-HumanAI.human-telemetry-driving-dataset-lite-version
Dataset Card for 15 Laps of 30Hz NGSIM-Style Telemetry
This is a Lite Version of a larger research dataset focusing on human driving signatures in high-fidelity simulations. It includes 15 full laps of telemetry captured at 30Hz within Unreal Engine 5, specifically formatted to match NGSIM standards.
Dataset Details
Dataset Description
This Lite Version dataset contains 15 laps of high-fidelity human driving telemetry. It is intended for researchers and… See the full description on the dataset page: https://huggingface.co/datasets/AtlasBuiltIt/human-telemetry-driving-dataset-lite-version.ad-creative-quality-human-vs-llm
Human Expert vs LLM Judge: Facebook Ad Creative Quality
500 real Facebook ads from 253 advertisers, each rated for creative quality by a human ad expert AND by a vision LLM — with the LLM's full reasoning.
The headline finding baked into this data: the human and the LLM agree on image quality only 26.8% of the time. The LLM judge rates 71.8% of ads "good"; the human expert rates only 20% "good". If you are using an LLM as a judge of ad creative (or any subjective visual quality)… See the full description on the dataset page: https://huggingface.co/datasets/AdControlCenter/ad-creative-quality-human-vs-llm.human_methylation_bench_ver1_test
Human DNA Methylation Dataset ver1
This dataset is a benchmark dataset for predicting the aging clock, curated from publicly available DNA methylation data. The original benchmark dataset was published by Dmitrii Kriukov et al. (2024) by integrating data from 65 individual studies.
To improve usability, we ensured unique sample IDs (excluding duplicate data, GSE118468 and GSE118469) and randomly split the data into training and testing subsets (train : test = 7 : 3) to… See the full description on the dataset page: https://huggingface.co/datasets/openaging/human_methylation_bench_ver1_test.subCat-human
SubCat: A Dataset of Subordinate Categories in Human Mind and LLMs for the Italian Language
A psycholinguistic italian dataset released with the paper How Humans and LLMs Organize Conceptual Knowledge: Exploring Subordinate Categories in Italian. It contains a list of subordiante categories, or exemplars, for 187 concrete words or, basic-level categories.
Dataset Creation
The dataset was created to study how Italian L1 speakers generate exemplars for common… See the full description on the dataset page: https://huggingface.co/datasets/ABSTRACTION-ERC/subCat-human.privacy-preserving-real-world-human-motion-sample
Privacy-Preserving Real-World Human Motion Sample
A market-validation sample of anonymous 2D skeleton/pose observations derived from a real-world indoor CCTV stream.
Why this sample exists
We are validating demand for continuously collected, privacy-oriented real-world human-motion data before expanding to multi-camera releases.
Current public sample
750 public observations
derived pose/skeleton data
anonymous track identifiers
no raw RGB video
no… See the full description on the dataset page: https://huggingface.co/datasets/Ragab-Adel/privacy-preserving-real-world-human-motion-sample.twitter-human-botshuman-jury-afd4ba
human-jury-afd4ba
Synthetic sensors test data: 36 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/rhan721/human-jury-afd4ba.innoduel-rlhf-real-world-human-preferences-sample
Real-World Human Pairwise Preferences — Public Sample
📦 This is a free, public sample of a commercial dataset.
It contains 1,350 rows curated for inspection. The full dataset has 1.5 million
human pairwise-preference decisions.
Full dataset: https://huggingface.co/datasets/NordosoftOy/innoduel-rlhf
Request access / licensing: see § Access to the full dataset — contact kari.nieminen@nordo.fi.
Use this sample to evaluate the data's quality, structure and… See the full description on the dataset page: https://huggingface.co/datasets/NordosoftOy/innoduel-rlhf-real-world-human-preferences-sample.chronos-human-ai-history-interpretationpaper: https://github.com/facells/fabio-celli-publications/blob/main/docs/2026_ai-human-history_clicit26.pdf
HAN-Humanoid-Object-Grasp-Dataset
HAN Humanoid Object Grasp Dataset
Overview
Dataset for training humanoid robots to detect and grasp objects.
Description
Contains object position data, hand joint states, and grasp success labels
collected from simulation environments.
Data Structure
object_x
object_y
object_z
hand_joint_1
hand_joint_2
grip_force
distance_to_object
grasp_success
Task Type
Supervised Learning – Binary Classification
Use Case
Robotic grasp… See the full description on the dataset page: https://huggingface.co/datasets/Caplin43/HAN-Humanoid-Object-Grasp-Dataset.humanoid_retargeting_tools_resources
Humanoid Retargeting Data
The humanoid retargeting data used in humanoid_retargeting_tools.
# make sure hf CLI is installed
curl -LsSf https://hf.co/cli/install.sh | bash
# download the dataset
hf download wty-yy/humanoid_retargeting_tools_resources --repo-type=dataset --local-dir resources
resources/data/g1/lafan1/: Converted LAFAN1 G1 retargeting dataset in NPZ format.
resources/data/g1/kimodo/: Generate data from kimodo_fock.
resources/data/g1/dailylife/: DailyLife dataset… See the full description on the dataset page: https://huggingface.co/datasets/wty-yy/humanoid_retargeting_tools_resources.Human_Gut_Microbiome_Data
Human Gut Microbiome Dataset (40-class)
Pre-processed, train/val/test-split human gut microbiome dataset used to train and evaluate
MicrobiomeFM, a Transformer-based foundation model for multi-class disease classification
from shotgun metagenomic data.
This release contains the 40-class filtered version of the dataset that the published
results were obtained on (test accuracy ≈ 96.05%, weighted F1 ≈ 0.950, macro F1 ≈ 0.691).
It is derived from the curatedMetagenomicData (cMD)… See the full description on the dataset page: https://huggingface.co/datasets/mohitraiyani27/Human_Gut_Microbiome_Data.human_methylation_bench_ver1_train
Human DNA Methylation Dataset ver1
This dataset is a benchmark dataset for predicting the aging clock, curated from publicly available DNA methylation data. The original benchmark dataset was published by Dmitrii Kriukov et al. (2024) by integrating data from 65 individual studies.
To improve usability, we ensured unique sample IDs (excluding duplicate data, GSE118468 and GSE118469) and randomly split the data into training and testing subsets (train : test = 7 : 3) to… See the full description on the dataset page: https://huggingface.co/datasets/openaging/human_methylation_bench_ver1_train.fda-peptide-human-evidence
Seven FDA-reviewed peptides: claims coded for identity, administration, outcome and replication
A claim-level comparison of the seven peptide pairs reviewed at the FDA Pharmacy Compounding Advisory Committee meeting of 23-24 July 2026. The data separates molecular identity, administration to people, claimed outcomes and independent replication.
Read the evidence-led article: https://lifesco.re/edge/which-peptide-claims-have-actually-been-tested-in-people/
Archived version and… See the full description on the dataset page: https://huggingface.co/datasets/lifescore/fda-peptide-human-evidence.humanoid-object-pushing-dataset-v1Dataset for controlled pushing and object relocation.
Description
Object weight and applied force signals mapped to pushing behaviors.
Task Description
Helps humanoid robots move objects safely by adapting push strength and body posture.
human-direction-63a8a0
human-direction-63a8a0
Synthetic sensors test data: 53 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Kenneth-Gonzalez/human-direction-63a8a0.
