datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
human_behavior_atlas
Human Behavior Atlas
A large-scale multimodal dataset for human behavior understanding, spanning emotion recognition, sentiment analysis, humor detection, mental health screening, and video question answering. The dataset integrates 16 source datasets into a unified schema with audio, video, and pre-extracted features.
This dataset was used to train OmniSapiens, a foundation model for social behavior processing.
Papers:
Human Behavior Atlas: Benchmarking Unified Psychological and… See the full description on the dataset page: https://huggingface.co/datasets/HumanBehaviorAtlas/human_behavior_atlas.mt_bench_human_judgments
Content
This dataset contains 3.3K expert-level pairwise human preferences for model responses generated by 6 models in response to 80 MT-bench questions.
The 6 models are GPT-4, GPT-3.5, Claud-v1, Vicuna-13B, Alpaca-13B, and LLaMA-13B. The annotators are mostly graduate students with expertise in the topic areas of each of the questions. The details of data collection can be found in our paper.
Agreement Calculation
This Colab notebook shows how to compute the… See the full description on the dataset page: https://huggingface.co/datasets/lmsys/mt_bench_human_judgments.human-coherence-preferences-images
Rapidata Image Generation Coherence Dataset
This dataset was collected in ~4 Days using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking it.
Overview
One of the largest human annotated coherence datasets for text-to-image models, this release contains over 1,200,000 human… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/human-coherence-preferences-images.human_behavior_atlas
Human Behavior Atlas
A large-scale multimodal dataset for human behavior understanding, spanning emotion recognition, sentiment analysis, humor detection, mental health screening, and video question answering. The dataset integrates 16 source datasets into a unified schema with audio, video, and pre-extracted features.
This dataset was used to train OmniSapiens, a foundation model for social behavior processing.
Papers:
Human Behavior Atlas: Benchmarking Unified Psychological… See the full description on the dataset page: https://huggingface.co/datasets/DennisDengHUst/human_behavior_atlas.human-alignment-preferences-images
Rapidata Image Generation Alignment Dataset
This dataset was collected in ~4 Days using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking it.
Overview
One of the largest human annotated alignment datasets for text-to-image models, this release contains over 1,200,000 human… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/human-alignment-preferences-images.OneMillion-Bench
$OneMillion-Bench
A bilingual (Global/Chinese) realistic expert-level benchmark for evaluating language agents across 5 professional domains. The benchmark contains 400 entries with detailed, weighted rubric-based grading criteria designed for fine-grained evaluation of domain expertise, analytical reasoning, and instruction following.
Dataset Structure
Each subdirectory is a Hugging Face subset (configuration), and all data is in the test split.
$OneMillion-Bench/
├──… See the full description on the dataset page: https://huggingface.co/datasets/humanlaya-data-lab/OneMillion-Bench.Flux_SD3_MJ_Dalle_Human_Coherence_Dataset
NOTE: A newer version of this dataset is available: Imagen3_Flux1.1_Flux1_SD3_MJ_Dalle_Human_Coherence_Dataset
Rapidata Image Generation Coherence Dataset
This Dataset is a 1/3 of a 2M+ human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment.
Link to the Preference dataset: https://huggingface.co/datasets/Rapidata/700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3
Link to the Text-2-Image Alignment dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Coherence_Dataset.IPHO2026
IPhO 2026 Curated Problems
This repository packages the official English problem, solution, and marking
materials for the LVI International Physics Olympiad (Bucaramanga, Colombia,
2026) as machine-readable, subquestion-level records.
Contents
Configuration
Rows
Description
all
41
All curated subquestions
theory
23
Theory papers T1–T3
experiment
18
Experimental paper E1
formalization_ready
29
Subset selected for theorem formalization… See the full description on the dataset page: https://huggingface.co/datasets/humanfia-lab/IPHO2026.KokushiMD-10
KokushiMD-10: Benchmark for Evaluating Large Language Models on Ten Japanese National Healthcare Licensing Examinations
Overview
KokushiMD-10 is the first comprehensive multimodal benchmark constructed from ten Japanese national healthcare licensing examinations. This dataset addresses critical gaps in existing medical AI evaluation by providing a linguistically grounded, multimodal, and multi-profession assessment framework for large language models (LLMs) in… See the full description on the dataset page: https://huggingface.co/datasets/humanalysis-square/KokushiMD-10.HumaniBench
HumaniBench: A Human-Centric Benchmark for Large Multimodal Models Evaluation
**HumaniBench** is a benchmark for evaluating large multimodal models (LMMs) using real-world, human-centric criteria. It consists of 32,000+ image–question pairs across 7 tasks:
✅ Open/closed VQA
🌍 Multilingual QA
📌 Visual grounding
💬 Empathetic captioning
🧠 Robustness, reasoning, and ethics
Each example is annotated with GPT-4o drafts, then verified by experts to ensure quality and… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/HumaniBench.rag-human-rights-from-files
Dataset Card for my-distiset-rag-files
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/sdiazlor/my-distiset-rag-files/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/sdiazlor/rag-human-rights-from-files.human_rank_eval
Dataset Card for HumanRankEval
This dataset supports the NAACL 2024 paper HumanRankEval: Automatic Evaluation of LMs as Conversational Assistants.
Dataset Description
Language models (LMs) as conversational assistants recently became popular tools that help people accomplish a variety of tasks. These typically result from adapting LMs pretrained on general domain text sequences through further instruction-tuning and possibly preference optimisation methods. The evaluation… See the full description on the dataset page: https://huggingface.co/datasets/huawei-noah/human_rank_eval.wikipedia-human-retrieval-ja
Japanese Wikipedia Human Retrieval dataset
This is a Japanese question answereing dataset with retrieval on Wikipedia articles
by trained human workers.
Contributors
Yusuke Oda
defined the dataset specification, data structure, and the scheme of data collection.
Baobab, Inc.
operated data collection, data checking, and formatting.
About the dataset
Each entry represents a single QA session:
given a question sentence, the responsible worker tried to search for… See the full description on the dataset page: https://huggingface.co/datasets/baobab-trees/wikipedia-human-retrieval-ja.human_anatomy_qa_with_difficulty
Truth, Trust, and Trouble (TTT) – Medical Anatomy QA Benchmark
This repository hosts the dataset introduced in the EMNLP Industry Track 2025 paper “Truth, Trust, and Trouble: Medical AI on the Edge.”
The dataset contains 1,077 high-quality, clinically validated True/False anatomy questions, designed to evaluate medical LLMs along three critical axes:
Honesty (factual alignment)
Helpfulness (semantic relevance & completeness)
Harmlessness (safety under clinical constraints)
This… See the full description on the dataset page: https://huggingface.co/datasets/ekplatebiryani/human_anatomy_qa_with_difficulty.Corp_FinanceLongContextReasoning
Corp_Finance Long-Context Reasoning — Expert-Authored Credit Agreement Benchmark (Showcase Sample)
A five-record public sample from Corp_Finance Long-Context Reasoning, a subject-matter-expert benchmarking dataset built by Human Edge for evaluating frontier model reasoning over leveraged finance and syndicated credit documentation.
Every question, reasoning trace, and answer in this dataset was authored by a practicing finance professional and independently reviewed by 2-3… See the full description on the dataset page: https://huggingface.co/datasets/HumanEdgeAI/Corp_FinanceLongContextReasoning.Corp_Taxes
Corp_Taxes — Expert-Authored US Corporate Tax Reasoning (Pilot Sample)
Seven evaluation tasks, authored and peer-reviewed by credentialed US corporate tax professionals.
Produced by Human Edge as a public demonstration of our expert-data methodology.
Curated by: Human Edge
Language: English
License: CC BY 4.0
Repository: HumanEdgeAI/Corp_Taxes
Why this dataset exists
Seven corporate tax scenarios, each written from scratch by a practicing US tax professional… See the full description on the dataset page: https://huggingface.co/datasets/HumanEdgeAI/Corp_Taxes.imabari_wiki_qa_v4_human_validated
Imabari QA v4 — Human Validated
Dataset Summary
Imabari QA v4 — Human Validated is a Japanese question-answering dataset for supervised fine-tuning (SFT).
It is derived from:
Source corpus: ikedachin/imabari_wiki_cpt_v3
QA generation model: Qwen3.8-27B-NVFP4
The dataset contains synthetic question-answer pairs together with generated reasoning traces stored in the thinking field.
The primary characteristic of this dataset is that the generated samples have been… See the full description on the dataset page: https://huggingface.co/datasets/ikedachin/imabari_wiki_qa_v4_human_validated.civil-human-rights-question-answering
Dataset Card for rag-prompt
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/sdiazlor/rag-prompt/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/sdiazlor/civil-human-rights-question-answering.human-ai-comparison
Dataset Card for Dataset Name
Contains Q&A datasets from both human and generative AI, with one AI answer and multiple human answers to each question.
MoA_Long_HumanQA
MoA: Mixture of Sparse Attention for Automatic Large Language Model Compression
This is the dataset used by the automatic sparse attention compression method MoA.
It enhances the calibration dataset by integrating long-range dependencies and model alignment.
MoA utilizes long-contextual datasets, which include question-answer pairs heavily dependent on long-range content.
The question-answer pairs are written by human in this dataset repository. Large language Models (LLMs) should… See the full description on the dataset page: https://huggingface.co/datasets/nics-efc/MoA_Long_HumanQA.human-gene-lof-rescue-evidence
Human Biallelic Loss-of-Function and Functional-Rescue Evidence
This dataset contains 88 curated human gene records linking three experimentally distinct observations:
biallelic human loss of function;
a consistent phenotype reported in independent affected families or cohorts;
functional rescue in affected humans or patient-derived human cells.
Each record therefore connects genotype → recurrent human phenotype → reversal of a disease-relevant defect. This convergent evidence… See the full description on the dataset page: https://huggingface.co/datasets/transhumanist-already-exists/human-gene-lof-rescue-evidence.Dojo-HumanFeedback-DPO
Dataset Description:
Dojo-HumanFeedback-DPO is a preference dataset designed to improve interface generation capabilities in large language models (LLMs). The dataset contains 12500 high-quality, synthetic chosen-rejected preference pairs, in the specific domain of generating frontend interfaces using HTML, CSS, and JavaScript.
The dataset format is optimized for Direct Preference Optimization (DPO), but can potentially be used in other machine learning contexts.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/tensorplex-labs/Dojo-HumanFeedback-DPO.Humanization_001
Creating Human Advance AI
Success is a game of winners.
— # Leroy Dyer (1972-Present)
Thinking Humanly:
AI aims to model human thought, a goal of cognitive science across fields like psychology and computer science.
Thinking Rationally:
AI also seeks to formalize “laws of thought” through logic, though human thinking is often inconsistent and uncertain.
Acting Humanly:
Turing's test evaluates AI by its ability to mimic human behavior convincingly… See the full description on the dataset page: https://huggingface.co/datasets/LeroyDyer/Humanization_001.LegalReasoning
Human Edge — Legal Reasoning Evaluation (Showcase Sample)
A public 6-task sample from a rubric-based legal reasoning evaluation dataset built by
Human Edge (humanedgetech.ai). Each task is authored and reviewed by
practicing senior lawyers and is designed to produce a verifiable, per-criterion reward signal
for post-training and evaluation of frontier language models on high-complexity legal work.
The sample contains one task per legal subdomain, drawn from a larger internal… See the full description on the dataset page: https://huggingface.co/datasets/HumanEdgeAI/LegalReasoning.rag-human-rights-from-prompt
Dataset Card for datset-rag-prompt
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/sdiazlor/datset-rag-prompt/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/sdiazlor/rag-human-rights-from-prompt.openai_humaneval-th
HumanEval-th
A Thai translation of all 164 problems of OpenAI's HumanEval. Every row corresponds
1:1, in order, to a row of the English original, so the Thai and English scores of a
model are directly comparable.
Only the prompt column is Thai. canonical_solution, test and entry_point are
Python rather than prose and were never translated; they are byte-identical to
openai/openai_humaneval in all 164 rows, and validate.py checks that on every run.
Every revision of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/iapp/openai_humaneval-th.mind2web-subset-human
Mind2Web Subset - Human Demonstrations
A collection of human-demonstrated web navigation tasks with detailed interaction traces. This dataset captures real browser interactions including clicks, typing, scrolling, DOM states, screenshots, and HTTP requests for web agent training and evaluation.
Overview
This dataset contains tasks performed by humans in real web environments, capturing:
Golden trajectories: Step-by-step sequences of actions (clicks, typing, navigation)… See the full description on the dataset page: https://huggingface.co/datasets/josancamon/mind2web-subset-human.LLM-Failure-Cases
Codatta LLM Failure Cases (Expert Critiques)
Overview
Codatta LLM Failure Cases is a specialized adversarial dataset designed to highlight and analyze scenarios where state-of-the-art Large Language Models (LLMs) produce incorrect, hallucinatory, or logically flawed responses.
This dataset originates from Codatta's "Airdrop Season 1" campaign, a crowdsourced data intelligence initiative where participants were tasked with finding prompts that caused leading LLMs… See the full description on the dataset page: https://huggingface.co/datasets/Humanbased-AI/LLM-Failure-Cases.MedicalReasoning
Project Aletheia — Expert-Grounded Medical QA Evaluation (6-Case Illustrative Sample)
A 6-record sample drawn from the complete Project Aletheia dataset — a 52-case, subject-matter-expert evaluation pilot built by Human Edge to ground medical question-answering preference data in licensed clinical judgment.
Every gold-standard response, segment agreement count, pairwise preference, and reasoning trace was authored or adjudicated by a licensed medical doctor, through a pipeline… See the full description on the dataset page: https://huggingface.co/datasets/HumanEdgeAI/MedicalReasoning.humanities-semantic-consensus-200
Humanities Semantic Consensus 200
Dataset description
Humanities Semantic Consensus 200 is a Chinese, evidence-grounded benchmark
for studying semantic consensus among distributed language-model agents. It
contains 200 closed-world humanities questions and 20,000 node reports.
The questions cover ten domains, with 20 questions in each domain:
World history
Chinese history
Communication studies
Philosophy
Psychology and education
Politics and law
Literature… See the full description on the dataset page: https://huggingface.co/datasets/yyfanfytfyt/humanities-semantic-consensus-200.
