datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HumanEval-XLThis dataset contains a viewer-friendly version of the dataset at FloatAI/HumanEval-XL. It is made available separately for the convenience of the vllm-code-harness package.
QIT
QIT Humanize-Physic Formalizations and Proofs
QIT (Quantum Information Theory) is a blind benchmark for formalizing theorems in quantum information. It evaluates whether an AI agent can faithfully translate natural-language and TeX problem statements into Lean 4 theorems and then construct formal proofs checked by the Lean kernel. Its 40 tasks cover quantum channels and Choi representations, entropy and coding, mixed-unitary obstructions and symmetry, norm and fidelity tools… See the full description on the dataset page: https://huggingface.co/datasets/humanfia-lab/QIT.humans-top
humans.top — LIVE Global ranking of influential people (open dataset)
This dataset ranks real, named living people by global influence — e.g. #1
Donald Trump, #2 Xi Jinping, #3 Vladimir Putin, alongside figures like Elon Musk,
Narendra Modi and Lionel Messi. Every row is a person: their live influence
rank, a concise biography in 15 languages, and Wikidata / Wikipedia links.
Published from the website humans.top (.top is the
domain name).
Available on (identical CC0… See the full description on the dataset page: https://huggingface.co/datasets/dsfox/humans-top.OneMillion-Bench
$OneMillion-Bench
A bilingual (Global/Chinese) realistic expert-level benchmark for evaluating language agents across 5 professional domains. The benchmark contains 400 entries with detailed, weighted rubric-based grading criteria designed for fine-grained evaluation of domain expertise, analytical reasoning, and instruction following.
Dataset Structure
Each subdirectory is a Hugging Face subset (configuration), and all data is in the test split.
$OneMillion-Bench/
├──… See the full description on the dataset page: https://huggingface.co/datasets/humanlaya-data-lab/OneMillion-Bench.QAlg
QAlg Humanize-Physic Formalizations and Proofs
QAlg (Quantum Algorithms) is a blind benchmark for formalizing theorems in quantum algorithms. It evaluates whether an AI agent can faithfully translate natural-language and TeX problem statements into Lean 4 theorems and then construct formal proofs checked by the Lean kernel. Its 36 tasks cover quantum circuits, linear algebra, the quantum Fourier transform, Hamiltonian simulation, hidden subgroups, QSP/QSVT, and parameterized… See the full description on the dataset page: https://huggingface.co/datasets/humanfia-lab/QAlg.human_behavior_atlas_tar
Human Behavior Atlas (HBA)
Human Behavior Atlas (HBA) is a unified benchmark for multimodal behavioral understanding.It aggregates and standardizes multiple behavioral datasets into a single training and evaluation framework, enabling consistent training and evaluation of foundation models on psychological and social behavior tasks (e.g., emotion, intent, sarcasm, mental health signals, nonverbal behavior).
Dataset on Hugging Face:… See the full description on the dataset page: https://huggingface.co/datasets/HumanBehaviorAtlas/human_behavior_atlas_tar.human-templated-captions-1bcsv delimiter is = ".,|,."
apparently python doesn't like multichar delimiters using the native csv so there's some issues with environments when loading.
This seemed like a good idea to avoid overlapping potential characters, but in practice it turned into additional overhead and bugs. I'll be manually converting the split to parquet and providing a proper file split soon.
Additionally with the parquet will introduce the large caption split; which are considerably longer captions for the… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/human-templated-captions-1b.human_assistant_conversationbc-humanevalThe HumanEval dataset in BabelCode format.HumanCreativityBenchmark
The Human Creativity Benchmark (HCB)
Expert evaluations of AI-generated creative work, built to separate two signals that single-score benchmarks collapse: convergence, where professionals align around shared, checkable standards, and divergence, where creative taste legitimately differs. Each AI output is judged by domain professionals through three complementary lenses — forced-choice pairwise comparisons, 1-5 scalar ratings on prompt adherence, usability, and visual appeal… See the full description on the dataset page: https://huggingface.co/datasets/contralabs/HumanCreativityBenchmark.icho-2026
IChO 2026 Lean 4 formalizations: verified model variants
This repository contains three independently generated Lean 4 proof sets for the
same 32 selected IChO 2026 theory subquestions. Practical papers P1–P3 remain
outside the corpus.
Proof-origin labels
Config
Records
proof_generator.label
Meaning
kimi-k3
32
Kimi-K3
Proofs generated in the clean K3 rerun with kimi-k3[1m] through Claude Code.
gpt
32
GPT
Proofs generated in a fresh answer-blind… See the full description on the dataset page: https://huggingface.co/datasets/humanfia-lab/icho-2026.kodcode-humaneval-like
KodCodeHumanEvalLike
Strict verified HumanEval-compatible conversion of
KodCode/KodCode-V1.
This is a derived dataset. It is not OpenAI HumanEval and should not be reported
as HumanEval. It follows the HumanEval-style JSONL schema and execution protocol
so code-generation pipelines can evaluate it with the same check(candidate)
interface.
Each JSONL record has the HumanEval-style fields:
task_id
prompt
canonical_solution
test
entry_point
The prompt is completion-style:
def… See the full description on the dataset page: https://huggingface.co/datasets/BOB12311/kodcode-humaneval-like.wikipedia-human-ai
Wikipedia Human/AI
A selection of ~10,000 paragraphs from Wikipedia, along with rewritten text by GPT 5 Nano.
human_behavior_atlas
Human Behavior Atlas (HBA)
Human Behavior Atlas (HBA) is a unified benchmark for multimodal behavioral understanding.It aggregates and standardizes multiple behavioral datasets into a single training and evaluation framework, enabling consistent training and evaluation of foundation models on psychological and social behavior tasks (e.g., emotion, intent, sarcasm, mental health signals, nonverbal behavior).
Dataset on Hugging Face:… See the full description on the dataset page: https://huggingface.co/datasets/droiden/human_behavior_atlas.emotion-negotiation-benchmarks
Emotion-Aware LLM Negotiation Benchmarks
Four high-stakes, edge-deployable negotiation benchmarks — the official evaluation suite for our research program on emotion-aware LLM agents. Each benchmark targets a distinct domain where (a) LLM-vs-LLM negotiation has real-world consequences, and (b) on-device deployment of small language models matters for privacy and latency.
The benchmarks were originally introduced with EmoMAS (ACL 2026 Main, top 9% of 12,148 submissions) and are… See the full description on the dataset page: https://huggingface.co/datasets/humanlong/emotion-negotiation-benchmarks.sorry-bench-human-judgment-202406
Dataset Card for 🧑⚖️SORRY-Bench Human Judgment Dataset (2024/06)
🏠Website
📑Paper
📚Dataset
💻Github
🧑⚖️Human Judgment Dataset
🤖Judge LLM
This dataset contains 7.2K annotations of human safety judgmentsfor LLM responses to unsafe instructions of our SORRY-Bench dataset.
Specifically, for each unsafe instruction of the 450 unsafe instructions in SORRY-Bench dataset, we annotate 16 diverse model responses (both ID and OOD) as either in… See the full description on the dataset page: https://huggingface.co/datasets/sorry-bench/sorry-bench-human-judgment-202406.mediationbench-sample
MediationBench Sample
This repository contains a small synthetic demonstration of the
MediationBench data contract. It is intended to
help developers inspect and prototype mediation agents. It is not a training
corpus, a leaderboard result, or evidence that AI mediation is safe or
effective with people.
What is included
The demo split contains two matched, fictional conversations that begin from
the same business-partnership dispute:
mediated: simulated parties… See the full description on the dataset page: https://huggingface.co/datasets/HumanAssisted/mediationbench-sample.humans-benchmark
HUMANS Benchmark Dataset
Authors: Woody Haosheng Gan¹, William Held²'³, Diyi Yang²
¹University of Southern California, ²Stanford University, ³OpenAthena
This dataset is part of the Putting HUMANS first: Efficient LAM Evaluation with Human Preference Alignment paper.
HUMANS (HUman-aligned Minimal Audio evaluatioN Subsets for Large Audio Models) Benchmark is designed to efficiently evaluate Large Audio Models using minimal subsets while predicting human preferences through learned… See the full description on the dataset page: https://huggingface.co/datasets/woodygan/humans-benchmark.humanevalHumanEval dataset annotated with the ground-truth programming solutions, to enable evaluations for retrieval and retrieval augmented code generation.
Please refer to code-rag-becnch for more details.
HumanML3D-500ms-FPP-descriptions-CoTs-1
HumanML3D 500ms First person perspective descriptions for CoTs
Introduction
This repository contains files of the Mr. Ri's and Ms. Tique's HumanML3D human motion dataset,
but also descriptions of the movements in first person perspective in 0.5 second time windows.
The descriptions were created synthetically with use of a multimodal LLM and are in json format. They can be found in comics_and_descriptions folder.
The dataset also contains motion capture data and… See the full description on the dataset page: https://huggingface.co/datasets/Wojtekb30/HumanML3D-500ms-FPP-descriptions-CoTs-1.sorry-bench-human-judgment-202503
Dataset Card for 🧑⚖️SORRY-Bench Human Judgment Dataset (2025/03)
🏠Website
📑Paper
📚Dataset
💻Github
🧑⚖️Human Judgment Dataset
🤖Judge LLM
🪧UPDATE: In this iteration, we removed the category "Impersonation" due to its ambiguous definition, and that most models more or less fulfill such requests.This dataset contains 7K annotations of human safety judgments for LLM responses to unsafe instructions of our SORRY-Bench dataset.
Specifically, for… See the full description on the dataset page: https://huggingface.co/datasets/sorry-bench/sorry-bench-human-judgment-202503.gohumanize-open-humanizer-dataset
GoHumanize Open Humanizer Dataset
2,957 training pairs and 300 test pairs for teaching a language model to rewrite
AI-styled English prose into natural human writing. Each pair is:
input: a passage rewritten by a large language model in the register typical of LLM output
(formal, smooth, hedged, connective phrases, no contractions);
output: the original human-written passage, from a public-domain book or, since version 2,
from a US federal government publication.
The human… See the full description on the dataset page: https://huggingface.co/datasets/gohumanize/gohumanize-open-humanizer-dataset.rag-human-rights-from-files
Dataset Card for my-distiset-rag-files
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/sdiazlor/my-distiset-rag-files/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/sdiazlor/rag-human-rights-from-files.human-vibehuman_rank_eval
Dataset Card for HumanRankEval
This dataset supports the NAACL 2024 paper HumanRankEval: Automatic Evaluation of LMs as Conversational Assistants.
Dataset Description
Language models (LMs) as conversational assistants recently became popular tools that help people accomplish a variety of tasks. These typically result from adapting LMs pretrained on general domain text sequences through further instruction-tuning and possibly preference optimisation methods. The evaluation… See the full description on the dataset page: https://huggingface.co/datasets/huawei-noah/human_rank_eval.synthetic_human_pointingSmall dataset containing synthetic images of artificially generated persons superimposed onto backgrounds with labelled output captions for VLM object detection and pointing
human_assistant_conversation_deduped
Deduplicated version of Isotonic/human_assistant_conversation
Deduped with max jaccard similarity of 0.75
human-ai-impact-bench-scenarios
HumanAI-Impact-Bench — Scenarios
Bilingual (English / Vietnamese) scenario set for evaluating how conversational
AI systems affect human emotion, autonomy, cognition, trust, and social
connection. Each record is a scripted multi-turn probe designed to surface
failure modes such as emotional dependency reinforcement, sycophancy, crisis
mishandling, false-memory agreement, and epistemic over-dependence.
Code / tooling: https://github.com/lamduong0/human-ai-impact-bench
License:… See the full description on the dataset page: https://huggingface.co/datasets/lamduong/human-ai-impact-bench-scenarios.task714_mmmlu_answer_generation_human_sexuality
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task714_mmmlu_answer_generation_human_sexuality
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task714_mmmlu_answer_generation_human_sexuality.gencode-human
GENCODE
GENCODE is a comprehensive annotation project that aims to provide high-quality annotations of the human and mouse genomes.
The project is part of the ENCODE (ENCyclopedia Of DNA Elements) scale-up project, which seeks to identify all functional elements in the human genome.
Disclaimer
This is an UNOFFICIAL release of the GENCODE by Paul Flicek, Roderic Guigo, Manolis Kellis, Mark Gerstein, Benedict Paten, Michael Tress, Jyoti Choudhary, et al.
The team… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/gencode-human.
