datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arena-human-preference-55kDataset for Kaggle competition on predicting human preference on Chatbot Arena battles.
The training dataset includes over 55,000 real-world user and LLM conversations and user preferences across over 70 state-of-the-art LLMs, such as GPT-4, Claude 2, Llama 2, Gemini, and Mistral models.
Each sample represents a battle consisting of 2 LLMs which answer the same question, with a user label of either prefer model A, prefer model B, tie, or tie (both bad).
Citation
Please cite the… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/arena-human-preference-55k.pizza_st_human_mouse_v1_train_with_labelshuman_ai_generated_text
Human or AI-Generated Text
The data can be valuable for educators, policymakers, and researchers interested in the evolving education landscape, particularly in detecting or identifying texts written by Humans or Artificial Intelligence systems.
File Name
model_training_dataset.csv
File Structure
id: Unique identifier for each record.
human_text: Human-written content.
ai_text: AI-generated texts.
instructions: Description of the task given to both Humans and… See the full description on the dataset page: https://huggingface.co/datasets/dmitva/human_ai_generated_text.Refined-TCGA-PRAD-Prostate-Cancer-Pathology-Dataset
Note: Our 245 TCGA cases are ones we identified as having potential for improvement.
We plan to upload them in two phases: the first batch of 138 cases, and the second batch of 107 cases in the quality review pipeline, we plan to upload them around early of January, 2025.
Dataset: A Second Opinion on TCGA PRAD Prostate Dataset Labels with ROI-Level Annotations
Overview
This dataset provides enhanced Gleason grading annotations for the TCGA PRAD prostate cancer… See the full description on the dataset page: https://huggingface.co/datasets/Humanbased-AI/Refined-TCGA-PRAD-Prostate-Cancer-Pathology-Dataset.human-templated-captions-1bcsv delimiter is = ".,|,."
apparently python doesn't like multichar delimiters using the native csv so there's some issues with environments when loading.
This seemed like a good idea to avoid overlapping potential characters, but in practice it turned into additional overhead and bugs. I'll be manually converting the split to parquet and providing a proper file split soon.
Additionally with the parquet will introduce the large caption split; which are considerably longer captions for the… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/human-templated-captions-1b.AI-and-Human-Generated-Text
AI & Human Generated Text
I am Using this dataset for AI Text Detection for https://exnrt.com.
Check Original DataSet GitHub Repository Here: https://github.com/panagiotisanagnostou/AI-GA
Description
The AI-GA dataset, short for Artificial Intelligence Generated Abstracts, comprises abstracts and titles. Half of these abstracts are generated by AI, while the remaining half are original. Primarily intended for research and experimentation in natural language… See the full description on the dataset page: https://huggingface.co/datasets/Ateeqq/AI-and-Human-Generated-Text.MM-Food-100K
Overview
This project aims to introduce and release a comprehensive food image dataset designed specifically for computer vision tasks, particularly food recognition, classification, and nutritional analysis. We hope this dataset will provide a reliable resource for researchers and developers to advance the field of food AI. By publishing on Hugging Face, we expect to foster community collaboration and accelerate innovation in applications such as smart recipe recommendations… See the full description on the dataset page: https://huggingface.co/datasets/Humanbased-AI/MM-Food-100K.Crypto-Address-Annotation-10K
Codatta Crypto Address Annotations (Sample)
Overview
This dataset is a 10,000-row sample of the comprehensive Codatta Crypto Address Annotations database. The full database serves as a massive repository of over 500 million labeled address pairs across multiple blockchains.
The data provides critical metadata aimed at solving the problem of fragmented and siloed blockchain information. It includes entity names, functional categories (e.g., Exchanges, DeFi, Scam)… See the full description on the dataset page: https://huggingface.co/datasets/Humanbased-AI/Crypto-Address-Annotation-10K.Humans_with_Collision
Humans with Collisions (HwC) Pose & Motion Dataset
This dataset contains the training, evaluation, and benchmark data for the paper:"PoseShield: Neural Collision Fields for Human Self-Collision Resolution (ECCV 2026)"
Paper (arXiv): arXiv:2606.29686
Code Repository: PoseShield on GitHub (or project repo)
Dataset Structure
The repository contains two main groups of data structured under the data/ directory:
1. HwC Pose Dataset (Single Poses)
Used… See the full description on the dataset page: https://huggingface.co/datasets/ZYYY99/Humans_with_Collision.landuse-sentence-relevance-golden-human-set
Land-use sentence relevance golden human set
This release contains the final 300-row V3 benchmark in English plus one
parallel CSV for each of the 84 non-English project-provided sat-3l-sm
language codes. There are 85 language files in total.
Files
Every file is at
data/translations/<iso>/v3-final-<iso>.csv. The nine columns are:
sentence, label, polygon_name, h3_cell, latitude, longitude,
source, region, source_url.
The Dataset Viewer exposes these files as 85… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/landuse-sentence-relevance-golden-human-set.humaneval-rerun-scoreswmt-da-human-evaluation
Dataset Summary
This dataset contains all DA human annotations from previous WMT News Translation shared tasks.
The data is organised into 8 columns:
lp: language pair
src: input text
mt: translation
ref: reference translation
score: z score
raw: direct assessment
annotators: number of annotators
domain: domain of the input text (e.g. news)
year: collection year
You can also find the original data for each year in the results section https://www.statmt.org/wmt{YEAR}/results.html… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-da-human-evaluation.wmt-mqm-human-evaluation
Dataset Summary
This dataset contains all MQM human annotations from previous WMT Metrics shared tasks and the MQM annotations from Experts, Errors, and Context.
The data is organised into 8 columns:
lp: language pair
src: input text
mt: translation
ref: reference translation
score: MQM score
system: MT Engine that produced the translation
annotators: number of annotators
domain: domain of the input text (e.g. news)
year: collection year
You can also find the original data here.… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-mqm-human-evaluation.ai-human-text-detection-v1
🧠 AI vs Human Text Detection Dataset (v1)
This dataset merges nine major public and academic corpora to form one of the most comprehensive resources for AI-generated text detection model training and evaluation.
🔗 Sources
The dataset consolidates, cleans, and standardizes multiple open datasets and research benchmarks, each focusing on human vs. AI-generated text classification:
Hello-SimpleAI / HC3 — Human–ChatGPT comparison corpus
gsingh1-py / train — Large-scale… See the full description on the dataset page: https://huggingface.co/datasets/silentone0725/ai-human-text-detection-v1.HumanCreativityBenchmark
The Human Creativity Benchmark (HCB)
Expert evaluations of AI-generated creative work, built to separate two signals that single-score benchmarks collapse: convergence, where professionals align around shared, checkable standards, and divergence, where creative taste legitimately differs. Each AI output is judged by domain professionals through three complementary lenses — forced-choice pairwise comparisons, 1-5 scalar ratings on prompt adherence, usability, and visual appeal… See the full description on the dataset page: https://huggingface.co/datasets/contralabs/HumanCreativityBenchmark.HumanoidRobotSoccer
Fall Prediction Dataset for Humanoid Robots
Dataset Summary
This dataset consists of 37.9 hours of real-world sensor data collected from 20 Nao humanoid robots over the course of one year in various test environments, including RoboCup soccer matches. The dataset includes 18.3 hours of walking data, featuring 2519 falls. It captures a wide range of activities such as omni-directional walking, collisions, standing up, and falls on various surfaces like artificial turf and… See the full description on the dataset page: https://huggingface.co/datasets/OliverUrbann/HumanoidRobotSoccer.VBench-I2V_human_annotationHumanEvalComm
HumanEvalComm: Benchmarking the Communication Skills of Code Generation for LLMs and LLM Agent
📄 Paper •
💻 GitHub Repository •
🤗 Dataset Viewer
Dataset Description
HumanEvalComm is a benchmark dataset for evaluating the communication skills of Large Language Models (LLMs) in code generation tasks. It is built upon the widely used HumanEval benchmark. HumanEvalComm contains 762 modified problem descriptions based on the 164 problems in the… See the full description on the dataset page: https://huggingface.co/datasets/jie-jw-wu/HumanEvalComm.emotion-negotiation-benchmarks
Emotion-Aware LLM Negotiation Benchmarks
Four high-stakes, edge-deployable negotiation benchmarks — the official evaluation suite for our research program on emotion-aware LLM agents. Each benchmark targets a distinct domain where (a) LLM-vs-LLM negotiation has real-world consequences, and (b) on-device deployment of small language models matters for privacy and latency.
The benchmarks were originally introduced with EmoMAS (ACL 2026 Main, top 9% of 12,148 submissions) and are… See the full description on the dataset page: https://huggingface.co/datasets/humanlong/emotion-negotiation-benchmarks.humancentric-scenes-ai
HumanCentric-Scenes-AI
A multimodal benchmark of 296 AI-generated human-centric scenes across four domains:
CCTV / surveillance imagery (Set 2, 85 images). Midjourney-generated stills that mimic low-resolution security-camera footage — parking lots, building interiors, outdoor public spaces — designed to test whether detection cues survive heavy compression and low-light noise.
Occupation × gender portraits (Set 3, 128 images). A balanced 64-occupation × 2-gender paired design… See the full description on the dataset page: https://huggingface.co/datasets/Nima0Kamali/humancentric-scenes-ai.vietnamese_news_human_ai
Detecting AI-Generated Vietnamese News Articles with Multilingual-E5 and BERT
This is the official dataset accompanying the paper Detecting AI-Generated Vietnamese News Articles with Multilingual-E5 and BERT, which was accepted at ICCIES 2025 and published in Computational Intelligence in Engineering Science (Springer CCIS, vol. 2587).
You can read the paper here: Detecting AI-Generated Vietnamese News Articles with Multilingual-E5 and BERT
Abstract
The emergence… See the full description on the dataset page: https://huggingface.co/datasets/ICCIES-2025-DetectAI/vietnamese_news_human_ai.humancentric-scenes-ai
HumanCentric-Scenes-AI
A multimodal benchmark of 296 AI-generated human-centric scenes across four domains:
CCTV / surveillance imagery (Set 2, 85 images). Midjourney-generated stills that mimic low-resolution security-camera footage — parking lots, building interiors, outdoor public spaces — designed to test whether detection cues survive heavy compression and low-light noise.
Occupation × gender portraits (Set 3, 128 images). A balanced 64-occupation × 2-gender paired design… See the full description on the dataset page: https://huggingface.co/datasets/rjmaftv33/humancentric-scenes-ai.IELTS_essay_human_feedbackhumanbreast_xenium_janesick # Human Breast Cancer Xenium · Sample 1 Rep1+Rep2
Curated, ready-to-load spatial transcriptomics dataset.
## Source
- Paper: [Janesick et al., Nat. Commun. 2023](https://www.nature.com/articles/s41467-023-43458-x)
- Canonical download: cf.10xgenomics.com/samples/xenium/1.0.1/Xenium_FFPE_Human_Breast_Cancer_Rep{1,2}
## Scale
| Property | Value |
|---|---|
| Technology | 10x Genomics Xenium (313-gene panel) |
| Species | Homo sapiens |
| Tissue |… See the full description on the dataset page: https://huggingface.co/datasets/Shaow/humanbreast_xenium_janesick.arc_agi_2_human_testing
ARC-AGI-2 Human testing data
This file contains data from human testing sessions on ARC-AGI tasks.
Each row represents a single test attempt by a human participant on a specific task-test pair in the "Public Train" or "Public Eval" ARC-AGI-2 datasets. Not all tasks in the released "Public Train"
sets were tested, so these results are not comprehensive. This data does not include tasks from "Semi Private Evaluation" or "Private Evaluation"
Column Descriptions… See the full description on the dataset page: https://huggingface.co/datasets/arcprize/arc_agi_2_human_testing.wmt-sqm-human-evaluation
Dataset Summary
In 2022, several changes were made to the annotation procedure used in the WMT Translation task. In contrast to the standard DA (sliding scale from 0-100) used in previous years, in 2022 annotators performed DA+SQM (Direct Assessment + Scalar Quality Metric). In DA+SQM, the annotators still provide a raw score between 0 and 100, but also are presented with seven labeled tick marks. DA+SQM helps to stabilize scores across annotators (as compared to DA).
The data is… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-sqm-human-evaluation.humanoid-robot-radarscore
Roboterradar Humanoid & Quadruped Robot Dataset
Curated editorial assessments of 21 commercially relevant humanoid robots (16) and
quadruped robots (5), with a frozen scoring methodology, evidence grades and a complete
source register.
This Hugging Face repository is a versioned distribution mirror. The canonical,
citeable publication is the Zenodo release:
Version 1.0.0 DOI: https://doi.org/10.5281/zenodo.21797689
Concept DOI for all versions:… See the full description on the dataset page: https://huggingface.co/datasets/roboterradar/humanoid-robot-radarscore.InstaArt-HumanAI
Instagram AI Art vs Human Art: Engagement & Comment Dataset
Dataset Summary
This dataset was created and contributed by Akshaya, Cynthia, Grace, and Soham as part of a project at UC San Diego.
This dataset supports research into how audiences engage with AI-generated art versus human-made art on Instagram, with a specific focus on comment sentiment, reaction types, and engagement patterns. It consists of 40 matched pairs of Instagram posts - one human art post and one… See the full description on the dataset page: https://huggingface.co/datasets/sparklessszzz/InstaArt-HumanAI.human-telemetry-driving-dataset-lite-version
Dataset Card for 15 Laps of 30Hz NGSIM-Style Telemetry
This is a Lite Version of a larger research dataset focusing on human driving signatures in high-fidelity simulations. It includes 15 full laps of telemetry captured at 30Hz within Unreal Engine 5, specifically formatted to match NGSIM standards.
Dataset Details
Dataset Description
This Lite Version dataset contains 15 laps of high-fidelity human driving telemetry. It is intended for researchers and… See the full description on the dataset page: https://huggingface.co/datasets/AtlasBuiltIt/human-telemetry-driving-dataset-lite-version.CounselBench-Adv-human-annotation
