CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lmarena-ai /arena-human-preference-55kDataset for Kaggle competition on predicting human preference on Chatbot Arena battles. The training dataset includes over 55,000 real-world user and LLM conversations and user preferences across over 70 state-of-the-art LLMs, such as GPT-4, Claude 2, Llama 2, Gemini, and Mistral models. Each sample represents a battle consisting of 2 LLMs which answer the same question, with a user label of either prefer model A, prefer model B, tie, or tie (both bad). Citation Please cite the… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/arena-human-preference-55k.tabulartext-classification10K<n<100K159 likes2.4k downloads2y agoHugging Face02HaoranoLee /pizza_st_human_mouse_v1_train_with_labelstextn<1K0 likes822 downloads9mo agoHugging Face03dmitva /human_ai_generated_text Human or AI-Generated Text The data can be valuable for educators, policymakers, and researchers interested in the evolving education landscape, particularly in detecting or identifying texts written by Humans or Artificial Intelligence systems. File Name model_training_dataset.csv File Structure id: Unique identifier for each record. human_text: Human-written content. ai_text: AI-generated texts. instructions: Description of the task given to both Humans and… See the full description on the dataset page: https://huggingface.co/datasets/dmitva/human_ai_generated_text.text1M<n<10M38 likes537 downloads3y agoHugging Face04Humanbased-AI /Refined-TCGA-PRAD-Prostate-Cancer-Pathology-Dataset Note: Our 245 TCGA cases are ones we identified as having potential for improvement. We plan to upload them in two phases: the first batch of 138 cases, and the second batch of 107 cases in the quality review pipeline, we plan to upload them around early of January, 2025. Dataset: A Second Opinion on TCGA PRAD Prostate Dataset Labels with ROI-Level Annotations Overview This dataset provides enhanced Gleason grading annotations for the TCGA PRAD prostate cancer… See the full description on the dataset page: https://huggingface.co/datasets/Humanbased-AI/Refined-TCGA-PRAD-Prostate-Cancer-Pathology-Dataset.geospatialn<1K16 likes355 downloads2y agoHugging Face05AbstractPhil /human-templated-captions-1bcsv delimiter is = ".,|,." apparently python doesn't like multichar delimiters using the native csv so there's some issues with environments when loading. This seemed like a good idea to avoid overlapping potential characters, but in practice it turned into additional overhead and bugs. I'll be manually converting the split to parquet and providing a proper file split soon. Additionally with the parquet will introduce the large caption split; which are considerably longer captions for the… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/human-templated-captions-1b.texttext-generation100M<n<1B1 likes331 downloads1y agoHugging Face06Ateeqq /AI-and-Human-Generated-Text AI & Human Generated Text I am Using this dataset for AI Text Detection for https://exnrt.com. Check Original DataSet GitHub Repository Here: https://github.com/panagiotisanagnostou/AI-GA Description The AI-GA dataset, short for Artificial Intelligence Generated Abstracts, comprises abstracts and titles. Half of these abstracts are generated by AI, while the remaining half are original. Primarily intended for research and experimentation in natural language… See the full description on the dataset page: https://huggingface.co/datasets/Ateeqq/AI-and-Human-Generated-Text.texttext-classification10K<n<100K24 likes280 downloads2y agoHugging Face07Humanbased-AI /MM-Food-100K Overview This project aims to introduce and release a comprehensive food image dataset designed specifically for computer vision tasks, particularly food recognition, classification, and nutritional analysis. We hope this dataset will provide a reliable resource for researchers and developers to advance the field of food AI. By publishing on Hugging Face, we expect to foster community collaboration and accelerate innovation in applications such as smart recipe recommendations… See the full description on the dataset page: https://huggingface.co/datasets/Humanbased-AI/MM-Food-100K.imageimage-classification100K<n<1M73 likes277 downloads1y agoHugging Face08Humanbased-AI /Crypto-Address-Annotation-10K Codatta Crypto Address Annotations (Sample) Overview This dataset is a 10,000-row sample of the comprehensive Codatta Crypto Address Annotations database. The full database serves as a massive repository of over 500 million labeled address pairs across multiple blockchains. The data provides critical metadata aimed at solving the problem of fragmented and siloed blockchain information. It includes entity names, functional categories (e.g., Exchanges, DeFi, Scam)… See the full description on the dataset page: https://huggingface.co/datasets/Humanbased-AI/Crypto-Address-Annotation-10K.texttoken-classification10K<n<100K1 likes246 downloads10mo agoHugging Face09ZYYY99 /Humans_with_Collision Humans with Collisions (HwC) Pose & Motion Dataset This dataset contains the training, evaluation, and benchmark data for the paper:"PoseShield: Neural Collision Fields for Human Self-Collision Resolution (ECCV 2026)" Paper (arXiv): arXiv:2606.29686 Code Repository: PoseShield on GitHub (or project repo) Dataset Structure The repository contains two main groups of data structured under the data/ directory: 1. HwC Pose Dataset (Single Poses) Used… See the full description on the dataset page: https://huggingface.co/datasets/ZYYY99/Humans_with_Collision.3drobotics1K<n<10K0 likes235 downloads2mo agoHugging Face10NoeFlandre /landuse-sentence-relevance-golden-human-set Land-use sentence relevance golden human set This release contains the final 300-row V3 benchmark in English plus one parallel CSV for each of the 84 non-English project-provided sat-3l-sm language codes. There are 85 language files in total. Files Every file is at data/translations/<iso>/v3-final-<iso>.csv. The nine columns are: sentence, label, polygon_name, h3_cell, latitude, longitude, source, region, source_url. The Dataset Viewer exposes these files as 85… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/landuse-sentence-relevance-golden-human-set.tabulartext-classification10K<n<100K0 likes224 downloads15h agoHugging Face11latkes /humaneval-rerun-scorestabular100K<n<1M0 likes221 downloads3mo agoHugging Face12RicardoRei /wmt-da-human-evaluation Dataset Summary This dataset contains all DA human annotations from previous WMT News Translation shared tasks. The data is organised into 8 columns: lp: language pair src: input text mt: translation ref: reference translation score: z score raw: direct assessment annotators: number of annotators domain: domain of the input text (e.g. news) year: collection year You can also find the original data for each year in the results section https://www.statmt.org/wmt{YEAR}/results.html… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-da-human-evaluation.tabular1M<n<10M10 likes202 downloads4y agoHugging Face13RicardoRei /wmt-mqm-human-evaluation Dataset Summary This dataset contains all MQM human annotations from previous WMT Metrics shared tasks and the MQM annotations from Experts, Errors, and Context. The data is organised into 8 columns: lp: language pair src: input text mt: translation ref: reference translation score: MQM score system: MT Engine that produced the translation annotators: number of annotators domain: domain of the input text (e.g. news) year: collection year You can also find the original data here.… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-mqm-human-evaluation.tabular100K<n<1M1 likes193 downloads4y agoHugging Face14silentone0725 /ai-human-text-detection-v1 🧠 AI vs Human Text Detection Dataset (v1) This dataset merges nine major public and academic corpora to form one of the most comprehensive resources for AI-generated text detection model training and evaluation. 🔗 Sources The dataset consolidates, cleans, and standardizes multiple open datasets and research benchmarks, each focusing on human vs. AI-generated text classification: Hello-SimpleAI / HC3 — Human–ChatGPT comparison corpus gsingh1-py / train — Large-scale… See the full description on the dataset page: https://huggingface.co/datasets/silentone0725/ai-human-text-detection-v1.text10K<n<100K8 likes191 downloads11mo agoHugging Face15contralabs /HumanCreativityBenchmark The Human Creativity Benchmark (HCB) Expert evaluations of AI-generated creative work, built to separate two signals that single-score benchmarks collapse: convergence, where professionals align around shared, checkable standards, and divergence, where creative taste legitimately differs. Each AI output is judged by domain professionals through three complementary lenses — forced-choice pairwise comparisons, 1-5 scalar ratings on prompt adherence, usability, and visual appeal… See the full description on the dataset page: https://huggingface.co/datasets/contralabs/HumanCreativityBenchmark.imagetext-to-image1K<n<10K2 likes189 downloads3mo agoHugging Face16OliverUrbann /HumanoidRobotSoccer Fall Prediction Dataset for Humanoid Robots Dataset Summary This dataset consists of 37.9 hours of real-world sensor data collected from 20 Nao humanoid robots over the course of one year in various test environments, including RoboCup soccer matches. The dataset includes 18.3 hours of walking data, featuring 2519 falls. It captures a wide range of activities such as omni-directional walking, collisions, standing up, and falls on various surfaces like artificial turf and… See the full description on the dataset page: https://huggingface.co/datasets/OliverUrbann/HumanoidRobotSoccer.tabular10M<n<100M1 likes159 downloads1y agoHugging Face17Vchitect /VBench-I2V_human_annotationtextn<1K0 likes136 downloads4d agoHugging Face18jie-jw-wu /HumanEvalComm HumanEvalComm: Benchmarking the Communication Skills of Code Generation for LLMs and LLM Agent 📄 Paper • 💻 GitHub Repository • 🤗 Dataset Viewer Dataset Description HumanEvalComm is a benchmark dataset for evaluating the communication skills of Large Language Models (LLMs) in code generation tasks. It is built upon the widely used HumanEval benchmark. HumanEvalComm contains 762 modified problem descriptions based on the 164 problems in the… See the full description on the dataset page: https://huggingface.co/datasets/jie-jw-wu/HumanEvalComm.textn<1K0 likes134 downloads2y agoHugging Face19humanlong /emotion-negotiation-benchmarks Emotion-Aware LLM Negotiation Benchmarks Four high-stakes, edge-deployable negotiation benchmarks — the official evaluation suite for our research program on emotion-aware LLM agents. Each benchmark targets a distinct domain where (a) LLM-vs-LLM negotiation has real-world consequences, and (b) on-device deployment of small language models matters for privacy and latency. The benchmarks were originally introduced with EmoMAS (ACL 2026 Main, top 9% of 12,148 submissions) and are… See the full description on the dataset page: https://huggingface.co/datasets/humanlong/emotion-negotiation-benchmarks.tabulartext-generationn<1K0 likes130 downloads4mo agoHugging Face20Nima0Kamali /humancentric-scenes-ai HumanCentric-Scenes-AI A multimodal benchmark of 296 AI-generated human-centric scenes across four domains: CCTV / surveillance imagery (Set 2, 85 images). Midjourney-generated stills that mimic low-resolution security-camera footage — parking lots, building interiors, outdoor public spaces — designed to test whether detection cues survive heavy compression and low-light noise. Occupation × gender portraits (Set 3, 128 images). A balanced 64-occupation × 2-gender paired design… See the full description on the dataset page: https://huggingface.co/datasets/Nima0Kamali/humancentric-scenes-ai.imageimage-classificationn<1K2 likes116 downloads3d agoHugging Face21ICCIES-2025-DetectAI /vietnamese_news_human_ai Detecting AI-Generated Vietnamese News Articles with Multilingual-E5 and BERT This is the official dataset accompanying the paper Detecting AI-Generated Vietnamese News Articles with Multilingual-E5 and BERT, which was accepted at ICCIES 2025 and published in Computational Intelligence in Engineering Science (Springer CCIS, vol. 2587). You can read the paper here: Detecting AI-Generated Vietnamese News Articles with Multilingual-E5 and BERT Abstract The emergence… See the full description on the dataset page: https://huggingface.co/datasets/ICCIES-2025-DetectAI/vietnamese_news_human_ai.texttext-classification100K<n<1M5 likes105 downloads9mo agoHugging Face22rjmaftv33 /humancentric-scenes-ai HumanCentric-Scenes-AI A multimodal benchmark of 296 AI-generated human-centric scenes across four domains: CCTV / surveillance imagery (Set 2, 85 images). Midjourney-generated stills that mimic low-resolution security-camera footage — parking lots, building interiors, outdoor public spaces — designed to test whether detection cues survive heavy compression and low-light noise. Occupation × gender portraits (Set 3, 128 images). A balanced 64-occupation × 2-gender paired design… See the full description on the dataset page: https://huggingface.co/datasets/rjmaftv33/humancentric-scenes-ai.imageimage-classificationn<1K2 likes105 downloads2mo agoHugging Face23chillies /IELTS_essay_human_feedbacktext1K<n<10K21 likes96 downloads3y agoHugging Face24Shaow /humanbreast_xenium_janesick # Human Breast Cancer Xenium · Sample 1 Rep1+Rep2 Curated, ready-to-load spatial transcriptomics dataset. ## Source - Paper: [Janesick et al., Nat. Commun. 2023](https://www.nature.com/articles/s41467-023-43458-x) - Canonical download: cf.10xgenomics.com/samples/xenium/1.0.1/Xenium_FFPE_Human_Breast_Cancer_Rep{1,2} ## Scale | Property | Value | |---|---| | Technology | 10x Genomics Xenium (313-gene panel) | | Species | Homo sapiens | | Tissue |… See the full description on the dataset page: https://huggingface.co/datasets/Shaow/humanbreast_xenium_janesick.tabular100K<n<1M0 likes95 downloads5mo agoHugging Face25arcprize /arc_agi_2_human_testing ARC-AGI-2 Human testing data This file contains data from human testing sessions on ARC-AGI tasks. Each row represents a single test attempt by a human participant on a specific task-test pair in the "Public Train" or "Public Eval" ARC-AGI-2 datasets. Not all tasks in the released "Public Train" sets were tested, so these results are not comprehensive. This data does not include tasks from "Semi Private Evaluation" or "Private Evaluation" Column Descriptions… See the full description on the dataset page: https://huggingface.co/datasets/arcprize/arc_agi_2_human_testing.tabular1K<n<10K9 likes94 downloads1y agoHugging Face26RicardoRei /wmt-sqm-human-evaluation Dataset Summary In 2022, several changes were made to the annotation procedure used in the WMT Translation task. In contrast to the standard DA (sliding scale from 0-100) used in previous years, in 2022 annotators performed DA+SQM (Direct Assessment + Scalar Quality Metric). In DA+SQM, the annotators still provide a raw score between 0 and 100, but also are presented with seven labeled tick marks. DA+SQM helps to stabilize scores across annotators (as compared to DA). The data is… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-sqm-human-evaluation.tabular100K<n<1M1 likes92 downloads4y agoHugging Face27roboterradar /humanoid-robot-radarscore Roboterradar Humanoid & Quadruped Robot Dataset Curated editorial assessments of 21 commercially relevant humanoid robots (16) and quadruped robots (5), with a frozen scoring methodology, evidence grades and a complete source register. This Hugging Face repository is a versioned distribution mirror. The canonical, citeable publication is the Zenodo release: Version 1.0.0 DOI: https://doi.org/10.5281/zenodo.21797689 Concept DOI for all versions:… See the full description on the dataset page: https://huggingface.co/datasets/roboterradar/humanoid-robot-radarscore.tabularn<1K0 likes88 downloads2mo agoHugging Face28sparklessszzz /InstaArt-HumanAI Instagram AI Art vs Human Art: Engagement & Comment Dataset Dataset Summary This dataset was created and contributed by Akshaya, Cynthia, Grace, and Soham as part of a project at UC San Diego. This dataset supports research into how audiences engage with AI-generated art versus human-made art on Instagram, with a specific focus on comment sentiment, reaction types, and engagement patterns. It consists of 40 matched pairs of Instagram posts - one human art post and one… See the full description on the dataset page: https://huggingface.co/datasets/sparklessszzz/InstaArt-HumanAI.tabulartext-classificationn<1K1 likes83 downloads6mo agoHugging Face29AtlasBuiltIt /human-telemetry-driving-dataset-lite-version Dataset Card for 15 Laps of 30Hz NGSIM-Style Telemetry This is a Lite Version of a larger research dataset focusing on human driving signatures in high-fidelity simulations. It includes 15 full laps of telemetry captured at 30Hz within Unreal Engine 5, specifically formatted to match NGSIM standards. Dataset Details Dataset Description This Lite Version dataset contains 15 laps of high-fidelity human driving telemetry. It is intended for researchers and… See the full description on the dataset page: https://huggingface.co/datasets/AtlasBuiltIt/human-telemetry-driving-dataset-lite-version.tabulartabular-classification1K<n<10K0 likes82 downloads5mo agoHugging Face30izi-ano /CounselBench-Adv-human-annotationtext1K<n<10K0 likes78 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.