CoolFace
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01humanfia-lab /QIT QIT Humanize-Physic Formalizations and Proofs QIT (Quantum Information Theory) is a blind benchmark for formalizing theorems in quantum information. It evaluates whether an AI agent can faithfully translate natural-language and TeX problem statements into Lean 4 theorems and then construct formal proofs checked by the Lean kernel. Its 40 tasks cover quantum channels and Choi representations, entropy and coding, mixed-unitary obstructions and symmetry, norm and fidelity tools… See the full description on the dataset page: https://huggingface.co/datasets/humanfia-lab/QIT.tabulartext-generationn<1K0 likes840 downloads2mo agoHugging Face02dsfox /humans-top humans.top — LIVE Global ranking of influential people (open dataset) This dataset ranks real, named living people by global influence — e.g. #1 Donald Trump, #2 Xi Jinping, #3 Vladimir Putin, alongside figures like Elon Musk, Narendra Modi and Lionel Messi. Every row is a person: their live influence rank, a concise biography in 15 languages, and Wikidata / Wikipedia links. Published from the website humans.top (.top is the domain name). Available on (identical CC0… See the full description on the dataset page: https://huggingface.co/datasets/dsfox/humans-top.imagetext-generationn<1K1 likes776 downloads13h agoHugging Face03contralabs /HumanCreativityBenchmark The Human Creativity Benchmark (HCB) Expert evaluations of AI-generated creative work, built to separate two signals that single-score benchmarks collapse: convergence, where professionals align around shared, checkable standards, and divergence, where creative taste legitimately differs. Each AI output is judged by domain professionals through three complementary lenses — forced-choice pairwise comparisons, 1-5 scalar ratings on prompt adherence, usability, and visual appeal… See the full description on the dataset page: https://huggingface.co/datasets/contralabs/HumanCreativityBenchmark.imagetext-to-image1K<n<10K2 likes211 downloads3mo agoHugging Face04humanlong /emotion-negotiation-benchmarks Emotion-Aware LLM Negotiation Benchmarks Four high-stakes, edge-deployable negotiation benchmarks — the official evaluation suite for our research program on emotion-aware LLM agents. Each benchmark targets a distinct domain where (a) LLM-vs-LLM negotiation has real-world consequences, and (b) on-device deployment of small language models matters for privacy and latency. The benchmarks were originally introduced with EmoMAS (ACL 2026 Main, top 9% of 12,148 submissions) and are… See the full description on the dataset page: https://huggingface.co/datasets/humanlong/emotion-negotiation-benchmarks.tabulartext-generationn<1K0 likes140 downloads4mo agoHugging Face05sorry-bench /sorry-bench-human-judgment-202406gated Dataset Card for 🧑‍⚖️SORRY-Bench Human Judgment Dataset (2024/06) 🏠Website 📑Paper 📚Dataset 💻Github 🧑‍⚖️Human Judgment Dataset 🤖Judge LLM This dataset contains 7.2K annotations of human safety judgmentsfor LLM responses to unsafe instructions of our SORRY-Bench dataset. Specifically, for each unsafe instruction of the 450 unsafe instructions in SORRY-Bench dataset, we annotate 16 diverse model responses (both ID and OOD) as either in… See the full description on the dataset page: https://huggingface.co/datasets/sorry-bench/sorry-bench-human-judgment-202406.tabulartext-classification1K<n<10K5 likes112 downloads2y agoHugging Face06sorry-bench /sorry-bench-human-judgment-202503gated Dataset Card for 🧑‍⚖️SORRY-Bench Human Judgment Dataset (2025/03) 🏠Website 📑Paper 📚Dataset 💻Github 🧑‍⚖️Human Judgment Dataset 🤖Judge LLM 🪧UPDATE: In this iteration, we removed the category "Impersonation" due to its ambiguous definition, and that most models more or less fulfill such requests.This dataset contains 7K annotations of human safety judgments for LLM responses to unsafe instructions of our SORRY-Bench dataset. Specifically, for… See the full description on the dataset page: https://huggingface.co/datasets/sorry-bench/sorry-bench-human-judgment-202503.tabulartext-classification1K<n<10K1 likes99 downloads2y agoHugging Face07gohumanize /gohumanize-open-humanizer-dataset GoHumanize Open Humanizer Dataset 2,957 training pairs and 300 test pairs for teaching a language model to rewrite AI-styled English prose into natural human writing. Each pair is: input: a passage rewritten by a large language model in the register typical of LLM output (formal, smooth, hedged, connective phrases, no contractions); output: the original human-written passage, from a public-domain book or, since version 2, from a US federal government publication. The human… See the full description on the dataset page: https://huggingface.co/datasets/gohumanize/gohumanize-open-humanizer-dataset.tabulartext-generation1K<n<10K0 likes74 downloads9h agoHugging Face08rasbt /human-writing-prompts-6k Human Writing Prompts 6K This dataset contains 6,500 unique English writing prompts for experiments on human-style text generation. The prompts were constructed from broad topics extracted from human-written source texts. The source texts themselves are not included. Split Prompts Train 5,000 Validation 500 Test 1,000 The source-document groups do not cross split boundaries. Each row retains the source collection, document, URL, and license metadata of the… See the full description on the dataset page: https://huggingface.co/datasets/rasbt/human-writing-prompts-6k.tabulartext-generation1K<n<10K0 likes64 downloads1mo agoHugging Face09ephipi /human-ai-parallel-detection Dataset Card for human-ai-parallel-detection Dataset Description Dataset Summary The human-ai-parallel-detection dataset contains 600 balanced instances for evaluating methods to distinguish between human-written and AI-generated text continuations. Each instance includes a 500-word human-written prompt followed by parallel continuations from humans, GPT-4o, and LLaMA-70B-Instruct. The dataset includes both style embedding features and LLM-as-judge predictions… See the full description on the dataset page: https://huggingface.co/datasets/ephipi/human-ai-parallel-detection.tabulartext-classificationn<1K1 likes48 downloads1y agoHugging Face10Experimental-Orange /HumanAgencyBench_Evaluation_Results HumanAgencyBench evaluation results Paper: HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants Code: https://github.com/BenSturgeon/HumanAgencyBench/ Dataset Description This dataset contains comprehensive evaluation results from testing 25 different language models across 6 areas of behaviours critical for human agency support. Each model was evaluated on 3,000 prompts (500 per category), resulting in 75,000 total evaluations designed… See the full description on the dataset page: https://huggingface.co/datasets/Experimental-Orange/HumanAgencyBench_Evaluation_Results.tabulartext-generation10K<n<100K0 likes47 downloads25d agoHugging Face11nics-efc /MoA_Long_HumanQA MoA: Mixture of Sparse Attention for Automatic Large Language Model Compression This is the dataset used by the automatic sparse attention compression method MoA. It enhances the calibration dataset by integrating long-range dependencies and model alignment. MoA utilizes long-contextual datasets, which include question-answer pairs heavily dependent on long-range content. The question-answer pairs are written by human in this dataset repository. Large language Models (LLMs) should… See the full description on the dataset page: https://huggingface.co/datasets/nics-efc/MoA_Long_HumanQA.tabularquestion-answering1K<n<10K4 likes44 downloads2y agoHugging Face12ababa134 /fuzzeval-humaneval-mbpp FuzzEval unit tests for HumanEval-f and MBPP-f Automatically generated unit tests for a reproduction of the ICML 2026 paper "Towards Functional Correctness of Large Code Models with Selective Generation" (Jeong, Kim & Park — arXiv:2505.13553, official repo trustml-lab/selective-code-generation). The paper's FuzzEval paradigm replaces a benchmark's handful of hand-written unit tests with hundreds of unit tests obtained by fuzzing the reference solution. This dataset is our… See the full description on the dataset page: https://huggingface.co/datasets/ababa134/fuzzeval-humaneval-mbpp.tabulartext-generationn<1K0 likes36 downloads2mo agoHugging Face13HumanEdgeAI /LegalReasoning Human Edge — Legal Reasoning Evaluation (Showcase Sample) A public 6-task sample from a rubric-based legal reasoning evaluation dataset built by Human Edge (humanedgetech.ai). Each task is authored and reviewed by practicing senior lawyers and is designed to produce a verifiable, per-criterion reward signal for post-training and evaluation of frontier language models on high-complexity legal work. The sample contains one task per legal subdomain, drawn from a larger internal… See the full description on the dataset page: https://huggingface.co/datasets/HumanEdgeAI/LegalReasoning.tabulartext-generationn<1K0 likes32 downloads2mo agoHugging Face14NordosoftOy /innoduel-rlhf-real-world-human-preferences-sample Real-World Human Pairwise Preferences — Public Sample 📦 This is a free, public sample of a commercial dataset. It contains 1,350 rows curated for inspection. The full dataset has 1.5 million human pairwise-preference decisions. Full dataset: https://huggingface.co/datasets/NordosoftOy/innoduel-rlhf Request access / licensing: see § Access to the full dataset — contact kari.nieminen@nordo.fi. Use this sample to evaluate the data's quality, structure and… See the full description on the dataset page: https://huggingface.co/datasets/NordosoftOy/innoduel-rlhf-real-world-human-preferences-sample.tabulartext-generation1K<n<10K0 likes31 downloads1mo agoHugging Face15Not-Humanity-Exam /Imprint-Train-v3tabulartext-classification100K<n<1M0 likes30 downloads7mo agoHugging Face16Not-Humanity-Exam /Imprint-Train-v2tabulartext-classification10K<n<100K0 likes18 downloads7mo agoHugging Face17Not-Humanity-Exam /Imprint-Train-v1tabulartext-classification1K<n<10K0 likes16 downloads7mo agoHugging Face18Alberto1231 /human_study human_study Free-form human Reddit responses derived from snap-stanford/user_study_annotations. The original Reddit post text comes from the HumanLM authors' reddit_post_dict_testset.json. Each row contains an original Reddit post in prompt and the response written by a human-study participant in target. Model responses, generated personas, comparison judgments, and worker identifiers are intentionally excluded. Deduplication and splits Source annotation… See the full description on the dataset page: https://huggingface.co/datasets/Alberto1231/human_study.tabulartext-generationn<1K0 likes14 downloads2mo agoHugging Face19harari /human_grch38_segment_sample Human GRCh38 Genome Segments Dataset Description This dataset contains 16,384 base pair segments from the human reference genome (GRCh38) prepared for Sparse Autoencoder (SAE) training with the Evo2 model. The segments are extracted using a sliding window approach with 75% overlap. Dataset Details Total segments: 718,648 Segment size: 16,384 base pairs Stride: 4,096 base pairs (75% overlap) Source genome: GRCh38.primary_assembly (GENCODE Release 41)… See the full description on the dataset page: https://huggingface.co/datasets/harari/human_grch38_segment_sample.tabulartext-generation1K<n<10K0 likes9 downloads1y agoHugging Face20enescingoz /humaneval-apple-silicon Mac Coding Bench Results v1 — Speed + Code Quality Benchmarks on Apple Silicon Speed and code quality benchmarks for quantized LLMs running locally on Apple Silicon Macs. The dataset pairs inference speed measurements (tokens/sec) with HumanEval+ functional correctness scores for 21 models, across three hardware configurations (M1, M2 Max, M5) totaling 123 benchmark results. Key Highlights Qwen 3.6 35B-A3B achieves 89.6% HumanEval+ pass@1 at 16.7 tok/s — best quality… See the full description on the dataset page: https://huggingface.co/datasets/enescingoz/humaneval-apple-silicon.tabulartext-generationn<1K0 likes8 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.