datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
COLING-2025-GENAI-3
🚨 RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors 🚨
🌐 Website, 🖥️ Github, 📝 Paper
RAID is the largest & most comprehensive dataset for evaluating AI-generated text detectors.
It contains over 10 million documents spanning 11 LLMs, 11 genres, 4 decoding strategies, and 12 adversarial attacks.
It is designed to be the go-to location for trustworthy third-party evaluation of both open-source and closed-source generated text detectors.
Load… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/COLING-2025-GENAI-3.NIST-GenAI-Profile
NIST Generative AI Profile Question Answering Dataset
Dataset Summary
This dataset contains question-and-answer records derived from the National Institute of Standards and Technology publication NIST AI 600-1, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. The source publication is a cross-sectoral companion resource to the NIST AI Risk Management Framework and focuses on risks that are unique to, or exacerbated by… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/NIST-GenAI-Profile.facts-grounding-processed
Dataset Summary
The dataset contains prompts, context documents, and target answers that challenge models to stay grounded in provided context rather than hallucinating.Processing steps added extra features like:
prompt – consolidated instruction + user request + context
has_url_in_context – boolean flag for URLs in context
len_system, len_user, len_context – token/word length statistics
row_id – unique identifier for tracking
Dataset Structure
Splits:
train – 688… See the full description on the dataset page: https://huggingface.co/datasets/GenAIDevTOProd/facts-grounding-processed.GenAIDetectHU
HuDetect-1k — a Hungarian AI-text detection benchmark (v1.3)
1,000 Hungarian documents, 500 written by people and 500 generated by large
language models, for evaluating AI-text detectors on Hungarian. Every text is
wholly human or wholly machine-written.
Developed by https://preds.hu
The Preds AI-detector got the best results for hungarian language: Preds
File
hudetect1k_v1.3.csv — UTF-8, RFC 4180 quoting, header + 1,000 rows.
column
content
text
the… See the full description on the dataset page: https://huggingface.co/datasets/mihalykiss/GenAIDetectHU.GenAi-Public-Responsegenai-fraud-benchmark-2026
📊 Aegis-Graph: 2026 GenAI Fraud Benchmark
Official high-fidelity dataset for academic integrity verification. Managed by ACLAS.
Official Website: aclas.college
🏛️ Institutional Grounding
Aegis-Graph is a sovereign protocol engineered and governed by Atlanta College of Liberal Arts and Sciences (ACLAS).
Official Website: aclas.college
Keywords: ACLAS, ACLAS College, Atlanta College of Liberal Arts and Sciences, Academic Integrity, Sovereign AI.
GenAIHealthCare
GenAIHealthCare
tags: EHR, Diagnosis, Prediction
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'GenAIHealthCare' dataset is a curated collection of Electronic Health Record (EHR) entries designed to support machine learning models focused on improving diagnostic prediction in the healthcare domain. Each entry in the dataset comprises a patient record that includes historical health data, symptoms, diagnostic codes, and… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/GenAIHealthCare.genai-terminology-en-ja生成AIの日英専門用語集です。正確さは保証しませんが、GPT-4などの頭に入れておくと綺麗に訳せると思います。
SC-train-valid-test_SDG-Descriptions-GenAI-01genai-tools-platforms-data
🧠 Generative AI Tools & Platforms (2025)
Author: Tarek Masryo · KaggleLicense: CC BY 4.0 (Attribution)
A clean, structured catalog of 113 Generative AI tools and platforms released up to 2025.
Each row represents one tool/platform and includes:
Vendor/maintainer
Canonical category and primary modality
API availability and status
Open-source indicator
Release year + derived timeline features
Modality flags and breadth count
Designed for ecosystem mapping, benchmarking, and… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/genai-tools-platforms-data.SFT_Science_AI_Genskill_embeddingsfinetuneCOLING-2025-GENAI-MULTIGenAI4ESG_Sample_Datasetmed_gen_aiGenAI4ESG_Sample_DatasetGenAI_final_projectMastitis_raw_QA_gen_chatgpt4oCOLING-2025-GENAI-MONOgenai-training-datasetdatasetCantilevergenaidesign/Cantilever
my-genai-datasetGenAI_HW1SC-train-valid-test_SDG-Papers-GenAI-01IT_terms_AI_genDL-gen-aiCantileverCSV''
Cantilever9
