datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
human_ai_generated_text
Human or AI-Generated Text
The data can be valuable for educators, policymakers, and researchers interested in the evolving education landscape, particularly in detecting or identifying texts written by Humans or Artificial Intelligence systems.
File Name
model_training_dataset.csv
File Structure
id: Unique identifier for each record.
human_text: Human-written content.
ai_text: AI-generated texts.
instructions: Description of the task given to both Humans and… See the full description on the dataset page: https://huggingface.co/datasets/dmitva/human_ai_generated_text.AI-and-Human-Generated-Text
AI & Human Generated Text
I am Using this dataset for AI Text Detection for https://exnrt.com.
Check Original DataSet GitHub Repository Here: https://github.com/panagiotisanagnostou/AI-GA
Description
The AI-GA dataset, short for Artificial Intelligence Generated Abstracts, comprises abstracts and titles. Half of these abstracts are generated by AI, while the remaining half are original. Primarily intended for research and experimentation in natural language… See the full description on the dataset page: https://huggingface.co/datasets/Ateeqq/AI-and-Human-Generated-Text.AI-human-textThis is a processed dataset of Human vs AI Text roughly 400k rows. This is taken from the Kaggle dataset https://www.kaggle.com/datasets/shanegerami/ai-vs-human-text/data then processed and split into training and test sets.
ai-human-text-detection-v1
🧠 AI vs Human Text Detection Dataset (v1)
This dataset merges nine major public and academic corpora to form one of the most comprehensive resources for AI-generated text detection model training and evaluation.
🔗 Sources
The dataset consolidates, cleans, and standardizes multiple open datasets and research benchmarks, each focusing on human vs. AI-generated text classification:
Hello-SimpleAI / HC3 — Human–ChatGPT comparison corpus
gsingh1-py / train — Large-scale… See the full description on the dataset page: https://huggingface.co/datasets/silentone0725/ai-human-text-detection-v1.multi_domain_ai_human_text
multi_domain_ai_human_text — Datasheet
Balanced, multi-domain AI-vs-human text detection benchmark with dedicated
out-of-distribution and adversarial evaluation panels. Built by
scripts/build_paper_dataset.py from an 11-corpus unified aggregation.
Splits
Split
AI
Human
Total
Purpose
train
300,000
300,000
600,000
training (balanced, English, clean)
validation
2,996
2,999
5,995
model selection
test
4,991
4,999
9,990
in-distribution test… See the full description on the dataset page: https://huggingface.co/datasets/acmc/multi_domain_ai_human_text.human-ai-generated-text
Dataset Card for human-ai-generated-text
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/ardavey/human-ai-generated-text/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/ardavey/human-ai-generated-text.human_ai_generated_text
Human or AI-Generated Text
The data can be valuable for educators, policymakers, and researchers interested in the evolving education landscape, particularly in detecting or identifying texts written by Humans or Artificial Intelligence systems.
File Name
model_training_dataset.csv
File Structure
id: Unique identifier for each record.
human_text: Human-written content.
ai_text: AI-generated texts.
instructions: Description of the task given to both… See the full description on the dataset page: https://huggingface.co/datasets/jist2009/human_ai_generated_text.AI-human-textThis is a processed dataset of Human vs AI Text roughly 400k rows. This is taken from the Kaggle dataset https://www.kaggle.com/datasets/shanegerami/ai-vs-human-text/data then processed and split into training and test sets.
human_ai_text_classification
Human AI Text Classification
Dataset Summary
human_ai_text_classification is a binary text classification dataset for distinguishing human-written text from AI-generated text.
It was created by combining three public datasets, standardizing them into a common schema, balancing the class labels, removing duplicate texts, and performing a stratified 80/20 train-test split.
Labels:
0 = human-written text
1 = AI-generated text
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/inokusan/human_ai_text_classification.ru_kz_ai_human_textsai-human-text-classification
AI vs Human Sentence Classification Dataset
Dataset Summary
sentence_dataset is a sentence-level binary classification dataset containing approximately 9.84 million sentences labelled as either AI-generated (1) or human-written (0).
It was constructed by extracting individual sentences from two source datasets and merging them:
Dataset 1 — ai_vs_human_content_v2_20000.csv: 20,000 rows of short text and code snippets with rich metadata (prompt, topic, source… See the full description on the dataset page: https://huggingface.co/datasets/Nerdy37/ai-human-text-classification.sentence-level-ai-human-textai-vs-real-text-generated-datasetArabic-English-Turkish-AI-Human-Textai_and_human_textai-vs-human-textsAI-human-textThis is a processed dataset of Human vs AI Text roughly 400k rows. This is taken from the Kaggle dataset https://www.kaggle.com/datasets/shanegerami/ai-vs-human-text/data then processed and split into training and test sets.
ru_kz_ai_human_texts_v2ai_human_textscombined_ai_human_arabic_text_datasetai_human_textRahul_AI_Human_Texthuman_and_ai_textai-and-human-text-classificationAI_and_human_textsdlr-hw-2-human-ai-texts
