datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ai-text-detection-pile
Dataset Card for AI Text Dectection Pile
Dataset Summary
This is a large scale dataset intended for AI Text Detection tasks, geared toward long-form text and essays. It contains samples of both human text and AI-generated text from GPT2, GPT3, ChatGPT, GPTJ.
Here is the (tentative) breakdown:
Human Text
Dataset
Num Samples
Link
Reddit WritingPromps
570k
Link
OpenAI Webtext
260k
Link
HC3 (Human Responses)
58k
Link
ivypanda-essays
TODO
TODO… See the full description on the dataset page: https://huggingface.co/datasets/artem9k/ai-text-detection-pile.ai-human-text-detection-v1
🧠 AI vs Human Text Detection Dataset (v1)
This dataset merges nine major public and academic corpora to form one of the most comprehensive resources for AI-generated text detection model training and evaluation.
🔗 Sources
The dataset consolidates, cleans, and standardizes multiple open datasets and research benchmarks, each focusing on human vs. AI-generated text classification:
Hello-SimpleAI / HC3 — Human–ChatGPT comparison corpus
gsingh1-py / train — Large-scale… See the full description on the dataset page: https://huggingface.co/datasets/silentone0725/ai-human-text-detection-v1.ai-text-detection-trainingai-text-detection-pile-cleaned
AI Text Detection Pile - Cleaned Dataset
Dataset Description
This is a cleaned and processed version of the AI Text Detection Pile dataset, specifically optimized for training AI vs Human text classification models. The dataset has been carefully preprocessed to remove duplicates, filter by optimal text length, normalize encoding, and ensure balanced class distribution for robust model training.
Dataset Details
Total Samples: 721,626 (cleaned from original… See the full description on the dataset page: https://huggingface.co/datasets/srikanthgali/ai-text-detection-pile-cleaned.ai-text-detection-pile
Dataset Card for AI Text Dectection Pile
Dataset Summary
This is a large scale dataset intended for AI Text Detection tasks, geared toward long-form text and essays. It contains samples of both human text and AI-generated text from GPT2, GPT3, ChatGPT, GPTJ.
Here is the (tentative) breakdown:
Human Text
Dataset
Num Samples
Link
Reddit WritingPromps
570k
Link
OpenAI Webtext
260k
Link
HC3 (Human Responses)
58k
Link
ivypanda-essays
TODO
TODO… See the full description on the dataset page: https://huggingface.co/datasets/Mharis205/ai-text-detection-pile.ai-text-detection-pile
Dataset Card for AI Text Dectection Pile
Dataset Summary
This is a large scale dataset intended for AI Text Detection tasks, geared toward long-form text and essays. It contains samples of both human text and AI-generated text from GPT2, GPT3, ChatGPT, GPTJ.
Here is the (tentative) breakdown:
Human Text
Dataset
Num Samples
Link
Reddit WritingPromps
570k
Link
OpenAI Webtext
260k
Link
HC3 (Human Responses)
58k
Link
ivypanda-essays
TODO
TODO… See the full description on the dataset page: https://huggingface.co/datasets/kevknowscode/ai-text-detection-pile.ai-text-detection-benchmarkai-text-detection-datasetisogram-ai-text-detection-splits
Isogram AI Text Detection Permissive Splits
This dataset contains train/validation/test splits for binary AI-generated text detection.
It is built from sources whose dataset-level licenses were checked as permissive or
public-domain-compatible.
Schema
text: essay text.
label: 0 for human-written text, 1 for AI-generated text.
source_dataset: upstream dataset identifier.
source_detail: source label retained from the upstream data.
source_license: row-level… See the full description on the dataset page: https://huggingface.co/datasets/sinatras/isogram-ai-text-detection-splits.ai-text-detection-pile-cleaned
AI Text Detection Pile - Cleaned Dataset
Dataset Description
This is a cleaned and processed version of the AI Text Detection Pile dataset, specifically optimized for training AI vs Human text classification models. The dataset has been carefully preprocessed to remove duplicates, filter by optimal text length, normalize encoding, and ensure balanced class distribution for robust model training.
Dataset Details
Total Samples: 721,626 (cleaned from… See the full description on the dataset page: https://huggingface.co/datasets/R-obi/ai-text-detection-pile-cleaned.AITextDetectionDatasetai-text-detection-pile-cleaned
AI Text Detection Pile - Cleaned Dataset
Dataset Description
This is a cleaned and processed version of the AI Text Detection Pile dataset, specifically optimized for training AI vs Human text classification models. The dataset has been carefully preprocessed to remove duplicates, filter by optimal text length, normalize encoding, and ensure balanced class distribution for robust model training.
Dataset Details
Total Samples: 721,626 (cleaned from original… See the full description on the dataset page: https://huggingface.co/datasets/Mharis205/ai-text-detection-pile-cleaned.ai_text_detection_dataset_dl_hw_2_v6AI_text_detectionai-text-detection-progressivelonglamp-ai-detection
author-ai-text-detect/longlamp-ai-detection
A fork of LongLaMP/LongLaMP (user-setting configs) with model
completions attached. Every row is a LongLaMP sample with its original fields intact — the
author field, input, output and the full profile — plus a completions list holding
the model-written texts for that same task input.
The human reference is output; each entry of completions is machine-written. There is no
label column because the nesting already says which is which.… See the full description on the dataset page: https://huggingface.co/datasets/author-ai-text-detect/longlamp-ai-detection.
