datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
eli5-human-vs-ai
ELI5 Human vs AI (long-form)
This dataset is for training and evaluating AI-writing detectors. It was built
as a clean way to compare known AI text against known human text: every human
answer predates ChatGPT by more than three years, so it is genuinely human by
construction, and every AI answer was written by a named 2026 model, so its origin
is certain too. Most detection datasets have to guess at their labels; this one
does not.
ELI5 answers were chosen because they are… See the full description on the dataset page: https://huggingface.co/datasets/mild-rgb/eli5-human-vs-ai.aita-human-vs-ai
AITA Human-vs-AI corpus (2026 generators)
A second human-vs-AI corpus, a companion to
mild-rgb/eli5-human-vs-ai,
in a deliberately different register: first-person judgment narratives from
r/AmItheAsshole, versus same-title posts written by seven 2026 models. 2,900
questions, one human post and one AI post each; ~414 documents per generator.
The human side is redacted — reconstruct it from Scruples
The human posts are verbatim r/AmItheAsshole text, obtained via… See the full description on the dataset page: https://huggingface.co/datasets/mild-rgb/aita-human-vs-ai.
