datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HumanVsAICode
Human vs. AI-Generated Code
Dataset Summary
This dataset is a large-scale collection of human-written and LLM-generated code designed to study differences in defect distribution, code quality, and security characteristics between human developers and modern AI code assistants.
It contains paired implementations of the same function across multiple authorship sources, spanning Python and Java, two widely adopted programming languages with distinct typing systems, paradigms… See the full description on the dataset page: https://huggingface.co/datasets/OSS-forge/HumanVsAICode.eli5-human-vs-ai
ELI5 Human vs AI (long-form)
This dataset is for training and evaluating AI-writing detectors. It was built
as a clean way to compare known AI text against known human text: every human
answer predates ChatGPT by more than three years, so it is genuinely human by
construction, and every AI answer was written by a named 2026 model, so its origin
is certain too. Most detection datasets have to guess at their labels; this one
does not.
ELI5 answers were chosen because they are… See the full description on the dataset page: https://huggingface.co/datasets/mild-rgb/eli5-human-vs-ai.aita-human-vs-ai
AITA Human-vs-AI corpus (2026 generators)
A second human-vs-AI corpus, a companion to
mild-rgb/eli5-human-vs-ai,
in a deliberately different register: first-person judgment narratives from
r/AmItheAsshole, versus same-title posts written by seven 2026 models. 2,900
questions, one human post and one AI post each; ~414 documents per generator.
The human side is redacted — reconstruct it from Scruples
The human posts are verbatim r/AmItheAsshole text, obtained via… See the full description on the dataset page: https://huggingface.co/datasets/mild-rgb/aita-human-vs-ai.myanmar_quran_parallel_dataset_human_vs_ai
Myanmar Quran Parallel Dataset: Human vs AI
This dataset is a comprehensive multi-parallel corpus of the Holy Qur'an, containing all 6,236 verses.
It is designed as a high-quality linguistic resource for evaluating and aligning AI systems on formal, literary, and modern Myanmar (Burmese) language in a religious context.
Each verse aligns the original Uthmani Arabic text with trusted human translations and multiple AI-generated translations, enabling fine-grained comparison between… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_quran_parallel_dataset_human_vs_ai.human-vs-ai-spanish-65kkaz-lang-human-vs-ai
