datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
best-ai-humanizer-independent-benchmark
AI Humanizer Benchmark: Rankings & Public Audit Record
The best AI humanizers, ranked by a monthly benchmark. This benchmark measures how well each AI humanizer bypasses the major AI detectors (GPTZero, Originality.ai, Copyleaks, Winston AI, and ZeroGPT) while preserving the original meaning and readability. We pay for every tool ourselves and run each one by hand on the most undetectable setting it advertises; there are no affiliate deals and no vendor-supplied numbers. Every… See the full description on the dataset page: https://huggingface.co/datasets/AIHumanizerBenchmarks/best-ai-humanizer-independent-benchmark.ai-humanizer-benchmark
AI Humanizer Benchmark — monthly cycle data
The complete raw data of AI Humanizer Benchmark, a monthly measured benchmark of AI humanizers. Every tool rewrites the same 33 freshly generated texts on its default settings; every output is scored by 7 commercial AI detectors (GPTZero, Originality.ai, Copyleaks, Winston AI, ZeroGPT, QuillBot, Grammarly) plus meaning preservation and readability.
This dataset is the official mirror of the GitHub data repository, published by the AI… See the full description on the dataset page: https://huggingface.co/datasets/ai-humanizer-benchmark/ai-humanizer-benchmark.humanizerbench
HumanizerBench: AI humanizer rankings and public audit record
The complete audit record of HumanizerBench, a monthly benchmark of AI humanizers. Every tool rewrites the same freshly generated texts on the most undetectable setting it advertises, and every output is scored by five commercial AI detectors alongside meaning preservation and readability. We pay for every tool ourselves. There are no affiliate deals and no vendor-supplied numbers.
Every input, every humanized output… See the full description on the dataset page: https://huggingface.co/datasets/HumanizerBench/humanizerbench.gohumanize-open-humanizer-dataset
GoHumanize Open Humanizer Dataset
2,957 training pairs and 300 test pairs for teaching a language model to rewrite
AI-styled English prose into natural human writing. Each pair is:
input: a passage rewritten by a large language model in the register typical of LLM output
(formal, smooth, hedged, connective phrases, no contractions);
output: the original human-written passage, from a public-domain book or, since version 2,
from a US federal government publication.
The human… See the full description on the dataset page: https://huggingface.co/datasets/gohumanize/gohumanize-open-humanizer-dataset.sft-humanizer-dataset-v4yi-humanizer-v19-test-100humanize-rl-v03yi-humanizer-dpo-v16-pairsyi-humanizer-v18-samples-100
Yi Humanizer v18 — 100 samples
100 humanized text samples generated by SwaYHell/yi-humanizer-v18-merged-v11-r8
via vLLM batch inference (double-merged: Yi + v11 + v18).
Generation params
Base: 01-ai/Yi-1.5-9B + v11 LoRA (merged) + v18 LoRA (merged)
Temperature: 1.0
Input length range: 300–800 words
N samples: 100
Schema (JSONL)
i: index
input: original AI text
output: humanized version
wi, wo: input/output word counts
ratio: wo/wi
temp: generation… See the full description on the dataset page: https://huggingface.co/datasets/SwaYHell/yi-humanizer-v18-samples-100.yi-humanizer-v18-full-pipeline-100yi-humanizer-v18-paraphrase-100yi-humanizer-v15-samples-100
Yi Humanizer v15 — 100 samples
100 humanized text samples generated by SwaYHell/yi-humanizer-v15-no-citations
via vLLM batch inference.
Generation params
Base model: 01-ai/Yi-1.5-9B + LoRA (merged for vLLM)
Temperature: 1.3
Input length range: 300–800 words
N samples: 100
Schema (JSONL)
i: index
input: original AI text
output: humanized version
wi, wo: input/output word counts
ratio: wo/wi
temp: generation temperature
model: LoRA model name
yi-humanizer-v18-AWQ-test-100
