CoolFace
Datasetpublic

agentlans/fake-wikipedia

Fake Wikipedia A synthetic and semi-synthetic dataset designed for studying hallucinations, omissions, and factual inconsistencies in language models. This dataset contains 10000 paragraphs derived from the agentlans/wikipedia-first-paragraph dataset, filtered for lengths between 1000 and 8000 characters. Using Qwen/Qwen3.5-4B, two distinct text variants were generated for each entry based solely on the article title: Fully Synthetic (fake): A completely hallucinated/made-up… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/fake-wikipedia.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes12downloads
settings

This repository belongs to agentlans on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namefake-wikipedia
visibilitypublic
licencecc-by-4.0
gatedno
owneragentlans
Account settings
agentlans/fake-wikipedia · CoolFace