agentlans/fake-wikipedia
Fake Wikipedia A synthetic and semi-synthetic dataset designed for studying hallucinations, omissions, and factual inconsistencies in language models. This dataset contains 10000 paragraphs derived from the agentlans/wikipedia-first-paragraph dataset, filtered for lengths between 1000 and 8000 characters. Using Qwen/Qwen3.5-4B, two distinct text variants were generated for each entry based solely on the article title: Fully Synthetic (fake): A completely hallucinated/made-up… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/fake-wikipedia.
This repository belongs to agentlans on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
