CoolFace
Datasetpublic

agentlans/fake-wikipedia

Fake Wikipedia A synthetic and semi-synthetic dataset designed for studying hallucinations, omissions, and factual inconsistencies in language models. This dataset contains 10000 paragraphs derived from the agentlans/wikipedia-first-paragraph dataset, filtered for lengths between 1000 and 8000 characters. Using Qwen/Qwen3.5-4B, two distinct text variants were generated for each entry based solely on the article title: Fully Synthetic (fake): A completely hallucinated/made-up… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/fake-wikipedia.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes12downloads

No commit history came back for main. The revision may not exist, or the source declined the request.