datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Ivy-Fake
IVY-FAKE: Unified Explainable Benchmark and Detector for AIGC Content
This repository provides the official implementation of IVY-FAKE and IVY-xDETECTOR, a unified explainable framework and benchmark for detecting AI-generated content (AIGC) across both images and videos.
🔍 Overview
IVY-FAKE is the first large-scale dataset designed for multimodal explainable AIGC detection. It contains:
150K+ training samples (images + videos)
18.7K evaluation samples
Fine-grained… See the full description on the dataset page: https://huggingface.co/datasets/Julien0123/Ivy-Fake.Ivy-Fake
IVY-FAKE: Unified Explainable Benchmark and Detector for AIGC Content
This repository provides the official implementation of IVY-FAKE and IVY-xDETECTOR, a unified explainable framework and benchmark for detecting AI-generated content (AIGC) across both images and videos.
🔍 Overview
IVY-FAKE is the first large-scale dataset designed for multimodal explainable AIGC detection. It contains:
150K+ training samples (images + videos)
18.7K evaluation samples
Fine-grained… See the full description on the dataset page: https://huggingface.co/datasets/yishiliu/Ivy-Fake.Tiefighter-13B-Fake-Distill-ShareGPTfake_tech_companies_market_reportsbcms-fake-news-articlesArabic_fake_news_dataset
Arabic_fake_news_dataset
Please note that this dataset needs more preprocessing.
Introduction
This repository contains the Arabic_fake_news_dataset, a collection of news articles scraped from the Egyptian platform متصدقش (Matsda2sh). The dataset is intended for studying and addressing the spread of fake news within the Egyptian community. It includes news articles classified as either fake or true, along with their corresponding titles.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/Arabic_fake_news_dataset.fake-news-detection-instructionHuman-fakebench-judgment-sonicFakeProfile600Excited to share 600 state-of-art profiles of fake voice profiles of real people. We are releasing this for training robust and diverse text-to-speech and text-to-music models.
Data fields
displayName:Full name of the speaker (e.g., "Éloïse Gagné").
language:Language code in all caps with underscore (e.g., EN_US).
locale:Regional locale code using ISO format (e.g., fr-CA for French, Canada).
gender:Gender of the speaker (e.g., "female").
imageUrl:URL to the speaker’s image/avatar.… See the full description on the dataset page: https://huggingface.co/datasets/sleeping-ai/FakeProfile600.FakeRomFakeRO_updatedfake-clinical-recordsSarav-real-fake-automobile-parts-datasetfake-wikipedia
Fake Wikipedia
A synthetic and semi-synthetic dataset designed for studying hallucinations, omissions, and factual inconsistencies in language models.
This dataset contains 10000 paragraphs derived from the agentlans/wikipedia-first-paragraph dataset, filtered for lengths between 1000 and 8000 characters.
Using Qwen/Qwen3.5-4B, two distinct text variants were generated for each entry based solely on the article title:
Fully Synthetic (fake): A completely hallucinated/made-up… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/fake-wikipedia.Fake_Testing_dataa_fake_day_in_the_life_cycleFakerom_updated_originalFakeNewsmoh_8_fake_rollouts
MOH-8 Fake Rollouts
480 math competition problems, each with 8 candidate solution rollouts from OSS 120B.
A controlled number of rollouts per problem are correct — use this to train/test a
verifier model that must identify which solutions are right.
Source
Problems and rollouts sampled from aimosprite/training-data-oss120b (the oss128-fixed-FINAL.jsonl file).
Only polymath-source problems in the 2/8–4/8 pass rate range (32–64 correct out of 128 attempts).
4 problems… See the full description on the dataset page: https://huggingface.co/datasets/aimosprite/moh_8_fake_rollouts.fake_knowledge使用 Gemini 生成的一个虚假知识数据库。用来测试大模型外挂数据库时是否真的从外挂的数据库中吸取了知识。
fakesv-qwen3fake_name_and_ssnFake_WiLlamafaker-step1-child0faker-step2-child0synthetic-bank-calls-1000-fakerfake_name_and_ssnfaker-step1-mastermock_faker_datasetFakeProfiles.v.2.0
