datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fake-real-newsfake-newsfake_news_corpus_spanish
Fake News Corpus Spanish
Citation
Gómez-Adorno, H., Posadas-Durán, J. P., Enguix, G. B., & Capetillo, C. P. (2021). Overview of FakeDeS at IberLEF 2021: Fake News Detection in Spanish Shared Task. Procesamiento del Lenguaje Natural, 67, 223-231.
Aragón, M. E., Jarquín, H., Gómez, M. M. Y., Escalante, H. J., Villaseñor-Pineda, L., Gómez-Adorno, H., ... & Posadas-Durán, J. P. (2020, September). Overview of mex-a3t at iberlef 2020: Fake news and aggressiveness analysis in… See the full description on the dataset page: https://huggingface.co/datasets/mariagrandury/fake_news_corpus_spanish.fake-news-detection-dataset-EnglishThis is a cleaned and splitted version of this dataset (https://www.kaggle.com/datasets/sadikaljarif/fake-news-detection-dataset-english)
Labels:
Fake News: 0
Real News: 1
You can find the cleansing script at: https://github.com/ErfanMoosaviMonazzah/Fake-News-Detection
fake_or_real_news
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Trinisha/fake_or_real_news.covid_fake_newsConstraint@AAAI2021 - COVID19 Fake News Detection in English
@misc{patwa2020fighting,
title={Fighting an Infodemic: COVID-19 Fake News Dataset},
author={Parth Patwa and Shivam Sharma and Srinivas PYKL and Vineeth Guptha and Gitanjali Kumari and Md Shad Akhtar and Asif Ekbal and Amitava Das and Tanmoy Chakraborty},
year={2020},
eprint={2011.03327},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
spanish-fake-news-fixed
Spanish Fake News Fixed
Este dataset contiene noticias etiquetadas en español, reparado para corregir saltos de línea internos.
FakeNewsSpanish_Kaggle2This dataset was obtained from: https://www.kaggle.com/datasets/zulanac/fake-and-real-news
Fake_and_Real_newsFakeNewsfake-newsfake-news-detector-euvsdisinfodata from https://euvsdisinfo.eu/
40-percent-cleaned-preprocessed-fake-real-newsKaggle based dataset for text classification task. The data has been cleaned and processed for preparation into any model for classification based tasks. This is just 40% of the entire dataset.
fake_news_en_opensources
Dataset Card for "Fake News Opensources"
Dataset Description
Homepage: https://github.com/AndyTheFactory/FakeNewsDataset
Repository: https://github.com/AndyTheFactory/FakeNewsDataset
Point of Contact: Andrei Paraschiv
Dataset Summary
a consolidated and cleaned up version of the opensources Fake News dataset
Fake News Corpus comprises 8,529,090 individual articles, classified into 12 classes: reliable, unreliable, political, bias, fake, conspiracy… See the full description on the dataset page: https://huggingface.co/datasets/andyP/fake_news_en_opensources.fake-news-detector-datasetPolitifact_fake_newscentral_de_fatos
Central de Fatos
Dataset Summary
In recent times, the interest for research dissecting the dissemination and prevention of misinformation in the online environment has spiked dramatically.
Given that scenario, a recurring obstacle is the unavailability of public datasets containing fact-checked instances.
In this work, we performed an extensive data collection of such instances from the better part of all major internationally recognized Brazilian fact-checking agencies.… See the full description on the dataset page: https://huggingface.co/datasets/fake-news-UFG/central_de_fatos.fake_news_combinedLabel Description
0 : Fake,
1 : Real
turkish-fake-news-detection
TR-FakeNews: Turkish Fake News Detection on Mainstream Media Dataset
This dataset contains 5325 news title and summaries related to significant events in Türkiye between 2015 and 2023.
Data Fields
title: a string format of the news headline.
description: a string format of the news summary.
status: a classification result 0 (fake) or 1 (real).
updated_log: Information about the data transformation process.
Resources: Indicates the source of the news.
Data Size… See the full description on the dataset page: https://huggingface.co/datasets/isakulaksiz/turkish-fake-news-detection.Thai-True-Fake-News
Thai Fake News Dataset
Language: Thai
Task: Fake News Classification
Dataset Description
This dataset contains news articles scraped from the Antifakenewscenter Thailand website using Selenium. It spans news published from 2017 to October 2024. The dataset is designed for fake news classification and consists of two main classes:
True News
Fake News
Each class contains 3002 samples, providing a balanced dataset for model training and evaluation.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/EXt1/Thai-True-Fake-News.turkish-fake-news-detection
MiDe22: An Annotated Multi-Event Tweet Dataset for Misinformation Detection
5,064 Turkish tweets with their misinformation labels for several recent events between 2020 and 2022, including the Russia-Ukraine war, COVID-19 pandemic, and Refugees. The dataset includes user engagements with the tweets in terms of likes, replies, retweets, and quotes. For user engagements please use contact at the end of the dataset card.
Data Fields
tweet: a string feature.
label: a… See the full description on the dataset page: https://huggingface.co/datasets/ogozcelik/turkish-fake-news-detection.fake-news-formated
Fake News Combined Dataset
This repository contains a single combined CSV of 165,620 news articles (real & fake), formatted for binary text‑classification tasks.
Each row has exactly five fields:
id,dataset_id,title,content,classification
id — 15‑character MD5 prefix (unique per article)
dataset_id — numeric source code (e.g. 2=WELFake, 5=CoAID, 6=Recovery, etc.)
title — headline or article title
content — full body text, sanitized to remove newlines
classification —… See the full description on the dataset page: https://huggingface.co/datasets/magnea/fake-news-formated.fake_real_newsFakeNewsNetFake-News-ClassificationDevelop a machine learning program to identify when an article might be fake news. Run by the UTK Machine Learning Club.
This is the Dataset to the Fake-News-Classifier competition in Kaggle. There is a Test csv to check for predictions.
Citation
William Lifferth. (2018). Fake News. Kaggle. https://kaggle.com/competitions/fake-news
english-fake-news-detection
MiDe22: An Annotated Multi-Event Tweet Dataset for Misinformation Detection
5,284 English tweets with their misinformation labels for several recent events between 2020 and 2022, including the Russia-Ukraine war, COVID-19 pandemic, and Refugees. The dataset also includes user engagements with the tweets in terms of likes, replies, retweets, and quotes. For user engagements please use contact at the end of the dataset card.
Data Fields
tweet: a string feature.
label: a… See the full description on the dataset page: https://huggingface.co/datasets/ogozcelik/english-fake-news-detection.Fake_News_Detection_System_29.5kfake-newsfake_news_testcleaned_fake_or_real_news
