datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fake_news
TODO: Add YAML tags here. Copy-paste the tags obtained with the online tagging app: https://huggingface.co/spaces/huggingface/datasets-tagging
annotations_creators:
- no-annotation
language_creators:
- found
language:
- en
license:
- unknown
multilinguality:
- monolingual
size_categories:
- 30k<n<50k
source_datasets:
- original
task_categories:
- text-classification
task_ids:
- fact-checking
- intent-classification
pretty_name: GonzaloA / Fake News
Dataset Card for… See the full description on the dataset page: https://huggingface.co/datasets/GonzaloA/fake_news.fake_news_filipino Low-Resource Fake News Detection Corpora in Filipino. The first of its kind. Contains 3,206 expertly-labeled news samples, half of which are real and half of which are fake.fake-real-newsFakeNewsCorpusSpanish
:newspaper: The Spanish Fake News Corpus
The Spanish Fake News Corpus Version 2.0 [[ FakeDeS Task @ Iberlef 2021 ]] :metal:
Corpus Description
The Spanish Fake News Corpus Version 2.0 contains pairs of fake and true publications about different events (all of them were written in Spanish) that were collected from November 2020 to March 2021. Different sources from the web were used to gather the information, but mainly of two types: 1) newspapers and media… See the full description on the dataset page: https://huggingface.co/datasets/sayalaruano/FakeNewsCorpusSpanish.fake-newsfake_news_english
Dataset Card for Fake News English
Dataset Summary
This dataset contains URLs of news articles classified as either fake or satire. The articles classified as fake also have the URL of a rebutting article.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
English
Dataset Structure
Data Instances
{
"article_number": 102 ,
"url_of_article":… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/fake_news_english.fake_news_corpus_spanish
Fake News Corpus Spanish
Citation
Gómez-Adorno, H., Posadas-Durán, J. P., Enguix, G. B., & Capetillo, C. P. (2021). Overview of FakeDeS at IberLEF 2021: Fake News Detection in Spanish Shared Task. Procesamiento del Lenguaje Natural, 67, 223-231.
Aragón, M. E., Jarquín, H., Gómez, M. M. Y., Escalante, H. J., Villaseñor-Pineda, L., Gómez-Adorno, H., ... & Posadas-Durán, J. P. (2020, September). Overview of mex-a3t at iberlef 2020: Fake news and aggressiveness analysis in… See the full description on the dataset page: https://huggingface.co/datasets/mariagrandury/fake_news_corpus_spanish.fake-news-detection-dataset-EnglishThis is a cleaned and splitted version of this dataset (https://www.kaggle.com/datasets/sadikaljarif/fake-news-detection-dataset-english)
Labels:
Fake News: 0
Real News: 1
You can find the cleansing script at: https://github.com/ErfanMoosaviMonazzah/Fake-News-Detection
FakeNewsSpanish_Kaggle1This dataset was obtained from: https://www.kaggle.com/datasets/arseniitretiakov/noticias-falsas-en-espaol
real-and-fake-newsfake_or_real_news
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Trinisha/fake_or_real_news.real-fake-news-workshopcovid_fake_newsConstraint@AAAI2021 - COVID19 Fake News Detection in English
@misc{patwa2020fighting,
title={Fighting an Infodemic: COVID-19 Fake News Dataset},
author={Parth Patwa and Shivam Sharma and Srinivas PYKL and Vineeth Guptha and Gitanjali Kumari and Md Shad Akhtar and Asif Ekbal and Amitava Das and Tanmoy Chakraborty},
year={2020},
eprint={2011.03327},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
urdu_fake_news
Dataset Card for Bend the Truth (Urdu Fake News)
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
news: a string in urdu
label: the label indicating whethere the provided news is real or fake.
category: The intent of the news being presented. The available 5… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/urdu_fake_news.spanish-fake-news-fixed
Spanish Fake News Fixed
Este dataset contiene noticias etiquetadas en español, reparado para corregir saltos de línea internos.
FakeNewsSpanish_Kaggle2This dataset was obtained from: https://www.kaggle.com/datasets/zulanac/fake-and-real-news
Unified-and-Balanced-Spanish-Fake-News-Corpus
Unified Spanish Misinformation and Satire Corpus (USMSC)
Dataset Description
This dataset is a comprehensive, deduplicated, and systematically structured corpus for domain-specific misinformation detection in Spanish social media text. It addresses the critical gap in Spanish-language resources by unifying multiple distinct datasets into a single, highly refined corpus.
Crucially, this dataset employs a three-class formulation (Fake, Real, Satire). Recent… See the full description on the dataset page: https://huggingface.co/datasets/gabrielhuav/Unified-and-Balanced-Spanish-Fake-News-Corpus.john_fake_newsISOT-Fake-News-Dataset-FineTuned-2022
Dataset for project: FakeLuke-ISOT
Dataset Description
A refined variant of the ISOT dataset.
For our binary task, only two tags are needed: type (fake or true) and text.
In order to achieve that we take both of the .csv files and we trim the article tags: title, type and publishing date.
Next, besides the text we add a new column “type” and we mark it with 0 for real news and 1 for fake news.
We further trim the Fake.csv file by eliminating all the empty columns, the… See the full description on the dataset page: https://huggingface.co/datasets/Phoenyx83/ISOT-Fake-News-Dataset-FineTuned-2022.FakeNewsNetFake_and_Real_newsFakeNewsfake-news-detection-dataset-english
Dataset Card for "fake-news-detection-dataset-english"
More Information needed
fakenews-fil-classification
Fakenews_fil_Classification
Deduplicated copy of kornwtp/fakenews-fil-classification.
Splits
split
rows
train
3,005
Fake_News_GossipCopDataset Source: Ahren09/MMSoc_GossipCop
This is a copied and reformatted version of the Ahren09/MMSoc_GossipCop
text: text of the article (str)
bert_embeddings: (768, )
roberta_embeddings: (768, )
label: (int)
0: real
1: fake
Datasets Distribution:
Train: 9988 (real: 7955, fake: 2033)
Test: 2672 (real: 2169, 503)
fake-newsfake-news-detector-euvsdisinfodata from https://euvsdisinfo.eu/
kobby_fake_newsDataset: ikekobby/40-percent-cleaned-preprocessed-fake-real-news
40-percent-cleaned-preprocessed-fake-real-newsKaggle based dataset for text classification task. The data has been cleaned and processed for preparation into any model for classification based tasks. This is just 40% of the entire dataset.
fake_news_en_opensources
Dataset Card for "Fake News Opensources"
Dataset Description
Homepage: https://github.com/AndyTheFactory/FakeNewsDataset
Repository: https://github.com/AndyTheFactory/FakeNewsDataset
Point of Contact: Andrei Paraschiv
Dataset Summary
a consolidated and cleaned up version of the opensources Fake News dataset
Fake News Corpus comprises 8,529,090 individual articles, classified into 12 classes: reliable, unreliable, political, bias, fake, conspiracy… See the full description on the dataset page: https://huggingface.co/datasets/andyP/fake_news_en_opensources.
