datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fake_news
TODO: Add YAML tags here. Copy-paste the tags obtained with the online tagging app: https://huggingface.co/spaces/huggingface/datasets-tagging
annotations_creators:
- no-annotation
language_creators:
- found
language:
- en
license:
- unknown
multilinguality:
- monolingual
size_categories:
- 30k<n<50k
source_datasets:
- original
task_categories:
- text-classification
task_ids:
- fact-checking
- intent-classification
pretty_name: GonzaloA / Fake News
Dataset Card for… See the full description on the dataset page: https://huggingface.co/datasets/GonzaloA/fake_news.fake-real-newsfake-newsfake_news_english
Dataset Card for Fake News English
Dataset Summary
This dataset contains URLs of news articles classified as either fake or satire. The articles classified as fake also have the URL of a rebutting article.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
English
Dataset Structure
Data Instances
{
"article_number": 102 ,
"url_of_article":… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/fake_news_english.fake_news_corpus_spanish
Fake News Corpus Spanish
Citation
Gómez-Adorno, H., Posadas-Durán, J. P., Enguix, G. B., & Capetillo, C. P. (2021). Overview of FakeDeS at IberLEF 2021: Fake News Detection in Spanish Shared Task. Procesamiento del Lenguaje Natural, 67, 223-231.
Aragón, M. E., Jarquín, H., Gómez, M. M. Y., Escalante, H. J., Villaseñor-Pineda, L., Gómez-Adorno, H., ... & Posadas-Durán, J. P. (2020, September). Overview of mex-a3t at iberlef 2020: Fake news and aggressiveness analysis in… See the full description on the dataset page: https://huggingface.co/datasets/mariagrandury/fake_news_corpus_spanish.fake-news-detection-dataset-EnglishThis is a cleaned and splitted version of this dataset (https://www.kaggle.com/datasets/sadikaljarif/fake-news-detection-dataset-english)
Labels:
Fake News: 0
Real News: 1
You can find the cleansing script at: https://github.com/ErfanMoosaviMonazzah/Fake-News-Detection
real-and-fake-newsfake_or_real_news
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Trinisha/fake_or_real_news.real-fake-news-workshopcovid_fake_newsConstraint@AAAI2021 - COVID19 Fake News Detection in English
@misc{patwa2020fighting,
title={Fighting an Infodemic: COVID-19 Fake News Dataset},
author={Parth Patwa and Shivam Sharma and Srinivas PYKL and Vineeth Guptha and Gitanjali Kumari and Md Shad Akhtar and Asif Ekbal and Amitava Das and Tanmoy Chakraborty},
year={2020},
eprint={2011.03327},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
urdu_fake_news
Dataset Card for Bend the Truth (Urdu Fake News)
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
news: a string in urdu
label: the label indicating whethere the provided news is real or fake.
category: The intent of the news being presented. The available 5… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/urdu_fake_news.spanish-fake-news-fixed
Spanish Fake News Fixed
Este dataset contiene noticias etiquetadas en español, reparado para corregir saltos de línea internos.
FakeNewsSpanish_Kaggle2This dataset was obtained from: https://www.kaggle.com/datasets/zulanac/fake-and-real-news
john_fake_newsFake_and_Real_newsFakeNewsfake-news-detection-dataset-english
Dataset Card for "fake-news-detection-dataset-english"
More Information needed
fakenews-fil-classification
Fakenews_fil_Classification
Deduplicated copy of kornwtp/fakenews-fil-classification.
Splits
split
rows
train
3,005
Fake_News_GossipCopDataset Source: Ahren09/MMSoc_GossipCop
This is a copied and reformatted version of the Ahren09/MMSoc_GossipCop
text: text of the article (str)
bert_embeddings: (768, )
roberta_embeddings: (768, )
label: (int)
0: real
1: fake
Datasets Distribution:
Train: 9988 (real: 7955, fake: 2033)
Test: 2672 (real: 2169, 503)
fake-newsfake-news-detector-euvsdisinfodata from https://euvsdisinfo.eu/
kobby_fake_newsDataset: ikekobby/40-percent-cleaned-preprocessed-fake-real-news
40-percent-cleaned-preprocessed-fake-real-newsKaggle based dataset for text classification task. The data has been cleaned and processed for preparation into any model for classification based tasks. This is just 40% of the entire dataset.
fake_news_en_opensources
Dataset Card for "Fake News Opensources"
Dataset Description
Homepage: https://github.com/AndyTheFactory/FakeNewsDataset
Repository: https://github.com/AndyTheFactory/FakeNewsDataset
Point of Contact: Andrei Paraschiv
Dataset Summary
a consolidated and cleaned up version of the opensources Fake News dataset
Fake News Corpus comprises 8,529,090 individual articles, classified into 12 classes: reliable, unreliable, political, bias, fake, conspiracy… See the full description on the dataset page: https://huggingface.co/datasets/andyP/fake_news_en_opensources.Fake_News_KDD2020Dataset Source: Fake News Detection Challenge KDD 2020
This is a copied and reformatted version of the Fake News Detection Challenge KDD 2020.
We use the raw train.csv from the official Kaggle Dataset and split the data into train and test sets.
text: text of the article (str)
embeddings: BERT embeddings (768, )
label: (int)
1: fake
0: true
Datasets Distribution:
Train: 4487
Test: 499
fake-news-detector-datasetPolitifact_fake_newscentral_de_fatos
Central de Fatos
Dataset Summary
In recent times, the interest for research dissecting the dissemination and prevention of misinformation in the online environment has spiked dramatically.
Given that scenario, a recurring obstacle is the unavailability of public datasets containing fact-checked instances.
In this work, we performed an extensive data collection of such instances from the better part of all major internationally recognized Brazilian fact-checking agencies.… See the full description on the dataset page: https://huggingface.co/datasets/fake-news-UFG/central_de_fatos.fake_news_combinedLabel Description
0 : Fake,
1 : Real
bcms-fake-news-articles
