datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
COVID-19-disinformation
COVID-19 Infodemic Multilingual Dataset
This repository contains a multilingual dataset related to the COVID-19 infodemic, annotated with fine-grained labels. The dataset is curated to address questions of interest to journalists, fact-checkers, social media platforms, policymakers, and the general public. The dataset includes tweets in Arabic, Bulgarian, Dutch, and English, focusing on both binary (misinformation detection) and multiclass classification (different types of… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/COVID-19-disinformation.mteb-nl-COVID-19-disinformationThis dataset contains Dutch tweets that were annotated for fine-grained disinformation analysis.
Citation Information
If you find our paper, benchmark or models helpful, please consider cite as follows:
@misc{banar2025mtebnle5nlembeddingbenchmark,
title={MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch},
author={Nikolay Banar and Ehsan Lotfi and Jens Van Nooten and Cristina Arhiliuc and Marija Kliocaite and Walter Daelemans},
year={2025}… See the full description on the dataset page: https://huggingface.co/datasets/clips/mteb-nl-COVID-19-disinformation.disinformation-research-gaps
Disinformation Research Gap Atlas
     
An evidence-grounded atlas of research gaps in the global disinformation literature —… See the full description on the dataset page: https://huggingface.co/datasets/bugraayantr/disinformation-research-gaps.africa-election-disinformation
Election Disinformation & Political Cyber Operations (Africa) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet, optimized-parquet - Sector: governance_security - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-election-disinformation.COVID_19_Disinformation_arfrugal-climate-disinformation
Dataset Card for frugal-climate-disinformation
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/dvilasuero/frugal-climate-disinformation/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/dvilasuero/frugal-climate-disinformation.disinformationedmo-disinformation-question-answer
EDMO Disinformation Question-and-answer Dataset
We present a question-and-answer dataset based on the research available at
EDMO related to disinformation.
The dataset was generated using ChatGPT, which was provided with sections of the
above research and tasked with generated questions related to disinformation and
the section of text provided.
Almost 240 PDF documents were used to generate this data, resulting in a total
of 5,991 questions.
The data is in the format required to… See the full description on the dataset page: https://huggingface.co/datasets/jameslloyd2/edmo-disinformation-question-answer.disinformation_wedging
Dataset Card for "disinformation_wedging"
More Information needed
german-disinformation-narratives-synthetic
Dataset Card for "german-disinformation-narratives-synthetic"
Dataset Summary
This dataset is built using GPT-4o and contains 41 narratives that are frequently encountered by German fact-checkers. Each narrative has been expanded using GPT-4o to generate multiple supporting, unrelated, or contradicting text units. These expansions are presented from various perspectives, including that of a politician, an angry citizen, a conspiracy theorist, and others.
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Sami92/german-disinformation-narratives-synthetic.
