datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
covid19_emergency_event
Dataset Card for EXCEPTIUS Corpus
Dataset Summary
This dataset presents a new corpus of legislative documents from 8 European countries (Beglium, France, Hunary, Italy, Netherlands, Norway, Poland, UK) in 7 languages (Dutch, English, French, Hungarian, Italian, Norwegian Bokmål, Polish) manually annotated for exceptional measures against COVID-19. The annotation was done on the sentence level.
Supported Tasks and Leaderboards
The dataset can be used for… See the full description on the dataset page: https://huggingface.co/datasets/joelniklaus/covid19_emergency_event.COVID-19-disinformation
COVID-19 Infodemic Multilingual Dataset
This repository contains a multilingual dataset related to the COVID-19 infodemic, annotated with fine-grained labels. The dataset is curated to address questions of interest to journalists, fact-checkers, social media platforms, policymakers, and the general public. The dataset includes tweets in Arabic, Bulgarian, Dutch, and English, focusing on both binary (misinformation detection) and multiclass classification (different types of… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/COVID-19-disinformation.covid19-ct-seg
COVID-19 CT Segmentation Dataset
Dataset Description
The COVID-19 CT Segmentation dataset for lung and COVID-19 infection segmentation from CT scans. This dataset contains CT scans with dense segmentation annotations.
Dataset Details
Modality: CT
Target: left lung, right lung, COVID-19 infection
Format: NIfTI (.nii.gz)
Dataset Structure
Each sample in the JSONL file contains:
{
"image": "path/to/image.nii.gz",
"mask":… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/covid19-ct-seg.COVID-19_EUR-LEX
[!NOTE]
Dataset origin: https://portulanclarin.net/repository/browse/covid-19-eur-lex-dataset-multilingual-cef-languages/4dbe6648186d11eb994702420a000409a958a80a367b4b33a599c8cd2aab7383/
Description
Multilingual (CEF languages) corpus acquired from website (https://eur-lex.europa.eu/legal-content) of the EU portal (9th July 2020). It contains 23 TMX files (EN-X, X is a CEF language) with 475,931 translation units pairs in total.
COVID-19_EU_presscorner_monolingual
[!NOTE]
Dataset origin: https://live.european-language-grid.eu/catalogue/corpus/19444
Description
This collection of monolingual documents was generated from content available at https://ec.europa.eu/commission/presscorner.
It includes 1354 documents in total in the following languages: EN 335 DE 266 ES 115 FR 123 IT 273 EL 120 SV 122
Citation
EU presscorner monolingual collections of COVID-19 related documents. (2020, June 15). Version 1.0. [Dataset (Text corpus)].… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/COVID-19_EU_presscorner_monolingual.COVID-19_weibo_emotionCOVID-19 Epidemic Weibo Emotional Dataset, the content of Weibo in this dataset is the epidemic Weibo obtained by using relevant keywords to filter during the epidemic, and its content is related to COVID-19.
Each tweet is labeled as one of the following six categories: neutral (no emotion), happy (positive), angry (angry), sad (sad), fear (fear), surprise (surprise)
The COVID-19 Weibo training dataset includes 8,606 Weibos, the validation set contains 2,000 Weibos, and the test dataset… See the full description on the dataset page: https://huggingface.co/datasets/souljoy/COVID-19_weibo_emotion.COVID-19-20
COVID-19-20 Lung CT Lesion Segmentation Challenge
Non-contrast chest CT with radiologist-verified binary COVID-19 lesion
masks, from the MICCAI 2020 COVID-19 Lung CT Lesion Segmentation Challenge
(a.k.a. COVID-19-20).
⚠️ Scope of this upload (training split only)
This repository contains the public training split: 199 CT volumes, each with a
ground-truth lesion mask. It is a faithful subset of the full challenge:
Split
Cases
Masks
Included here
Train… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/COVID-19-20.africa-burkinafaso-covid19-subnational
Burkina Faso: Coronavirus (Covid-19) Subnational | Africa (original)
Size category: 1K<n<10K - Formats: parquet - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets help researchers examine… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-burkinafaso-covid19-subnational.covid19-cxr-bm3d-denoising
Dataset Card for COVID-19 CXR BM3D Denoising
Dataset Summary
This dataset contains paired chest X-ray (CXR) images for supervised image denoising.It is built from the COVID-19 Radiography Database and, for each original CXR, provides:
the original (clean) image,
a synthetically noisy version,
a BM3D-denoised version (pseudo–ground truth).
The goal is to train and evaluate models that learn to reproduce BM3D from noisy CXRs.
Important: This is a derived… See the full description on the dataset page: https://huggingface.co/datasets/adlito/covid19-cxr-bm3d-denoising.COVID-19_HEALTH
[!NOTE]
Dataset origin: https://live.european-language-grid.eu/catalogue/corpus/21222
Description
The dataset contains 327 X-Y TMX files, where X and Y belong to the set {CEF language plus IS and NO} (3905604 TUs in total). Acquisition of data (from multi/bi-lingual websites), normalization, cleaning, deduplication and identification of parallel documents have been done by ILSP-FC tool. Multilingual embeddings (LASER) were used for alignment of segments. Merging/filtering of… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/COVID-19_HEALTH.COVID-19_OSHA-EUROPA
[!NOTE]
Dataset origin: https://live.european-language-grid.eu/catalogue/corpus/21210
Description
Multilingual (CEF languages) corpus acquired (17th july 2020) from website (https://osha.europa.eu/) of the European Agency for Safety and Health at Work (17th July 2020). It contains 24 TMX files (EN-X, where X is a CEF language plus IS and NB but not Irish) with 1169 TUs in total.
Citation
COVID-19 OSHA-EUROPA dataset v1. Multilingual (CEF languages plus IS and NB but not… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/COVID-19_OSHA-EUROPA.COVID-19_HEALTH_Wikipedia
[!NOTE]
Dataset origin: https://live.european-language-grid.eu/catalogue/corpus/19430/
Description
Multilingual (52 EN-X language pairs) corpus acquired from Wikipedia on health and COVID-19 domain (2nd May 2020). It contains 81671 TUs in total.
Citation
COVID-19 - HEALTH Wikipedia dataset. Multilingual (52 EN-X language pairs) (2020, May 02). Version 1.0. [Dataset (Text corpus)]. Source: European Language Grid.… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/COVID-19_HEALTH_Wikipedia.PhoNER_COVID19COVID-19_EU_presscorner_multilingual
[!NOTE]
Dataset origin: https://portulanclarin.net/repository/browse/covid-19-eu-presscorner-v2-dataset-multilingual-cef-languages/8b165624184811eb821802420a000409d0d66a6591694c9f80bece8037a339b9/
Description
Multilingual (CEF languages) corpus acquired from website (https://ec.europa.eu/commission/presscorner/) of the EU portal (8th July 2020). It contains 23 TMX files (EN-X, where X is a CEF language) with 151895 TUs in total.
COVID-19_EC-EUROPA
[!NOTE]
Dataset origin: https://portulanclarin.net/repository/browse/covid-19-ec-europa-v1-dataset-multilingual-cef-languages/5c340abe146e11eb8d9d02420a0004093a2977c08692490faf372638eb994046/
Description
Multilingual (CEF languages) corpus acquired from website (https://ec.europa.eu/*coronavirus-response) of the EU portal (20th May 2020). It contains 23 TMX files (EN-X, where X is a CEF language) with 53311 TUs in total.
bimcv_covid19_all_cxrCOVID-19-EUROPARLv2
[!NOTE]
Dataset origin: https://live.european-language-grid.eu/catalogue/corpus/19451
Description
Multilingual (24 CEF languages) corpus acquired from the website (https://www.europarl.europa.eu/) of the European Parliament (9th May 2020). It contains 14631 TUs in total.
Citation
COVID-19 EUROPARL dataset v2. Multilingual (24 CEF languages) (2020, May 10). Version 2.0. [Dataset (Text corpus)]. Source: European Language Grid.… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/COVID-19-EUROPARLv2.COVID-19coronavirusCOVID-19-EUROPARL
[!NOTE]
Dataset origin: https://live.european-language-grid.eu/catalogue/corpus/19432
Description
Multilingual (24 CEF languages) corpus acquired from the website (https://www.europarl.europa.eu/) of the European Parliament (25th April 2020). It contains 9388 TUs in total.
Citation
COVID-19 EUROPARL dataset v1. Multilingual (24 CEF languages) (2020, May 02). Version 1.0. [Dataset (Text corpus)]. Source: European Language Grid.… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/COVID-19-EUROPARL.COVID-19-vaccine-attitude-tweets
Dataset Card for COVID-19-vaccine-attitude-tweets
Dataset Summary
The dataset consists of 2564 manually annotated tweets related to COVID-19 vaccines. The dataset can be used to discover the attitude expressed in the tweet towards the subject of COVID-19 vaccines. Tweets are in English. The dataset was curated in such a way as to maximize the likelihood of tweets with a strong emotional tone. We have assumed the existence of three classes:
PRO (label 0): positive, the… See the full description on the dataset page: https://huggingface.co/datasets/webimmunization/COVID-19-vaccine-attitude-tweets.bimcv_covid19_sampledafrica-nigeria-covid19-subnational
Nigeria: Coronavirus (Covid-19) Subnational | Africa (original)
Size category: 10K<n<100K - Formats: parquet - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets help researchers examine… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-nigeria-covid19-subnational.global-covid19-twitter-datasetafrica-gambia-covid19-subnational
Gambia: Coronavirus (Covid-19) Subnational | Africa (original)
Size category: 1K<n<10K - Formats: parquet - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets help researchers examine… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-gambia-covid19-subnational.COVID-19-CT-Chinese
COVID-19 CT Dataset — Chest CT with Chinese Findings and Conclusions
A chest CT dataset of 368 studies / 3,680 axial slices paired with the original
Chinese radiology reports (free-text findings and conclusion) written by
attending radiologists during the early COVID-19 outbreak in China.
This is the dataset released with the TNNLS paper
Medical-VLBERT: Medical Visual Language BERT for COVID-19 CT Report Generation
With Alternate Learning
(arXiv:2108.05067).
Project page:… See the full description on the dataset page: https://huggingface.co/datasets/guangyil/COVID-19-CT-Chinese.COVID-19-USAHELLOv2
[!NOTE]
Dataset origin: https://live.european-language-grid.eu/catalogue/corpus/21347
Description
Multilingual (EN, AR, ES, FA, FR, IT, KO, PT, RU, TL, TR, UK, UR, VI, ZH) corpus acquired from the website https://usahello.org/, a free online center for information and education for refugees, asylum seekers, immigrants and welcoming communities (9th August 2020). It contains 41165 TUs in total.
Citation
COVID-19 USAHELLO dataset v2. Multilingual (EN, AR, ES, FA, FR, IT… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/COVID-19-USAHELLOv2.africa-burkinafaso-covid19-city-level
Burkina Faso: Coronavirus (Covid-19) City level | Africa (original)
Size category: 1K<n<10K - Formats: parquet - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets help researchers examine… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-burkinafaso-covid19-city-level.Covid19VoyagerPBMCPhoNer_Covid19
Dataset Card for "PhoNer_Covid19"
More Information needed
africa-mauritania-covid19-subnational
Mauritania: Coronavirus (Covid-19) Subnational | Africa (original)
Size category: 10K<n<100K - Formats: parquet - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets help researchers examine… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-mauritania-covid19-subnational.
