datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
covid19_emergency_event
Dataset Card for EXCEPTIUS Corpus
Dataset Summary
This dataset presents a new corpus of legislative documents from 8 European countries (Beglium, France, Hunary, Italy, Netherlands, Norway, Poland, UK) in 7 languages (Dutch, English, French, Hungarian, Italian, Norwegian Bokmål, Polish) manually annotated for exceptional measures against COVID-19. The annotation was done on the sentence level.
Supported Tasks and Leaderboards
The dataset can be used for… See the full description on the dataset page: https://huggingface.co/datasets/joelniklaus/covid19_emergency_event.covid19-ct-seg
COVID-19 CT Segmentation Dataset
Dataset Description
The COVID-19 CT Segmentation dataset for lung and COVID-19 infection segmentation from CT scans. This dataset contains CT scans with dense segmentation annotations.
Dataset Details
Modality: CT
Target: left lung, right lung, COVID-19 infection
Format: NIfTI (.nii.gz)
Dataset Structure
Each sample in the JSONL file contains:
{
"image": "path/to/image.nii.gz",
"mask":… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/covid19-ct-seg.covid19-ct-seg
COVID-19 CT Segmentation Dataset
Dataset Description
The COVID-19 CT Segmentation dataset for lung and COVID-19 infection segmentation from CT scans. This dataset contains CT scans with dense segmentation annotations.
Dataset Details
Modality: CT
Target: left lung, right lung, COVID-19 infection
Format: NIfTI (.nii.gz)
Dataset Structure
Each sample in the JSONL file contains:
{
"image": "path/to/image.nii.gz",
"mask":… See the full description on the dataset page: https://huggingface.co/datasets/Zhao-zi/covid19-ct-seg.COVID-19_qa_pairs
Dataset Card for COVID-19_qa_pairs dataset
Dataset Summary
This datasets includes 604 question-answer pairs related to COVID-19 pandemic machine translated in Greek language.
The data is extracted from the official website of WHO.
Data Fields
question: Query question
document: Answer to the question
Bias, Risks, and Limitations
This dataset is the result of machine translation.
Licensing Information
The dataset is licensed under the… See the full description on the dataset page: https://huggingface.co/datasets/panosgriz/COVID-19_qa_pairs.COVID-19-el-corpus
Dataset Card for
Dataset Summary
This corpus contains Greek-language texts about the COVID-19 pandemic including relevant information, FAQs, etc. The texts were collected from official websites (WHO, ECDC, NPHO, covid19.gov.gr) and articles from the greek Wikipedia. Total number of words: 204,748.
Data Fields
Each instance contains:
content: Plain text
id: Instance ID
title: A document title (only in instances related to Wikipedia articles)
Licensing… See the full description on the dataset page: https://huggingface.co/datasets/panosgriz/COVID-19-el-corpus.covid19-vaccine-ae-detection-synthetic
COVID-19 Vaccine Adverse Event Detection (Synthetic)
This dataset contains synthetic examples generated to train and evaluate large language models (LLMs) to detect whether a passage of text describes an adverse event (AE) following COVID-19 vaccination.
Dataset Structure
The dataset is split into:
train.jsonl: 500 examples for training
test.jsonl: 100 examples for testing
Each entry follows the Alpaca-style instruction format with the following fields:
instruction:… See the full description on the dataset page: https://huggingface.co/datasets/podiche/covid19-vaccine-ae-detection-synthetic.
