datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
trec-covid-qrels
Dataset Card for BEIR Benchmark
Dataset Summary
BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks:
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018
Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus
News Retrieval: TREC-NEWS, Robust04
Argument Retrieval: Touche-2020, ArguAna
Duplicate Question Retrieval: Quora, CqaDupstack
Citation-Prediction: SCIDOCS
Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/trec-covid-qrels.covid_fake_newsConstraint@AAAI2021 - COVID19 Fake News Detection in English
@misc{patwa2020fighting,
title={Fighting an Infodemic: COVID-19 Fake News Dataset},
author={Parth Patwa and Shivam Sharma and Srinivas PYKL and Vineeth Guptha and Gitanjali Kumari and Md Shad Akhtar and Asif Ekbal and Amitava Das and Tanmoy Chakraborty},
year={2020},
eprint={2011.03327},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
COVID-19_weibo_emotionCOVID-19 Epidemic Weibo Emotional Dataset, the content of Weibo in this dataset is the epidemic Weibo obtained by using relevant keywords to filter during the epidemic, and its content is related to COVID-19.
Each tweet is labeled as one of the following six categories: neutral (no emotion), happy (positive), angry (angry), sad (sad), fear (fear), surprise (surprise)
The COVID-19 Weibo training dataset includes 8,606 Weibos, the validation set contains 2,000 Weibos, and the test dataset… See the full description on the dataset page: https://huggingface.co/datasets/souljoy/COVID-19_weibo_emotion.long-covid-classification-data
Data Description
Long-COVID related articles have been manually collected by information specialists.Please find further information here.
Size
Training
Development
Test
Total
Positive Examples
215
76
70
345
Negative Examples
199
62
68
345
Total
414
238
138
690
Citation
@article{10.1093/database/baac048,author = {Langnickel, Lisa and Darms, Johannes and Heldt, Katharina and Ducks, Denise and Fluck, Juliane},title = "{Continuous development… See the full description on the dataset page: https://huggingface.co/datasets/llangnickel/long-covid-classification-data.covid-conspiracy-redditRegional-COVID-19-dataset-Sri-Lanka
Regional COVID-19 Cases in Sri Lanka Dataset
Overview
This repository contains a dataset documenting the COVID-19 cases in Districts of Sri Lanka. The data is organized by year, quarter, month, day, number of cases, and district.
Dataset Information
Year: 2021
Quarters: Qtr 1
Months: January
Days (shows date of update): 1
Data Columns
Year: The year in which the data was recorded.
Quarter: The quarter of the year.
Month: The month… See the full description on the dataset page: https://huggingface.co/datasets/DePubudu/Regional-COVID-19-dataset-Sri-Lanka.covid_fact_checked_google_apiThis dataset was gathered from the Google Fact Checker API, using an automatic web scraper. 10,000 facts were pulled, but for the sake of simplicity, only ones were the ratings were singular words "false" or "true", were kept, which filtered it down to ~3000 fact checks, with about 90% of the facts being false.
annotations_creators:
expert-generated
language_creators:
crowdsourced
languages:
en-US
licenses:
unknown
multilinguality:
monolingual
pretty_name: polifact-covid-fact-checker… See the full description on the dataset page: https://huggingface.co/datasets/justinqbui/covid_fact_checked_google_api.CDC-COVID-FAQDataset extracted from https://www.cdc.gov/coronavirus/2019-ncov/hcp/faq.html#Treatment-and-Management.
COVID-19coronavirusfr_covid_news
Dataset Card for COVID-19 French News dataset
Dataset Summary
The COVID-19 French News dataset is a French-language dataset containing just over 40k unique news articles from more than 50 different French-speaking online newspapers. The dataset has been prepared using news-please - an integrated web crawler and information extractor for news. The current version supports abstractive summarization and topic classification. Dataset Card not finished yet.… See the full description on the dataset page: https://huggingface.co/datasets/gustavecortal/fr_covid_news.covid_news
Dataset Card for COVID News Articles (2020 - 2022)
Dataset Summary
The dataset encapsulates approximately half a million news articles collected over a period of 2 years during the Coronavirus pandemic onset and surge. It consists of 3 columns - title, content and category. title refers to the headline of the news article. content refers to the article in itself and category denotes the overall context of the news article at a high level. The dataset encapsulates… See the full description on the dataset page: https://huggingface.co/datasets/osanseviero/covid_news.vaccine-hesitancy-for-covid-19-public-use-microdat
Vaccine Hesitancy for COVID-19: Public Use Microdata Areas (PUMAs)
Description
Due to the change in the survey instrument regarding intention to vaccinate, our estimates for “hesitant or unsure” or “hesitant” derived from April 14-26, 2021, are not directly comparable with prior Household Pulse Survey data and should not be used to examine trends in hesitancy.
To support state and local communication and outreach efforts, ASPE developed state, county, and sub-state level… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/vaccine-hesitancy-for-covid-19-public-use-microdat.Twitter-COVID-19General description:
This dataset comprisses a set of tweets crawled during the COVID-19 pandemic (from March 2020 to June 2021). Tweets are located in two different regions: Spain and USA. This adds value to the collection, as it contains data in two languages.
This data was used as part of a broader study that aimed to determine the evolution of different personality traits and disorders during the pandemic. Thus, weak labels for different dimensions, such as sentiment, personality… See the full description on the dataset page: https://huggingface.co/datasets/citiusLTL/Twitter-COVID-19.covid-vaccine-conspiracy
COVID-19 Vaccine Conspiracy Theory Detection Dataset
Dataset Summary
This dataset contains 581 manually labeled social media comments related to COVID-19 vaccines,
annotated for the presence of conspiracy theories. It supports NLP research on automated vaccine
misinformation detection, a critical public health challenge.
This dataset has been cited and reused by external researchers.
Associated Paper
Amin, M. H., Madanu, H., Lavu, S., Mansourifar, H.… See the full description on the dataset page: https://huggingface.co/datasets/AminHasibul/covid-vaccine-conspiracy.provisional-covid-19-death-counts-rates-and-percen
Provisional COVID-19 death counts, rates, and percent of total deaths, by jurisdiction of residence
Description
This file contains COVID-19 death counts, death rates, and percent of total deaths by jurisdiction of residence. The data is grouped by different time periods including 3-month period, weekly, and total (cumulative since January 1, 2020). United States death counts and rates include the 50 states, plus the District of Columbia and New York City. New York state… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/provisional-covid-19-death-counts-rates-and-percen.Message-Severity-COVID-19
Datset Details
This is the dataset from the following paper: Identifying the Perceived Severity of Patient-Generated Telemedical Queries Regarding COVID: Developing and Evaluating a Transfer Learning–Based Solution
Citation
If you use this dataset in your work, please cite the following paper:
```
@article{gatto2022identifying,
title={Identifying the perceived severity of patient-generated telemedical queries regarding covid: Developing and evaluating a transfer… See the full description on the dataset page: https://huggingface.co/datasets/PortalPal-AI/Message-Severity-COVID-19.trec-covid-pl-qrelsPart of BEIR-PL: Zero Shot Information Retrieval Benchmark for the Polish Language.
Link to arxiv: https://arxiv.org/pdf/2305.19840.pdf
Contact: konrad.wojtasik@pwr.edu.pl
long-covid-recovery-interaction-geometry-v0.1
What this dataset does
This dataset tests whether failed Long Covid recovery is better predicted by interaction geometry than by any single biological burden signal.
The target is not disease diagnosis.
The target is failed recovery.
Core stability idea
A patient may show high biological burden and still recover if repair capacity remains strong and omics layers stay coherent.
A patient may show moderate burden and fail to recover if repair capacity weakens and… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/long-covid-recovery-interaction-geometry-v0.1.long-covid-recovery-trajectory-competition-v0.1
What this dataset does
This dataset tests whether a model can identify which Long Covid recovery trajectory is currently emerging from a biological state.
The task is trajectory forecasting.
It is not diagnosis.
Core stability idea
Two patients can show similar symptom burden while moving toward different futures.
The benchmark tests whether models can infer direction rather than classify current severity.
Prediction target
The target column is:… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/long-covid-recovery-trajectory-competition-v0.1.provisional-percent-of-deaths-for-covid-19-influen
Provisional Percent of Deaths for COVID-19, Influenza, and RSV
Description
This file contains the provisional percent of total deaths by week for COVID-19, Influenza, and Respiratory Syncytial Virus for deaths occurring among residents in the United States. Provisional data are based on non-final counts of deaths based on the flow of mortality data in National Vital Statistics System.
Dataset Details
Publisher: Centers for Disease Control and Prevention… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/provisional-percent-of-deaths-for-covid-19-influen.long-covid-closed-loop-recovery-control-v1.0
What this dataset does
This dataset tests whether a model can identify successful closed-loop recovery control in a synthetic Long Covid recovery setting.
The task is not diagnosis.
The task is control success prediction.
Core stability idea
A recovery plan may begin well but fail if feedback is poor, timing is wrong, adaptation is slow, or the control sequence diverges.
This dataset tests whether models can identify when a recovery control loop is likely to… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/long-covid-closed-loop-recovery-control-v1.0.trec-covid_neomme_260m_li
trec-covid_neomme_260m_li
Multi-vector (late-interaction) embeddings of BEIR trec-covid (beir/trec-covid), encoded with
Hcompany/NeoMME-260M-Retriever-ST-late at revision 023be2a8ab9d797f5aa76f5bf8b5dde78d819659.
Source data: ir_datasets beir/trec-covid (ir_datasets 0.6.3), which downloads trec-covid.zip (md5 ce62140cb23feb9becf6270d0d1fe6d1). BEIR also publishes this corpus on the Hub as BeIR/trec-covid, whose card gives this dataset's license; the data here was loaded through… See the full description on the dataset page: https://huggingface.co/datasets/robro612/trec-covid_neomme_260m_li.united-states-covid-19-county-level-data-sources
United States COVID-19 County Level Data Sources - ARCHIVED
Description
The Public Health Emergency (PHE) declaration for COVID-19 expired on May 11, 2023. As a result, the Aggregate Case and Death Surveillance System will be discontinued. Although these data will continue to be publicly available, this dataset will no longer be updated.
On October 20, 2022, CDC began retrieving aggregate case and death data from jurisdictional and state partners weekly instead of daily.… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/united-states-covid-19-county-level-data-sources.provisional-covid-19-deaths-by-sex-and-age
Provisional COVID-19 Deaths by Sex and Age
Description
Effective September 27, 2023, this dataset will no longer be updated. Similar data are accessible from wonder.cdc.gov.
Deaths involving COVID-19, pneumonia, and influenza reported to NCHS by sex, age group, and jurisdiction of occurrence.
Dataset Details
Publisher: Centers for Disease Control and Prevention
Temporal Coverage: 2020-01-01/2023-07-29
Geographic Coverage: United States, Puerto Rico
Last… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/provisional-covid-19-deaths-by-sex-and-age.trec-covid_lateon_hpool_regularized
trec-covid_lateon_hpool_regularized
Multi-vector (late-interaction) embeddings of BEIR trec-covid (beir/trec-covid), encoded with
lightonai/LateOn-hpool-regularized at revision 3f9c577e4959ffd717a2565ba8db62f3d5f9d34c.
Source data: ir_datasets beir/trec-covid (ir_datasets 0.6.3), which downloads trec-covid.zip (md5 ce62140cb23feb9becf6270d0d1fe6d1). BEIR also publishes this corpus on the Hub as BeIR/trec-covid, whose card gives this dataset's license; the data here was loaded… See the full description on the dataset page: https://huggingface.co/datasets/robro612/trec-covid_lateon_hpool_regularized.COVID-19-conspiracy-theories-tweets
Dataset Summary
This dataset consists of 6591 tweets generated by GPT-3.5 model. The tweets are juxtaposed with a conspiracy theory related to COVID-19 pandemic. Each item consists of a label that represents the item's output class. The possible labels are support/deny/neutral.
support: the tweet suggests support for the conspiracy theory
deny: the tweet contradicts the conspiracy theory
neutral: the tweet is mostly informative, and does not show emotions against the conspiracy… See the full description on the dataset page: https://huggingface.co/datasets/webimmunization/COVID-19-conspiracy-theories-tweets.covid-19-reported-patient-impact-and-hospital-capa
COVID-19 Reported Patient Impact and Hospital Capacity by State
Description
After May 3, 2024, this dataset and webpage will no longer be updated because hospitals are no longer required to report data on COVID-19 hospital admissions, and hospital capacity and occupancy data, to HHS through CDC’s National Healthcare Safety Network. Data voluntarily reported to NHSN after May 1, 2024, will be available starting May 10, 2024, at COVID Data Tracker Hospitalizations.
The… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/covid-19-reported-patient-impact-and-hospital-capa.rus-trec-covid-qrelsitalian_long_covid_tweetswho-covid-19-epidemiological-update-edition-163
World Health Organization (WHO) Epidemiological Update - Edition 163 (for embeddings)
train.pnf is taken from the WHO website
test.csv was generated by GPT-3.5-turbo
All text is chunked to a length of 500 tokens with 10% overlap.
