clinical-text
nomic-embed-text-v1.5_Clinical-Trials_Matryoshka_finalClinicalEncoder26AM-Diagnosable-Colbert-L2-for-multilingual-medical-textsnomic-embed-text-v1.5_Clinical-Trials_Matryoshka2nomic-embed-text-v1.5_Clinical-Trials_Matryoshkaroberta-base-biomedical-clinical-es-finetuned-text_classificationclinical_text-summtiny-clinicalbert-medical-text-classificationdistil-clinicalbert-medical-text-classification
clinical_trial_texts
Dataset Card for "clinical_trial_texts"
These are the text of clinical trials dowloaded from https://ClinicalTrials.gov/AllAPIJSON.zip on Dec 3rd 2022.
Total trials is 434977
Number of tokens is 2,184,397,556 (2.1bn tokens).
The tokens here are from the default BERT tokenizer in hugginface.
This data can be used for pretraining in the clinical trial and biomedical domains.
If you use this data please acknowledge @domenicrosati and link to this dataset
More Information needed
clinical-synthetic-text-kg
Data Description
We release the synthetic data generated using the method described in the paper Knowledge-Infused Prompting: Assessing and Advancing Clinical Text Data Generation with Large Language Models
(ACL 2024 Findings). The external knowledge we use is based on external knowledge graphs.
Generated Datasets
The original train/validation/test data, and the generated synthetic training data are listed as follows. For each dataset, we generate 5000 synthetic… See the full description on the dataset page: https://huggingface.co/datasets/ritaranx/clinical-synthetic-text-kg.clinical_trial_texts
Dataset Card for "clinical_trial_texts"
These are the text of clinical trials dowloaded from https://ClinicalTrials.gov/AllAPIJSON.zip on Dec 3rd 2022.
Total trials is 434977
Number of tokens is 2,184,397,556 (2.1bn tokens).
The tokens here are from the default BERT tokenizer in hugginface.
This data can be used for pretraining in the clinical trial and biomedical domains.
If you use this data please acknowledge @domenicrosati and link to this dataset
More Information needed
clinical-synthetic-text-llm
Data Description
We release the synthetic data generated using the method described in the paper Knowledge-Infused Prompting: Assessing and Advancing Clinical Text Data Generation with Large Language Models
(ACL 2024 Findings). The external knowledge we use is based on LLM-generated topics and writing styles.
Generated Datasets
The original train/validation/test data, and the generated synthetic training data are listed as follows. For each dataset, we generate 5000… See the full description on the dataset page: https://huggingface.co/datasets/ritaranx/clinical-synthetic-text-llm.VN-Clinical-TextText-Clinical-Records
