datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
clinical-synthetic-text-llm
Data Description
We release the synthetic data generated using the method described in the paper Knowledge-Infused Prompting: Assessing and Advancing Clinical Text Data Generation with Large Language Models
(ACL 2024 Findings). The external knowledge we use is based on LLM-generated topics and writing styles.
Generated Datasets
The original train/validation/test data, and the generated synthetic training data are listed as follows. For each dataset, we generate 5000… See the full description on the dataset page: https://huggingface.co/datasets/ritaranx/clinical-synthetic-text-llm.clinical-trial-llm-v2Clinical_trials_anchor-contextORpositive-ground-truth_LLM_LORA-junk_handled_ft
Dataset details:-
This dataset is the 2nd iteration following bugs in 1st dataset.
The initial data suffered with followoing cases:-
(i) The failed reference_answers generation(due error totalling 23) primarly because of 2 reasons/exceptions:- (a) There was normal limit(300 in 1st request) and worst case limit(450 in 3rd request) number of tokens for consolidated 4 refernce_answers per chunk and its 4 corresponding answers. However certain answers breached this higher… See the full description on the dataset page: https://huggingface.co/datasets/vab46/Clinical_trials_anchor-contextORpositive-ground-truth_LLM_LORA-junk_handled_ft.Clinical_trials_anchor-contextORpositive-ground-truth_LLM_LORA_ft
Dataset details:-
This dataset is basically mapping of final anchor-positive pair data with their refernce answer.
The given input data considered because:-
(i) it had the had purest anchor-positive pairs with semantically bound anchors with context/positive.
(ii) gave us the best result on final embedding fine tuning model.
The anchor-context(positive)-reference_answer data has been generated via Qwen-2.5-7B teacher model with temperature 0.1 and a strict system prompt.… See the full description on the dataset page: https://huggingface.co/datasets/vab46/Clinical_trials_anchor-contextORpositive-ground-truth_LLM_LORA_ft.clinical-trial-llmclinical-trial-llm-Open_condition_Cleaned_dup_NCT_IDclinical-trial-llm-cancerclinical-trial-llm-CAR-T-CELL_Mix_CANCER_Cleaned_dup_NCT_IDclinical-trial-llm-cancer-v3clinical-trial-llm-CAR-T-CELL_Mix_CANCER_v1clinical-trial-llm-1kclinical-trial-llm-cancer-v4clinical-trial-llm-CAR-T-CELL_Mix_CANCER_v1clinical-trial-llm-cancer-StdName-NCTID-v5ICD-10-LLM-generated-Synthetic-Clinical-Note-G00-G99
MedGemma ICD-10 Clinical Notes Dataset
Synthetic clinical notes (english) generated by MedGemma-4B-IT for fine-tuning ICD-10-CM diagnosis code prediction models. Focused on Chapter 6: Diseases of the Nervous System (G00-G99).
Dataset Summary
Split
Examples
Unique ICD-10 Codes
Train
3,325
665
Eval
250
50
Each example is a realistic clinical note paired with its ICD-10-CM diagnosis code, formatted as a chat conversation for instruction fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/singhankit16/ICD-10-LLM-generated-Synthetic-Clinical-Note-G00-G99.clinical-trial-llm-CAR-T-CELL-V1clinical-trial-llm-1k-v2clinical-trial-llm-cancer-v2clinical-trial-llm-cancer-restructure
