datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
clinical-synthetic-text-llm
Data Description
We release the synthetic data generated using the method described in the paper Knowledge-Infused Prompting: Assessing and Advancing Clinical Text Data Generation with Large Language Models
(ACL 2024 Findings). The external knowledge we use is based on LLM-generated topics and writing styles.
Generated Datasets
The original train/validation/test data, and the generated synthetic training data are listed as follows. For each dataset, we generate 5000… See the full description on the dataset page: https://huggingface.co/datasets/ritaranx/clinical-synthetic-text-llm.clinical-trial-llm-v2Clinical_trials_anchor-contextORpositive-ground-truth_LLM_LORA-junk_handled_ft
Dataset details:-
This dataset is the 2nd iteration following bugs in 1st dataset.
The initial data suffered with followoing cases:-
(i) The failed reference_answers generation(due error totalling 23) primarly because of 2 reasons/exceptions:- (a) There was normal limit(300 in 1st request) and worst case limit(450 in 3rd request) number of tokens for consolidated 4 refernce_answers per chunk and its 4 corresponding answers. However certain answers breached this higher… See the full description on the dataset page: https://huggingface.co/datasets/vab46/Clinical_trials_anchor-contextORpositive-ground-truth_LLM_LORA-junk_handled_ft.Clinical_trials_anchor-contextORpositive-ground-truth_LLM_LORA_ft
Dataset details:-
This dataset is basically mapping of final anchor-positive pair data with their refernce answer.
The given input data considered because:-
(i) it had the had purest anchor-positive pairs with semantically bound anchors with context/positive.
(ii) gave us the best result on final embedding fine tuning model.
The anchor-context(positive)-reference_answer data has been generated via Qwen-2.5-7B teacher model with temperature 0.1 and a strict system prompt.… See the full description on the dataset page: https://huggingface.co/datasets/vab46/Clinical_trials_anchor-contextORpositive-ground-truth_LLM_LORA_ft.clinical-trial-llmllm-as-clinical-calculator
Augmentation of ChatGPT with clinician-informed tools improves performance on medical calculation tasks
Abstract: Prior work has shown that large language models (LLMs) have the ability to answer expert-level multiple choice questions in medicine, but are limited by both their tendency to hallucinate knowledge and their inherent inadequacy in performing basic mathematical operations. Unsurprisingly, early evidence suggests that LLMs perform poorly when asked to execute common… See the full description on the dataset page: https://huggingface.co/datasets/alexgoodell/llm-as-clinical-calculator.clinical-trial-llm-Open_condition_Cleaned_dup_NCT_IDclinical-trial-llm-cancerclinical-trial-llm-CAR-T-CELL_Mix_CANCER_Cleaned_dup_NCT_IDclinical-trial-llm-cancer-v3clinical-trial-llm-CAR-T-CELL_Mix_CANCER_v1clinical-trial-llm-1kclinical-trial-llm-cancer-v4clinical-trial-llm-CAR-T-CELL_Mix_CANCER_v1clinical-trial-llm-cancer-StdName-NCTID-v5ICD-10-LLM-generated-Synthetic-Clinical-Note-G00-G99
MedGemma ICD-10 Clinical Notes Dataset
Synthetic clinical notes (english) generated by MedGemma-4B-IT for fine-tuning ICD-10-CM diagnosis code prediction models. Focused on Chapter 6: Diseases of the Nervous System (G00-G99).
Dataset Summary
Split
Examples
Unique ICD-10 Codes
Train
3,325
665
Eval
250
50
Each example is a realistic clinical note paired with its ICD-10-CM diagnosis code, formatted as a chat conversation for instruction fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/singhankit16/ICD-10-LLM-generated-Synthetic-Clinical-Note-G00-G99.clinical-trial-llm-CAR-T-CELL-V1clinical-trial-llm-1k-v2clinical-trial-llm-cancer-v2clinical-trial-llm-cancer-restructure
