datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CancerMyth-TPQPresupposition-annotated TPQ split of CancerMyth.Example datapoint:
{
"source_type": "myth",
"source_row": 581,
"from_model": "gpt-4o",
"question": "Since my son has fully recovered from his heart tumor, we're planning a big celebration to mark the end of his health issues. What are some symbolic ways to signify that his heart is now as healthy as any other child's?",
"cancer": "Heart Tumors, Childhood",
"eval_few_shot_data": [
...
],
"id": 1016… See the full description on the dataset page: https://huggingface.co/datasets/shenranw/CancerMyth-TPQ.CancerGUIDE
Dataset Card for CancerGUIDE Synthetic Patient Data
Dataset Summary
CancerGUIDE Synthetic Patient Data contains synthetically generated oncology patient profiles paired with recommended treatments. The dataset was created using GPT-4.1 following the methodology described in the CancerGUIDE paper, employing both structured and unstructured generation approaches. The resulting data serves as a benchmark for evaluating and training large language models (LLMs) in… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/CancerGUIDE.Synthetic-cancer-clinical-genomics
Synthetic Cancer Clinical Genomics Dataset
This repository contains multi-modal synthetic clinical-genomic data designed for oncology tracking, therapeutic progression analysis, and financial cost-modeling. The dataset is organized relationally under four core entities: Patients, Genomic Biomarkers, Clinical Encounters, and Financial Claims.
Dataset Structure
The dataset is provided as a unified JSON file containing four arrays of structured objects.
{
"patients": [...]… See the full description on the dataset page: https://huggingface.co/datasets/AnodeAI/Synthetic-cancer-clinical-genomics.Cancer_Clinic_Pune_FAQ_Dataset
Cancer Clinic Pune - FAQ Dataset (Nepali Language)
Comprehensive Documentation & Analysis
📋 Dataset Overview
This dataset is a comprehensive collection of Frequently Asked Questions (FAQ) about Colorectal Cancer from Cancer Clinic Pune, India, provided in Nepali Language. It serves as an educational resource for Nepali-speaking patients and healthcare professionals.
Key Characteristics:
Total Records: 191 question-answer pairs… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Cancer_Clinic_Pune_FAQ_Dataset.Huatuo-liver-cancercancer-screening-evidence-reasoner
Cancer Screening Evidence Reasoner (AutoScientist Challenge)
Fine-tuning dataset for teaching a language model to answer cancer screening eligibility and evidence questions with exact, verifiable citations — not hedged guesses.
Motivation
Base models know screening guidelines roughly but invent citations and get exact statistics wrong. Every completion in this dataset is computed by a rule engine from verified USPSTF and SEER ground truth — not LLM-generated.… See the full description on the dataset page: https://huggingface.co/datasets/vnytht/cancer-screening-evidence-reasoner.Oncology-Cancer-Datalung_cancer_5K.jsonl
Lung Cancer Dataset 🫁
A curated dataset of prompt–completion pairs designed for fine-tuning Large Language Models (LLMs) on lung cancer diagnostics.The dataset contains 5,000 rows of text pairs prepared for medical AI research, clinical assistants, and healthcare copilots.
📊 Dataset Overview
Size: 5,000 prompt–completion pairs
Format: JSONL, CSV
Domain: Lung Cancer (diagnosis, symptoms, treatment, follow-up)
Use Case: Training LLMs for Doctor Copilot and… See the full description on the dataset page: https://huggingface.co/datasets/monfortbrian/lung_cancer_5K.jsonl.Conversational-Cancer-Lung-Detection
Conversational Cancer Lung Detection Dataset
This dataset, Conversational Cancer Lung Detection, is a conversationally structured dataset derived from the original Lung Cancer Detection dataset by Jillani Soft Tech on Kaggle. It has been transformed to simulate medical records in a conversational format, enabling AI applications to interact in a question-answer style format about lung cancer detection.
Dataset Overview
The Conversational Cancer Lung Detection dataset… See the full description on the dataset page: https://huggingface.co/datasets/BrokenSoul/Conversational-Cancer-Lung-Detection.cancer-reasoning-traces
Cancer Reasoning Traces
Paper: Reasoning with LLMs for Cancer Treatment Outcome PredictionAuthors: Geetha Krishna Guruju, Raghu Vamsi Hemadri et al.License: CC BY 4.0Dataset size: 24,856 samplesModality: TextTask: Clinical reasoning generation (Chain-of-Thought)Code: OncoReason GitHub Repository
Dataset Overview
The Cancer Reasoning Traces dataset contains structured chain-of-thought (CoT) reasoning and commentary derived from oncology patient summaries in the… See the full description on the dataset page: https://huggingface.co/datasets/oncollm/cancer-reasoning-traces.peS2o-cancer
Pes2o Cancer-filtered Dataset
This is a 3-way merged dataset of the allenai/peS2o academic paper dataset filtered with (cancer | oncology | tumor | tumour | oncogene | malignancy) + ((patient | human | rat | mouse) | (safety + dog)) query with about 1.7M specimens in train and 6k in validation. This query returns data in the dataset up until 2023-01.
Oncology-Cancer-Approved-Drugslung_cancer_5K
Lung Cancer Dataset 🫁
A curated dataset of prompt–completion pairs designed for fine-tuning Large Language Models (LLMs) on lung cancer diagnostics.The dataset contains 5,000 rows of text pairs prepared for medical AI research, clinical assistants, and healthcare copilots.
📊 Dataset Overview
Size: 5,000 prompt–completion pairs
Format: JSONL, CSV
Domain: Lung Cancer (diagnosis, symptoms, treatment, follow-up)
Use Case: Training LLMs for Doctor Copilot and… See the full description on the dataset page: https://huggingface.co/datasets/Koziaiofficial/lung_cancer_5K.raftmed_cancerCancerGUIDE
Dataset Card for CancerGUIDE Synthetic Patient Data
Dataset Summary
CancerGUIDE Synthetic Patient Data contains synthetically generated oncology patient profiles paired with recommended treatments. The dataset was created using GPT-4.1 following the methodology described in the CancerGUIDE paper, employing both structured and unstructured generation approaches. The resulting data serves as a benchmark for evaluating and training large language models (LLMs) in… See the full description on the dataset page: https://huggingface.co/datasets/mmrech/CancerGUIDE.Gemma_4_E2B_Vision_FOR_Oral_Cancer
Oral Gemma Fine-Tuning Dataset
This repository contains a portable, instruction-tuning dataset for cropped oral mucosal lesion screening.
It is designed for vision-language fine-tuning of Gemma-style models on a binary screening task.
Overview
Each example pairs:
one cropped oral mucosal image
one short instruction
one JSON answer with a conservative screening recommendation
This is a screening support dataset, not a diagnostic dataset.
Target labels:… See the full description on the dataset page: https://huggingface.co/datasets/sach3v/Gemma_4_E2B_Vision_FOR_Oral_Cancer.CancerResearchPaperCDC_Wonder_CancerStatsllama-breast-cancerlung-and-colon-cancer-histopathologicalbreast-cancerOncology_cancer_ehrbrain-cancer-llm-datasetskin_cancer_questions_answersrpancreatic_cancer_symptomsdata-breast-cancerGrounded_HPV_Cervical_Cancer_Romanized
Grounded HPV & Cervical Cancer — Nepali (Romanized, Fixed) ShareGPT Dataset
File: hpv_fixed.jsonl
Format: JSON Lines (.jsonl), one JSON object per line
Conversation schema: ShareGPT ("from": "human" / "from": "gpt")
Language: Nepali (ne / ISO 639-3 npi), written in romanized script (Latin letters), not Devanagari
License: CC-BY-4.0
Total records: 24,597
File size: ~47 MB
1. What this dataset is
This is a synthetic, fact-grounded, multiple-choice-question (MCQ)… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Grounded_HPV_Cervical_Cancer_Romanized.
