datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
clinvar_variant_summarysource data from https://ftp.ncbi.nlm.nih.gov/pub/clinvar/tab_delimited/variant_summary.txt.gz
TCGA-Cancer-Variant-and-Clinical-Data
TCGA Cancer Variant and Clinical Data
Dataset Description
This dataset combines genetic variant information at the protein level with clinical data from The Cancer Genome Atlas (TCGA) project, curated by the International Cancer Genome Consortium (ICGC). It provides a comprehensive view of protein-altering mutations and clinical characteristics across various cancer types.
Dataset Summary
The dataset includes:
Protein sequence data for both mutated and… See the full description on the dataset page: https://huggingface.co/datasets/seq-to-pheno/TCGA-Cancer-Variant-and-Clinical-Data.Car-Price-Datasetcommon-variety-d04dad
common-variety-d04dad
Synthetic products test data: 45 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/KarenSmith/common-variety-d04dad.plant-variety-database
Plant Variety Database
An open dataset that joins cultivar-level seed-catalog data with USDA hardiness zones and per-zone monthly planting calendars — 1,972 varieties × 13 zones × 12 months, fully sourced, CC BY 4.0.
The hero rows aren't the 1,972 varieties (USDA PLANTS already has ~98K species). They're the joins:
20,728 variety × zone planting-calendar entries (indoor sow / transplant / direct sow / harvest windows)
21,880 companion-plant pairings with relationship and reason
2… See the full description on the dataset page: https://huggingface.co/datasets/WindRiverGreens/plant-variety-database.clinical-quad-enrollment-protocol-deviation-site-variance-endpoint-integrity-v0.1
Clinical Quad Enrollment–Protocol Deviations–Site Variance–Endpoint Integrity v0.1
What this is
A quad-coupling dataset for trial collapse driven by the interaction of:
enrollment pattern changes
rising protocol deviations
site-to-site variance
endpoint integrity degradation
Task
Input: one quad state rowOutput: label
0 — Stable1 — Drift2 — Collapse
Why it matters
Trials often fail through operational pressure:
recruitment becomes spiky or slow… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-enrollment-protocol-deviation-site-variance-endpoint-integrity-v0.1.clinical-quad-enrollment-protocol-deviation-site-variance-endpoint-integrity-v0.2Clinical Quad Enrollment Protocol Deviation Site Variance Endpoint Integrity v0.2
What this dataset does
It tests whether a model can detect when endpoint integrity degrades under four coupled operational pressures.
Quad nodes
enrollment_pattern
protocol_deviation_rate
site_variance_level
endpoint_integrity
Labels
0 coherent
endpoints clean
enrollment stable
deviations not high
site variance not high
1 tradeoff
strain exists
endpoint softens or system drifts… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-enrollment-protocol-deviation-site-variance-endpoint-integrity-v0.2.clinical-quad-population-shift-protocol-deviation-site-variance-endpoint-fragility-v0.1
Clinical Quad Population Shift × Protocol Deviation × Site Variance × Endpoint Fragility v0.1
What this is
A quad-coupling dataset for trial collapse that happens when:
the enrolled population drifts from the intended cohort
protocol deviations rise
site-to-site variance widens
the primary endpoint is fragile to measurement or baseline imbalance
Task
Input: one row describing the quad stateOutput: label
0 — Stable1 — Drift2 — Collapse
Why it matters… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-population-shift-protocol-deviation-site-variance-endpoint-fragility-v0.1.legal-costs-budget-phase-scope-variance-coherence-risk-v0.1What this dataset does
You receive
budget by phase
assumptions
actual work
variance notes
client updates
approval status
You decide
coherent
or
incoherent
Daily use
overspend early warning
client surprise risk
variance justification QC
clinical-quad-trial-pop-variance-realworld-subgroup-signal-generalization-claim-drift-v0.1What this repo does
This dataset models population mismatch narrative drift in clinical trial reporting. It predicts when the interaction between trial population variance, real-world variance, subgroup signal strength, and generalization claim intensity indicates that narrative claims extend beyond what the data supports.
Core quad
trial_population_variance_index
real_world_variance_index
subgroup_signal_strength_index
generalization_claim_index
Prediction target
label_claim_drift
Row… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-trial-pop-variance-realworld-subgroup-signal-generalization-claim-drift-v0.1.alzheimers-variant-tutorial-data
alzheimers-variant-tutorial-data
Dataset Summary
This dataset contains summary statistics for 1,000 genomic variants associated with Alzheimer's disease. Each row represents a single-nucleotide polymorphism (SNP) mapped to the hg19 reference genome.
Dataset Structure
Number of variants: 1,000
Genome build: hg19
Data Fields
Based on the header of variants.csv:
Column
Type
Description
snpid
string
Unique identifier in chr:pos_ref_alt… See the full description on the dataset page: https://huggingface.co/datasets/Genentech/alzheimers-variant-tutorial-data.clinical-quad-population-shift-protocol-deviation-site-variance-endpoint-fragility-v0.2Clinical Quad Population Shift Protocol Deviation Site Variance Endpoint Fragility v0.2
What this dataset does
It tests whether a model can detect when clinical trial endpoints lose credibility under quad coupling.
Quad nodes
population_shift
protocol_deviation_rate
site_variance_level
endpoint_fragility
Labels
0 coherent
Stable population
Low deviations
Low site variance
Endpoint robust
1 tradeoff
Some drift exists
Endpoint still usable
Risk is present but not… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-population-shift-protocol-deviation-site-variance-endpoint-fragility-v0.2.workout-routinesars-cov-2-variant-proportions
SARS-CoV-2 Variant Proportions
Description
To identify and track SARS-CoV-2 variants, CDC uses genomic surveillance. CDC's national genomic surveillance system collects SARS-CoV-2 specimens for sequencing through the National SARS-CoV-2 Strain Surveillance (NS3) program, as well as SARS-CoV-2 sequences generated by commercial or academic laboratories contracted by CDC and state or local public health laboratories. Viral genomic sequences are analyzed and classified as a… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/sars-cov-2-variant-proportions.BIG-IDEAs-Lab-Glycemic-Variability-and-Wearable-Device-Datamissense-variant-features
missense-variant-features
Benchmark dataset for missense genetic variant analysis and pathogenic classification (CFTR, PAH, Cancer).
clinical-quad-baseline-risk-adherence-pk-variance-genetic-factor-response-loss-v0.1What this repo does
This dataset models efficacy response divergence as a basin exit event in patient response space. It predicts when the interaction between baseline risk, adherence, pharmacokinetic variance, and genetic response factor increases the probability of response loss under treatment.
Core quad
baseline_risk_score
adherence_index
pk_variance_index
genetic_response_factor
Prediction target
label_response_loss
Row structure
Each row represents a patient response state snapshot… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-baseline-risk-adherence-pk-variance-genetic-factor-response-loss-v0.1.clinical-quad-dose-adherence-exposure-variability-outcome-failure-v0.1Clinical Quad Biomarker Subpopulation Dose Endpoint Drift v0.1
Each row is a patient state snapshot.
Core quad
Biomarker statusSubpopulation flagDosing strategyEndpoint drift
Target
label_signal_fail_next_30d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
clinical-quad-dispense-delay-visit-reschedule-half-life-adherence-exposure-variance-v0.1What this repo does
This dataset models exposure variance risk driven by dose timing instability in clinical trials. It predicts when the interaction between dispense delays, visit rescheduling, drug half-life, and patient adherence creates clinically meaningful pharmacokinetic exposure variability.
Core quad
dispense_delay_days
visit_reschedule_count
drug_half_life_hr
patient_adherence_index
Prediction target
label_exposure_variance
Row structure
Each row represents a dosing cycle timing… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-dispense-delay-visit-reschedule-half-life-adherence-exposure-variance-v0.1.clinical-quad-batch-potency-exposure-variance-outcome-noise-v0.1Clinical Quad Batch Potency Exposure Variance Outcome Noise v0.1
Each row is a patient snapshot tied to manufacturing batch.
Core quad
Batch potencyStorage conditionsExposure varianceOutcome noise
Target
label_signal_loss_next_30d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
missense-variant-effects
Genetic Variant Pathogenicity Dataset
Dataset Description
This dataset contains annotated genetic variants (mutations) designed for tabular binary classification tasks. The objective is to predict whether a given genetic variant is Pathogenic (disease-causing) or Benign (harmless) based on a rich set of bioinformatics annotations, evolutionary conservation scores, and functional prediction tools.
Task: Binary Classification
Target Column: Pathologic/Benign… See the full description on the dataset page: https://huggingface.co/datasets/bedylmz/missense-variant-effects.clinical-quad-device-change-measurement-drift-training-variance-endpoint-noise-v0.1Clinical Quad Device Change Measurement Drift Training Variance Endpoint Noise v0.1
Each row is a site monthly snapshot.
Core quad
Device changeMeasurement driftTraining varianceEndpoint noise
Target
label_primary_fail_next_90d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
This dataset identifies a measurable coupling pattern associated with systemic instability.
The sample demonstrates the… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-device-change-measurement-drift-training-variance-endpoint-noise-v0.1.f1-quad-pit-duration-crew-fatigue-race-pressure-weather-variability-pit-error-v0.1What this repo does
This dataset models pit stop failure risk in Formula One. It predicts when the interaction between stop duration strain, cumulative crew fatigue, race pressure intensity, and weather variability produces elevated probability of pit execution error.
Core quad
pit_duration_s
crew_fatigue_index
race_pressure_index
weather_variability_index
Prediction target
label_pit_error
Binary forward label predicting unsafe release, delayed wheel fit, or procedural error during the pit… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/f1-quad-pit-duration-crew-fatigue-race-pressure-weather-variability-pit-error-v0.1.TCGA-Cancer-Variant-and-Clinical-Data
TCGA Cancer Variant and Clinical Data
Dataset Description
This dataset combines genetic variant information at the protein level with clinical data from The Cancer Genome Atlas (TCGA) project, curated by the International Cancer Genome Consortium (ICGC). It provides a comprehensive view of protein-altering mutations and clinical characteristics across various cancer types.
Dataset Summary
The dataset includes:
Protein sequence data for both mutated and… See the full description on the dataset page: https://huggingface.co/datasets/hammad655/TCGA-Cancer-Variant-and-Clinical-Data.nndss-table-ii-tetanus-to-varicella
NNDSS - Table II. Tetanus to Varicella
Description
NNDSS - Table II. Tetanus to Varicella - 2018. In this Table, provisional cases of selected notifiable diseases (≥1,000 cases reported during the preceding year), and selected low frequency diseases are displayed. The Table includes total number of cases reported in the United States, by region and by states or territory.
Note:
This table contains provisional cases of selected national notifiable diseases from the… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/nndss-table-ii-tetanus-to-varicella.PredictiveMaintenanceclinical-quad-adjudication-drift-rater-variability-missingness-bias-effect-attenuation-v0.1Clinical Quad Adjudication Drift Rater Variability Missingness Bias Effect Attenuation v0.1
Each row is a regional monthly snapshot.
Core quad
Adjudication driftRater variabilityMissingness biasEffect attenuation
Target
label_primary_fail_next_90d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
This dataset identifies a measurable coupling pattern associated with systemic instability.
The sample… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-adjudication-drift-rater-variability-missingness-bias-effect-attenuation-v0.1.MPRA_VarCREsuperkart-sales-datademo-data
