datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SQL-Queries-Datasetclinvar_variant_summarysource data from https://ftp.ncbi.nlm.nih.gov/pub/clinvar/tab_delimited/variant_summary.txt.gz
TCGA-Cancer-Variant-and-Clinical-Data
TCGA Cancer Variant and Clinical Data
Dataset Description
This dataset combines genetic variant information at the protein level with clinical data from The Cancer Genome Atlas (TCGA) project, curated by the International Cancer Genome Consortium (ICGC). It provides a comprehensive view of protein-altering mutations and clinical characteristics across various cancer types.
Dataset Summary
The dataset includes:
Protein sequence data for both mutated and… See the full description on the dataset page: https://huggingface.co/datasets/seq-to-pheno/TCGA-Cancer-Variant-and-Clinical-Data.Car-Price-DatasetFrench_Wolof_Various_Parallel_Corpuscommon-variety-d04dad
common-variety-d04dad
Synthetic products test data: 45 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/KarenSmith/common-variety-d04dad.bigcodebench-typo-variants
BigCodeBench Typo Variants
This dataset contains typo-injected variants of the BigCodeBench coding benchmark to evaluate the robustness of code generation models to typographical errors in problem descriptions.
Dataset Description
BigCodeBench is a benchmark for evaluating large language models on diverse and challenging coding tasks. This dataset provides 7 variants with different levels of typos injected into the instruction prompts:
Original (0% typos): Clean baseline… See the full description on the dataset page: https://huggingface.co/datasets/jeqcho/bigcodebench-typo-variants.Telugu_EmotionDo cite the below reference for using the dataset:
@article{marreddy2022resource, title={Am I a Resource-Poor Language? Data Sets, Embeddings, Models and Analysis for four different NLP tasks in Telugu Language},
author={Marreddy, Mounika and Oota, Subba Reddy and Vakada, Lakshmi Sireesha and Chinni, Venkata Charan and Mamidi, Radhika},
journal={Transactions on Asian and Low-Resource Language Information Processing}, publisher={ACM New York, NY} }
If you want to use the four classes (angry… See the full description on the dataset page: https://huggingface.co/datasets/vardhan28/Telugu_Emotion.clinical-quad-signal-detection-drift-ae-coding-variance-unblinding-risk-dsmb-decision-delay-v0.1
Clinical Quad: Signal Detection Drift × AE Coding Variance × Unblinding Risk × DSMB Decision Delay
This dataset targets safety governance collapse.
Signals weaken or shift.AE coding diverges across sites.Unblinding pressure rises.The DSMB response slows.
The quad can turn a manageable safety issue into a governance failure.
Variables
signal_detection_drift (low | medium | high)
ae_coding_variance (low | medium | high)
unblinding_risk (low | medium | high)… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-signal-detection-drift-ae-coding-variance-unblinding-risk-dsmb-decision-delay-v0.1.plant-variety-database
Plant Variety Database
An open dataset that joins cultivar-level seed-catalog data with USDA hardiness zones and per-zone monthly planting calendars — 1,972 varieties × 13 zones × 12 months, fully sourced, CC BY 4.0.
The hero rows aren't the 1,972 varieties (USDA PLANTS already has ~98K species). They're the joins:
20,728 variety × zone planting-calendar entries (indoor sow / transplant / direct sow / harvest windows)
21,880 companion-plant pairings with relationship and reason
2… See the full description on the dataset page: https://huggingface.co/datasets/WindRiverGreens/plant-variety-database.clinical-quad-enrollment-protocol-deviation-site-variance-endpoint-integrity-v0.1
Clinical Quad Enrollment–Protocol Deviations–Site Variance–Endpoint Integrity v0.1
What this is
A quad-coupling dataset for trial collapse driven by the interaction of:
enrollment pattern changes
rising protocol deviations
site-to-site variance
endpoint integrity degradation
Task
Input: one quad state rowOutput: label
0 — Stable1 — Drift2 — Collapse
Why it matters
Trials often fail through operational pressure:
recruitment becomes spiky or slow… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-enrollment-protocol-deviation-site-variance-endpoint-integrity-v0.1.clinical-quad-enrollment-protocol-deviation-site-variance-endpoint-integrity-v0.2Clinical Quad Enrollment Protocol Deviation Site Variance Endpoint Integrity v0.2
What this dataset does
It tests whether a model can detect when endpoint integrity degrades under four coupled operational pressures.
Quad nodes
enrollment_pattern
protocol_deviation_rate
site_variance_level
endpoint_integrity
Labels
0 coherent
endpoints clean
enrollment stable
deviations not high
site variance not high
1 tradeoff
strain exists
endpoint softens or system drifts… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-enrollment-protocol-deviation-site-variance-endpoint-integrity-v0.2.clinical-quad-pk-sampling-window-deviation-bioanalytical-variance-dose-adjustment-interim-v0.1Clarus Clinical Quad Coupling PK Integrity v0.1
PurposeDetect PK integrity distortion driven by four interacting nodes.
Quad nodes
Sampling window deviation
Bioanalytical or stability variance
Dose adjustment decisions
Governance interim or submission timing
InputOne vignette.
OutputStrict JSON only.
Required keys
pk_integrity_risk
risk_type
driver_nodes
recommended_action
action_detail
rationale
confidence
Filesdata/train.csvdata/test.csvscorer.py
Run scoringCreate… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-pk-sampling-window-deviation-bioanalytical-variance-dose-adjustment-interim-v0.1.clinical-quad-population-shift-protocol-deviation-site-variance-endpoint-fragility-v0.1
Clinical Quad Population Shift × Protocol Deviation × Site Variance × Endpoint Fragility v0.1
What this is
A quad-coupling dataset for trial collapse that happens when:
the enrolled population drifts from the intended cohort
protocol deviations rise
site-to-site variance widens
the primary endpoint is fragile to measurement or baseline imbalance
Task
Input: one row describing the quad stateOutput: label
0 — Stable1 — Drift2 — Collapse
Why it matters… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-population-shift-protocol-deviation-site-variance-endpoint-fragility-v0.1.legal-costs-budget-phase-scope-variance-coherence-risk-v0.1What this dataset does
You receive
budget by phase
assumptions
actual work
variance notes
client updates
approval status
You decide
coherent
or
incoherent
Daily use
overspend early warning
client surprise risk
variance justification QC
clinical-quad-protocol-deviation-staffing-drift-adjudication-variance-missingness-bias-v0.1Clarus Clinical Quad Coupling Protocol Deviation Staffing Drift Adjudication Variance Missingness Bias v0.1
What this dataset isThis dataset tests whether a model can detect protocol deviation events driven by quad coupling.
Quad coupling nodes
Operational staffing drift or site capacity constraint
Protocol compliance breakdown
Endpoint adjudication variance or bias risk
Data missingness that distorts safety or efficacy interpretation under governance rules
Input
One vignette in… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-protocol-deviation-staffing-drift-adjudication-variance-missingness-bias-v0.1.TP53_protein_variants
Dataset Summary
This dataset will help you to develop a machine learning-based model to predict the pathogenic variants (Positive labels) by utilizing their amino acid sequences.
Used as an example to benchmark biomerida as part of the Bio-Hakathon Mena region
causal-inference-variant-interpretation-genomics-v01
Dataset
ClarusC64/causal-inference-variant-interpretation-genomics-v01
This dataset tests one capability.
Can a model distinguish association from causation when interpreting genetic variants.
Core rule
Genomic evidence has tiers.
A claim must respect
evidence strength
effect size
penetrance
inheritance logic
Association does not equal causation.
Risk does not equal destiny.
Uncertain does not equal pathogenic.
Canonical labels
WITHIN_SCOPE… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/causal-inference-variant-interpretation-genomics-v01.clinical-quad-trial-pop-variance-realworld-subgroup-signal-generalization-claim-drift-v0.1What this repo does
This dataset models population mismatch narrative drift in clinical trial reporting. It predicts when the interaction between trial population variance, real-world variance, subgroup signal strength, and generalization claim intensity indicates that narrative claims extend beyond what the data supports.
Core quad
trial_population_variance_index
real_world_variance_index
subgroup_signal_strength_index
generalization_claim_index
Prediction target
label_claim_drift
Row… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-trial-pop-variance-realworld-subgroup-signal-generalization-claim-drift-v0.1.Korean-Corpus-From-Various-Task-1es-paremias-variantesParemias con sus variantes para entrenar modelo de embeddings.
Obtenidos de web: https://cvc.cervantes.es/lengua/refranero/listado.aspx
prompt-variations
Prompt Variations and LLM Responses
Prompt variants and model responses used to evaluate the
Stability-Generalization Score (SGS) across eleven LLMs (eight
open-source + three closed-source) on six QA / instruction benchmarks
under six families of stylistic perturbations.
Splits
split
rows
source dataset
truthful_qa
99,888
TruthfulQA
natural_questions
41,040
Natural Questions
alpaca
13,872
Alpaca
simpleqa_verified
13,872
SimpleQA Verified… See the full description on the dataset page: https://huggingface.co/datasets/naghamo/prompt-variations.NIST-SP-800-171r2-Context-VarianceThis dataset contains the requirements found in Chapter 3 of the NIST SP 800-171 document titled Protecting Controlled Unclassified Information in Nonfederal Systems and Organizations. The context of each row has information related to the expected response.
Variants_IOB
Variants IOB files
Train/Test/Dev/All IOB files for variants in text
Data produced using this data as a starting point
Original source data came from the paper "Dataset from a human-in-the-loop approach to identify functionally important protein residues from literature.", and was filtered into a "light"-version focusing only on variants / mutants. Our motivation here being that we hope to be able to use this data for the training of a NER model, tagging variants in text.
alzheimers-variant-tutorial-data
alzheimers-variant-tutorial-data
Dataset Summary
This dataset contains summary statistics for 1,000 genomic variants associated with Alzheimer's disease. Each row represents a single-nucleotide polymorphism (SNP) mapped to the hg19 reference genome.
Dataset Structure
Number of variants: 1,000
Genome build: hg19
Data Fields
Based on the header of variants.csv:
Column
Type
Description
snpid
string
Unique identifier in chr:pos_ref_alt… See the full description on the dataset page: https://huggingface.co/datasets/Genentech/alzheimers-variant-tutorial-data.linical-quad-protocol-deviation-staffing-drift-adjudication-variance-missingness-bias-v0.1
Clinical Quad: Protocol Deviations × Staffing Drift × Adjudication Variance × Missingness Bias
This dataset targets a “trial looks clean on paper” failure mode.
Sites drift in staffing.Protocol deviations rise.Endpoint adjudication becomes inconsistent.Missing data stops being random.
The four-way coupling can create false stability or false efficacy.
Variables
protocol_deviation (low | medium | high)
staffing_drift (yes | no)
adjudication_variance (low | high)… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/linical-quad-protocol-deviation-staffing-drift-adjudication-variance-missingness-bias-v0.1.clinical-quad-population-shift-protocol-deviation-site-variance-endpoint-fragility-v0.2Clinical Quad Population Shift Protocol Deviation Site Variance Endpoint Fragility v0.2
What this dataset does
It tests whether a model can detect when clinical trial endpoints lose credibility under quad coupling.
Quad nodes
population_shift
protocol_deviation_rate
site_variance_level
endpoint_fragility
Labels
0 coherent
Stable population
Low deviations
Low site variance
Endpoint robust
1 tradeoff
Some drift exists
Endpoint still usable
Risk is present but not… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-population-shift-protocol-deviation-site-variance-endpoint-fragility-v0.2.BIG-IDEAs-Lab-Glycemic-Variability-and-Wearable-Device-Dataes-paremias-variantes-antonimosAmpliación del dataset es-paremias-variantes con la columna "Frases Antonimas" que pretende ser una frase que tiene el significado completamente opuesto al de la frase variante.
❗Esta columna se ha generado de forma automática utilizando el modelo Phi-4 a través de LMStudio.
La inspección manual de algunos ejemplos valida que sea una frase que mantiene el significado opuesto, pero no todos los ejemplos han sido revisados.
Sientete libre de abrir pull-request con mejoras a este dataset ✨workout-routine
