datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Identification_of_Somatic_Driver_Mutations_in_Indian_Oral_OSCC
Indian OSCC Somatic Driver Mutation Dataset
This repository contains the processed datasets used in the study: "Identification of Somatic Driver Mutations in Indian Oral Squamous Cell Carcinoma Using XGBoost and Integrative Genomic Features".
Files
train_MAF.csv – Somatic mutation data used for training
gene_variants.csv – Curated cancer driver gene list
ROH.csv – Runs of homozygosity intervals
SBS_Signature_HNSC.csv – Gene-level SBS13 annotation
synthetic_mutations.csv… See the full description on the dataset page: https://huggingface.co/datasets/aarushidas/Identification_of_Somatic_Driver_Mutations_in_Indian_Oral_OSCC.DRD2-mutationsAll possible mutations to a specific coding sequence of the human DRD2 gene, classified as missense or synonymous, PTVs are classed as missense.
Dataset can be used as a 1st pass evaluation of DNA language models, which should assign higher likelihoods/probabilities to synonymous mutations then to missense mutations.
psen1-mutations-alzforum
