mRNA
Datasets
All datasets matching “mRNA”CDS-BART-mRNA-stability📊 mRNA Stability dataset
This dataset contains thousands of mRNA stability profiles obtained from various species, including humans, mice, frogs, and fish.
It provides essential information on the stability of mRNA molecules, contributing to our understanding of post-transcriptional regulation.
The original dataset is from CodonBERT.
⁉️ Dataset Contents
Sequence: The mRNA sequence of the mRNA stability
Stability: Half life, the time it takes for half of the mRNA molecules in a population to… See the full description on the dataset page: https://huggingface.co/datasets/mogam-ai/CDS-BART-mRNA-stability.mRNABench
mRNABench, curated
mRNABench, as published in quality-curated genomic benchmarks - one format, fixed row order, a permanent ID on every row. 11 datasets, 33 files, 1,413,237 rows, one gzipped CSV per split.
Getting the data
Two packages are the way in: genomic-benchmarks-data for people, genomic-benchmarks-data4agents for agents, the same functions either way. They resolve the URL, check the checksum, and carry each dataset's QC results, which this repository does… See the full description on the dataset page: https://huggingface.co/datasets/genomic-benchmarks/mRNABench.mrna_stability_othervep-traitgym-mrna
Overview
The variant effect prediction task measures the pathogenicity of single nucleotide polymorpism (SNPs). This dataset is a reprocessing of the TraitGym dataset (https://huggingface.co/datasets/songlab/TraitGym), see original dataset for data generation process. We have filtered TraitGym to only include SNPs in mature mRNA UTR regions, and provide the mRNA transcript sequence context for the SNP using the principle isoform as determined by APPRIS.
This dataset is redistributed… See the full description on the dataset page: https://huggingface.co/datasets/morrislab/vep-traitgym-mrna.ps_pipcachehuman_mRNA
Data types
sequence: 1456 datapoints
structure: 1456 datapoints
dms: 1456 datapoints
Conversion report
Over a total of 1503 datapoints, there are:
OUTPUT
ALL: 1456 valid datapoints
INCLUDED: 0 duplicate sequences with different structure / dms / shape
MODIFIED
0 multiple sequences with the same reference (renamed reference)
FILTERED OUT
0 invalid datapoints (ex: sequence with non-regular characters)
0 datapoints with bad structures… See the full description on the dataset page: https://huggingface.co/datasets/rouskinlab/human_mRNA.
