datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
archiveii
ArchiveII
ArchiveII is a dataset of RNA sequences and their secondary structures, widely used in RNA secondary structure prediction benchmarks.
ArchiveII contains 2975 RNA samples across 10 RNA families, with sequence lengths ranging from 28 to 2968 nucleotides.
This dataset is frequently used to evaluate RNA secondary structure prediction methods, including those that handle both pseudoknotted and non-pseudoknotted structures.
It is considered complementary to the RNAStrAlign… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/archiveii.rnacentral
RNAcentral
RNAcentral is a free, public resource that offers integrated access to a comprehensive and up-to-date set of non-coding RNA sequences provided by a collaborating group of Expert Databases representing a broad range of organisms and RNA types.
The development of RNAcentral is coordinated by European Bioinformatics Institute and is supported by Wellcome. Initial funding was provided by BBSRC.
Disclaimer
This is an UNOFFICIAL release of the RNAcentral by The… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/rnacentral.eternabench-switch
EternaBench-Switch
EternaBench-Switch is a synthetic RNA dataset consisting of 7,228 riboswitch constructs, designed to explore the structural behavior of RNA molecules that change conformation upon binding to ligands such as FMN, theophylline, or tryptophan.
These riboswitches exhibit different structural states in the presence or absence of their ligands, and the dataset includes detailed measurements of binding affinities (dissociation constants), activation ratios, and RNA… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/eternabench-switch.eternabench-cm
EternaBench-CM
EternaBench-CM is a synthetic RNA dataset comprising 12,711 RNA constructs that have been chemically mapped using SHAPE and MAP-seq methods.
These RNA sequences are probed to obtain experimental data on their nucleotide reactivity, which indicates whether specific regions of the RNA are flexible or structured.
The dataset provides high-resolution, large-scale data that can be used for studying RNA folding and stability.
Disclaimer
This is an UNOFFICIAL… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/eternabench-cm.eternabench-external.300
EternaBench-External
EternaBench-External consists of 31 independent RNA datasets from various biological sources, including viral genomes, mRNAs, and synthetic RNAs.
These sequences were probed using techniques such as SHAPE-CE, SHAPE-MaP, and DMS-MaP-seq to understand RNA secondary structures under different experimental and biological conditions.
This dataset serves as a benchmark for evaluating RNA structure prediction models, with a particular focus on generalization to… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/eternabench-external.300.bprna
bpRNA-1m
bpRNA-1m is a database of single molecule secondary structures annotated using bpRNA.
Disclaimer
This is an UNOFFICIAL release of the bpRNA-1m by Center for Quantitative Life Sciences of the Oregon State University.
The team releasing bpRNA did not write this dataset card for this dataset so this dataset card has been written by the MultiMolecule team.
Example Entry
id
sequence
secondary_structure
structural_annotation
functional_annotation… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/bprna.bprna-spot
bpRNA-spot
bpRNA-spot is a collection of the datasets used by SPOT-RNA for RNA secondary structure prediction.
The dataset is released as a composite repository, bpRNA-spot, and three numbered component repositories:
bpRNA-spot-0: the initial bpRNA split, TR0, VL0, and TS0.
bpRNA-spot-1: the PDB transfer-learning split, TR1, VL1, and TS1.
bpRNA-spot-2: the NMR-only evaluation split, TS2.
bpRNA-spot concatenates the components in order:
train: TR0 + TR1
validation: VL0 + VL1
test:… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/bprna-spot.bprna-new
bpRNA-new
bpRNA-new is a database of single molecule secondary structures annotated using bpRNA.
bpRNA-new is a dataset of RNA families from Rfam 14.2, designed for cross-family validation to assess generalization capability.
It focuses on families distinct from those in bpRNA-1m, providing a robust benchmark for evaluating model performance on unseen RNA families.
Disclaimer
This is an UNOFFICIAL release of the bpRNA-new by Kengo Sato, et al.
The team releasing… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/bprna-new.bprna-90
bpRNA-1m
bpRNA-1m is a database of single molecule secondary structures annotated using bpRNA.
Disclaimer
This is an UNOFFICIAL release of the bpRNA-1m by Center for Quantitative Life Sciences of the Oregon State University.
The team releasing bpRNA did not write this dataset card for this dataset so this dataset card has been written by the MultiMolecule team.
Example Entry
id
sequence
secondary_structure
structural_annotation
functional_annotation… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/bprna-90.eternabench-external.1200
EternaBench-External
EternaBench-External consists of 31 independent RNA datasets from various biological sources, including viral genomes, mRNAs, and synthetic RNAs.
These sequences were probed using techniques such as SHAPE-CE, SHAPE-MaP, and DMS-MaP-seq to understand RNA secondary structures under different experimental and biological conditions.
This dataset serves as a benchmark for evaluating RNA structure prediction models, with a particular focus on generalization to… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/eternabench-external.1200.gencode-human
GENCODE
GENCODE is a comprehensive annotation project that aims to provide high-quality annotations of the human and mouse genomes.
The project is part of the ENCODE (ENCyclopedia Of DNA Elements) scale-up project, which seeks to identify all functional elements in the human genome.
Disclaimer
This is an UNOFFICIAL release of the GENCODE by Paul Flicek, Roderic Guigo, Manolis Kellis, Mark Gerstein, Benedict Paten, Michael Tress, Jyoti Choudhary, et al.
The team… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/gencode-human.rnastralign
RNAStrAlign
RNAStrAlign is a comprehensive dataset of RNA sequences and their secondary structures.
RNAStrAlign aggregates data from multiple established RNA structure repositories, covering diverse RNA families such as 5S ribosomal RNA, tRNA, and group I introns.
It is considered complementary to the ArchiveII dataset.
Disclaimer
This is an UNOFFICIAL release of the RNAStrAlign by Zhen Tan, et al.
The team releasing RNAStrAlign did not write this dataset card for… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/rnastralign.rivas
RIVAS
The RIVAS dataset is a curated collection of RNA sequences and their secondary structures, designed for training and evaluating RNA secondary structure prediction methods.
The dataset combines sequences from published studies and databases like Rfam, covering diverse RNA families such as tRNA, SRP RNA, and ribozymes.
The secondary structure data is obtained from experimentally verified structures and consensus structures from Rfam alignments, ensuring high-quality annotations… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/rivas.eternabench-external.600
EternaBench-External
EternaBench-External consists of 31 independent RNA datasets from various biological sources, including viral genomes, mRNAs, and synthetic RNAs.
These sequences were probed using techniques such as SHAPE-CE, SHAPE-MaP, and DMS-MaP-seq to understand RNA secondary structures under different experimental and biological conditions.
This dataset serves as a benchmark for evaluating RNA structure prediction models, with a particular focus on generalization to… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/eternabench-external.600.eternabench-external.900
EternaBench-External
EternaBench-External consists of 31 independent RNA datasets from various biological sources, including viral genomes, mRNAs, and synthetic RNAs.
These sequences were probed using techniques such as SHAPE-CE, SHAPE-MaP, and DMS-MaP-seq to understand RNA secondary structures under different experimental and biological conditions.
This dataset serves as a benchmark for evaluating RNA structure prediction models, with a particular focus on generalization to… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/eternabench-external.900.rfam
Rfam
Rfam is a database of structure-annotated multiple sequence alignments, covariance models and family annotation for a number of non-coding RNA, cis-regulatory and self-splicing intron families.
The seed alignments are hand curated and aligned using available sequence and structure data, and covariance models are built from these alignments using the INFERNAL v1.1.4 software suite.
The full regions list is created by searching the RFAMSEQ database using the covariance model… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/rfam.gencode-mouse
GENCODE
GENCODE is a comprehensive annotation project that aims to provide high-quality annotations of the human and mouse genomes.
The project is part of the ENCODE (ENCyclopedia Of DNA Elements) scale-up project, which seeks to identify all functional elements in the human genome.
Disclaimer
This is an UNOFFICIAL release of the GENCODE by Paul Flicek, Roderic Guigo, Manolis Kellis, Mark Gerstein, Benedict Paten, Michael Tress, Jyoti Choudhary, et al.
The team… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/gencode-mouse.rnacentral.2048
RNAcentral
RNAcentral is a free, public resource that offers integrated access to a comprehensive and up-to-date set of non-coding RNA sequences provided by a collaborating group of Expert Databases representing a broad range of organisms and RNA types.
The development of RNAcentral is coordinated by European Bioinformatics Institute and is supported by Wellcome. Initial funding was provided by BBSRC.
Disclaimer
This is an UNOFFICIAL release of the RNAcentral by The… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/rnacentral.2048.archiveii.512
ArchiveII
ArchiveII is a dataset of RNA sequences and their secondary structures, widely used in RNA secondary structure prediction benchmarks.
ArchiveII contains 2975 RNA samples across 10 RNA families, with sequence lengths ranging from 28 to 2968 nucleotides.
This dataset is frequently used to evaluate RNA secondary structure prediction methods, including those that handle both pseudoknotted and non-pseudoknotted structures.
It is considered complementary to the RNAStrAlign… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/archiveii.512.bprna-spot-0
bpRNA-spot
bpRNA-spot is a collection of the datasets used by SPOT-RNA for RNA secondary structure prediction.
The dataset is released as a composite repository, bpRNA-spot, and three numbered component repositories:
bpRNA-spot-0: the initial bpRNA split, TR0, VL0, and TS0.
bpRNA-spot-1: the PDB transfer-learning split, TR1, VL1, and TS1.
bpRNA-spot-2: the NMR-only evaluation split, TS2.
bpRNA-spot concatenates the components in order:
train: TR0 + TR1
validation: VL0 + VL1
test:… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/bprna-spot-0.ryos-1
RYOS
RYOS is a database of RNA backbone stability in aqueous solution.
RYOS focuses on exploring the stability of mRNA molecules for vaccine applications.
This dataset is part of a broader effort to address one of the key challenges of mRNA vaccines: degradation during shipping and storage.
Statement
Deep learning models for predicting RNA degradation via dual crowdsourcing is published in Nature Machine Intelligence, which is a Closed Access / Author-Fee journal.… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/ryos-1.rnacentral.4096
RNAcentral
RNAcentral is a free, public resource that offers integrated access to a comprehensive and up-to-date set of non-coding RNA sequences provided by a collaborating group of Expert Databases representing a broad range of organisms and RNA types.
The development of RNAcentral is coordinated by European Bioinformatics Institute and is supported by Wellcome. Initial funding was provided by BBSRC.
Disclaimer
This is an UNOFFICIAL release of the RNAcentral by The… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/rnacentral.4096.chanrg
Comprehensive Hierarchical Annotation of Non-coding RNA Groups (CHANRG)
CHANRG is a database of non-coding RNA families and secondary structures.
Example Entry
id
sequence
secondary_structure
structural_annotation
functional_annotation
family
clan
architecture
super_family
split
AAAA02037454.1_2001-2135
GGATGCGATCATACCAGCACTAAAGCACCGGA...
(((((((((....((.(((((...((..((((...
SSSSSSSSSMMMMSSISSSSSIIISSIISSSS...
NNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNN...
RF00001… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/chanrg.rnacentral.1024
RNAcentral
RNAcentral is a free, public resource that offers integrated access to a comprehensive and up-to-date set of non-coding RNA sequences provided by a collaborating group of Expert Databases representing a broad range of organisms and RNA types.
The development of RNAcentral is coordinated by European Bioinformatics Institute and is supported by Wellcome. Initial funding was provided by BBSRC.
Disclaimer
This is an UNOFFICIAL release of the RNAcentral by The… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/rnacentral.1024.rivas-b
RIVAS
The RIVAS dataset is a curated collection of RNA sequences and their secondary structures, designed for training and evaluating RNA secondary structure prediction methods.
The dataset combines sequences from published studies and databases like Rfam, covering diverse RNA families such as tRNA, SRP RNA, and ribozymes.
The secondary structure data is obtained from experimentally verified structures and consensus structures from Rfam alignments, ensuring high-quality annotations… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/rivas-b.rnacentral.8192
RNAcentral
RNAcentral is a free, public resource that offers integrated access to a comprehensive and up-to-date set of non-coding RNA sequences provided by a collaborating group of Expert Databases representing a broad range of organisms and RNA types.
The development of RNAcentral is coordinated by European Bioinformatics Institute and is supported by Wellcome. Initial funding was provided by BBSRC.
Disclaimer
This is an UNOFFICIAL release of the RNAcentral by The… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/rnacentral.8192.archiveii.1024
ArchiveII
ArchiveII is a dataset of RNA sequences and their secondary structures, widely used in RNA secondary structure prediction benchmarks.
ArchiveII contains 2975 RNA samples across 10 RNA families, with sequence lengths ranging from 28 to 2968 nucleotides.
This dataset is frequently used to evaluate RNA secondary structure prediction methods, including those that handle both pseudoknotted and non-pseudoknotted structures.
It is considered complementary to the RNAStrAlign… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/archiveii.1024.bprna-spot-2
bpRNA-spot
bpRNA-spot is a collection of the datasets used by SPOT-RNA for RNA secondary structure prediction.
The dataset is released as a composite repository, bpRNA-spot, and three numbered component repositories:
bpRNA-spot-0: the initial bpRNA split, TR0, VL0, and TS0.
bpRNA-spot-1: the PDB transfer-learning split, TR1, VL1, and TS1.
bpRNA-spot-2: the NMR-only evaluation split, TS2.
bpRNA-spot concatenates the components in order:
train: TR0 + TR1
validation: VL0 + VL1
test:… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/bprna-spot-2.rnacentral.512
RNAcentral
RNAcentral is a free, public resource that offers integrated access to a comprehensive and up-to-date set of non-coding RNA sequences provided by a collaborating group of Expert Databases representing a broad range of organisms and RNA types.
The development of RNAcentral is coordinated by European Bioinformatics Institute and is supported by Wellcome. Initial funding was provided by BBSRC.
Disclaimer
This is an UNOFFICIAL release of the RNAcentral by The… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/rnacentral.512.rivas-a
RIVAS
The RIVAS dataset is a curated collection of RNA sequences and their secondary structures, designed for training and evaluating RNA secondary structure prediction methods.
The dataset combines sequences from published studies and databases like Rfam, covering diverse RNA families such as tRNA, SRP RNA, and ribozymes.
The secondary structure data is obtained from experimentally verified structures and consensus structures from Rfam alignments, ensuring high-quality annotations… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/rivas-a.
