multimolecule
archiveii
ArchiveII
ArchiveII is a dataset of RNA sequences and their secondary structures, widely used in RNA secondary structure prediction benchmarks.
ArchiveII contains 2975 RNA samples across 10 RNA families, with sequence lengths ranging from 28 to 2968 nucleotides.
This dataset is frequently used to evaluate RNA secondary structure prediction methods, including those that handle both pseudoknotted and non-pseudoknotted structures.
It is considered complementary to the RNAStrAlign… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/archiveii.rnacentral
RNAcentral
RNAcentral is a free, public resource that offers integrated access to a comprehensive and up-to-date set of non-coding RNA sequences provided by a collaborating group of Expert Databases representing a broad range of organisms and RNA types.
The development of RNAcentral is coordinated by European Bioinformatics Institute and is supported by Wellcome. Initial funding was provided by BBSRC.
Disclaimer
This is an UNOFFICIAL release of the RNAcentral by The… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/rnacentral.eternabench-switch
EternaBench-Switch
EternaBench-Switch is a synthetic RNA dataset consisting of 7,228 riboswitch constructs, designed to explore the structural behavior of RNA molecules that change conformation upon binding to ligands such as FMN, theophylline, or tryptophan.
These riboswitches exhibit different structural states in the presence or absence of their ligands, and the dataset includes detailed measurements of binding affinities (dissociation constants), activation ratios, and RNA… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/eternabench-switch.pdb-rna_secondary_structure
pdb-rna_secondary_structure
[!IMPORTANT]The pdb-rna_secondary_structure dataset is in beta test.
This dataset card may not accurately reflects the data content.
The data content and this dataset card may subject to change.
Please contact the MultiMolecule team on GitHub issues should you have any feedback.
[!CAUTION]
This dataset is converted from the dataset released by the authors of SPOT-RNA.
The MultiMolecule is aware of a potential issue in data quality.
We are working on… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/pdb-rna_secondary_structure.eternabench-cm
EternaBench-CM
EternaBench-CM is a synthetic RNA dataset comprising 12,711 RNA constructs that have been chemically mapped using SHAPE and MAP-seq methods.
These RNA sequences are probed to obtain experimental data on their nucleotide reactivity, which indicates whether specific regions of the RNA are flexible or structured.
The dataset provides high-resolution, large-scale data that can be used for studying RNA folding and stability.
Disclaimer
This is an UNOFFICIAL… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/eternabench-cm.bprna
bpRNA-1m
bpRNA-1m is a database of single molecule secondary structures annotated using bpRNA.
Disclaimer
This is an UNOFFICIAL release of the bpRNA-1m by Center for Quantitative Life Sciences of the Oregon State University.
The team releasing bpRNA did not write this dataset card for this dataset so this dataset card has been written by the MultiMolecule team.
Example Entry
id
sequence
secondary_structure
structural_annotation
functional_annotation… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/bprna.
