datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rnafold
RNA 3D Structure Library and Targets
Predict the 3D structure of an RNA molecule from its sequence, as C1' atom coordinates per residue, given a library of thousands of experimentally solved RNA structures. train_sequences.csv and train_labels.csv give every training target's sequence and C1' coordinates, structures/ holds one coordinate-only mmCIF per training assembly, and test_sequences.csv lists the held-out targets, which were solved after every training structure was… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/rnafold.rnacentral-pretokenized-shardedRNAgym
🧬 RNAGym
Benchmark suite for RNA fitness & structure prediction
RNAGym brings together >1 M mutational fitness measurements and curated RNA structure datasets in a single place. This Hugging Face repo lets you:
Download fitness benchmark data
Download secondary structure benchmark data
Download tertiary structure benchmark data
Contents
Folder
Purpose
fitness_prediction/
All mutation-fitness assays
secondary_structure_prediction/
Secondary… See the full description on the dataset page: https://huggingface.co/datasets/Marks-lab/RNAgym.rna-downstream-tasks
GB.RNA Benchmark Datasets
mRNA related tasks
Translation efficiency prediction from Chu et al.(2024) [1]
3 cell lines: Muscle, pc3, HEK
input sequence: 5'UTR
10-fold cross-validation split
mRNA expression level prediction from Chu et al.(2024) [1]
3 cell lines: Muscle, pc3, HEK
input sequence: 5'UTR
10-fold cross-validation split
Mean ribosome load prediction from Sample et al. (2019) [2]
input sequence: 5'UTR
ouput: mean ribosome load
the original data… See the full description on the dataset page: https://huggingface.co/datasets/genbio-ai/rna-downstream-tasks.RNAmd-v1
RNAmd v1
RNAmd is a public dataset of standardized RNA molecular-dynamics simulations,
RNA dynamic labels, quality-control summaries and benchmark-ready metadata for
compact functional and structured RNAs.
This Hugging Face Dataset repository is the core archival mirror for RNAmd v1.
It mirrors the analysis-ready metadata, labels, QC summaries, protocol files,
website index and RNA-only structures. The canonical RNAmd release server
remains the primary access point for processed… See the full description on the dataset page: https://huggingface.co/datasets/LlewynLuo/RNAmd-v1.RNA_chemical_ribonanzaRNAcentral-Struct
RNAcentral-Struct
RNAcentral-Struct is a structure subset derived from the official RNAcentral PDB cross-reference table. It contains every unique PDB entry referenced by RNAcentral current_release/id_mapping/database_mappings/pdb.tsv, downloaded as mmCIF from official PDB mirrors.
This dataset is intended for RNA structure-conditioned modeling, inverse folding, and cross-reference analysis. RNAcentral itself is a sequence and annotation resource; the 3D coordinates here come… See the full description on the dataset page: https://huggingface.co/datasets/StarLiu714/RNAcentral-Struct.e_coli_rnasrnaglib
Dataset Card for Dataset Name
This dataset contains the tasks provided by rnaglib, a benchmarking suite for RNA structure-function modelling.
Dataset Details
Dataset Description
This dataset is a Python package wrapping RNA benchmark datasets and tasks. Data access, preprocessing, and task-specific pipelines are implemented in code and not fully expressed through this metadata schema. See documentation at https://github.com/cgoliver/rnaglib.
Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/luiswyss/rnaglib.rnacentral
RNAcentral
RNAcentral is a free, public resource that offers integrated access to a comprehensive and up-to-date set of non-coding RNA sequences provided by a collaborating group of Expert Databases representing a broad range of organisms and RNA types.
The development of RNAcentral is coordinated by European Bioinformatics Institute and is supported by Wellcome. Initial funding was provided by BBSRC.
Disclaimer
This is an UNOFFICIAL release of the RNAcentral by The… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/rnacentral.jump-cp-0016-labelfree-RNARNA_degradationrna-junctions-dbMIRROR_Pruned_TCGA_RNASeq_DataHere we provide pruned TCGA transcriptomics data from manuscript "MIRROR: Multi-Modal Pathological Self-Supervised Representation Learning via Modality Alignment and Retention". Code is available at GitHub.
The TCGA [1] transcriptomics data were collected from Xena [2] and preprocessed using the proposed novel pipeline in MIRROR [3].
For raw transcriptomics data, we first apply RFE [4] with 5-fold cross-validation for each cohort to identify the most performant support set for the subtyping… See the full description on the dataset page: https://huggingface.co/datasets/Franklin2001/MIRROR_Pruned_TCGA_RNASeq_Data.pdb-rna_secondary_structure
pdb-rna_secondary_structure
[!IMPORTANT]The pdb-rna_secondary_structure dataset is in beta test.
This dataset card may not accurately reflects the data content.
The data content and this dataset card may subject to change.
Please contact the MultiMolecule team on GitHub issues should you have any feedback.
[!CAUTION]
This dataset is converted from the dataset released by the authors of SPOT-RNA.
The MultiMolecule is aware of a potential issue in data quality.
We are working on… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/pdb-rna_secondary_structure.rnahl-saluki-human
Overview
mRNA half-life is a measure of the degradation rate of mRNA molecules. This experiment reports the time that the expression level of a transcript takes to decrease by half. The original data source aggregates 39 human and 27 mouse transcriptome wide datasets -- this dataset contains the human datasets. Several data preprocessing steps are taken, reported in the original paper. Half-life measures per gene averaged across collected datasets, and PCA is performed on the gene x… See the full description on the dataset page: https://huggingface.co/datasets/morrislab/rnahl-saluki-human.rna-sbdd-v2
RNA-SBDD v2
A frozen benchmark for RNA structure-based drug design: 8,006 RNA pocket-ligand
complexes derived from RCSB, with a sequence-identity-disjoint split, the
evaluation artifacts, and the trained checkpoints the benchmark's numbers come
from.
This repository exists because the cluster the work ran on was retired. It is a
complete handoff — dataset, artifacts, weights, and the tooling to bring all of
it up somewhere else.
Code: git@Ced3-han:Ced3-han/RNASBDD.git, branch… See the full description on the dataset page: https://huggingface.co/datasets/CedLJH/rna-sbdd-v2.rna-stability-siegel
Overview
This dataset contains 3' UTR fragment measurements from the fast-UTR
massively parallel reporter assay reported by Siegel et al. The source library
contains 41,255 sequences tested in Jurkat T cells and BEAS-2B airway
epithelial cells. Each configuration retains rows with a T4 stability target,
a T4 effect target, or a reference needed to pair a measured effect. The
library contains native human 3' UTR fragments, natural variants, and designed
mutations of regulatory… See the full description on the dataset page: https://huggingface.co/datasets/morrislab/rna-stability-siegel.rna_evaluation
RNA Evalaution Tasks
Obtained from BEACON: Benchmark for Comprehensive RNA Tasks and Language Models
Tasks
Mean Ribosome Loading
Modification
ncRNA Family
Secondary Structure
Splicing
Citation
If you find this repo useful for your research, please consider citing the paper
@misc{ren2024beacon,
title={BEACON: Benchmark for Comprehensive RNA Tasks and Language Models},
author={Yuchen Ren and Zhiyuan Chen and Lifeng Qiao and Hongtai Jing and Yuchen… See the full description on the dataset page: https://huggingface.co/datasets/wanglab/rna_evaluation.rna-secondary-structure-predictionrnastralign
RNAStrAlign
RNAStrAlign is a comprehensive dataset of RNA sequences and their secondary structures.
RNAStrAlign aggregates data from multiple established RNA structure repositories, covering diverse RNA families such as 5S ribosomal RNA, tRNA, and group I introns.
It is considered complementary to the ArchiveII dataset.
Disclaimer
This is an UNOFFICIAL release of the RNAStrAlign by Zhen Tan, et al.
The team releasing RNAStrAlign did not write this dataset card for… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/rnastralign.RNAstralign
Data types
sequence: 27125 datapoints
structure: 27125 datapoints
family: 27125 datapoints
Conversion report
Over a total of 37149 datapoints, there are:
OUTPUT
ALL: 27125 valid datapoints
INCLUDED: 104 duplicate sequences with different structure / dms / shape
MODIFIED
0 multiple sequences with the same reference (renamed reference)
FILTERED OUT
3949 invalid datapoints (ex: sequence with non-regular characters)
9 datapoints with… See the full description on the dataset page: https://huggingface.co/datasets/rouskinlab/RNAstralign.recount3-RNA-seqrna-loc-fazal
Overview
mRNA localization annotates the subcellular compartments that mRNA are found in. This task is a multilabel classification -- mRNA can be found in more than one compartment. This dataset was computed from experimental APEX RNA seq data collected by Fazal et al. 2019.
This dataset is redistributed as part of mRNABench: https://github.com/morrislab/mRNABench
Data Format
Description of data columns:
target: Multihot labelling of cellular components that an mRNA… See the full description on the dataset page: https://huggingface.co/datasets/morrislab/rna-loc-fazal.operon-identification-long-read-rna-sequencing-protein-sequences
Dataset for operon identification from long-read RNA sequencing
A dataset of annotated operons across 5 distinct bacterial strains. The operons were annotated by running and analysing long-read RNA sequencing and identifying genes
located on the same transcripts.
The genome protein sequences have been extracted from GenBank. Each row contains whole bacterial genome represented by an ordered list
of protein sequences.
Usage
For a complete example on how to read and use… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/operon-identification-long-read-rna-sequencing-protein-sequences.rnahl-saluki-mouse
Overview
mRNA half-life is a measure of the degradation rate of mRNA molecules. This experiment reports the time that the expression level of a transcript takes to decrease by half. The original data source aggregates 39 human and 27 mouse transcriptome wide datasets -- this dataset contains the mouse datasets. Several data preprocessing steps are taken, reported in the original paper. Half-life measures per gene averaged across collected datasets, and PCA is performed on the gene x… See the full description on the dataset page: https://huggingface.co/datasets/morrislab/rnahl-saluki-mouse.gtex-single-cell-rnaseq
GTEx Single-Cell RNA-seq Dataset
This repository provides tools to create a Hugging Face dataset from GTEx single-nucleus RNA-seq data, transforming the hierarchical H5AD format into a flat, ML-ready structure.
Overview
Data Source
The data comes from GTEx's snRNA-seq atlas:
Source: GTEx Portal
Publication: Eraslan et al., Science 2022 - "Single-nucleus cross-tissue molecular reference maps toward understanding disease gene function"
Content: 209… See the full description on the dataset page: https://huggingface.co/datasets/ai-department-lpnu/gtex-single-cell-rnaseq.rna-lifecycle-ietswaart
RNA Lifecycle Prediction
Overview
An mRNA molecule’s path from transcription to translation involves traversing multiple cellular compartments. This dataset, processed by mRNABench from experimental data by Ietswaart et al. (2024), provides an isoform-resolved map of this process using direct RNA sequencing.
The underlying assay captures RNA flow dynamics by measuring the rates at which transcripts are released from chromatin, exported from the nucleus, and loaded onto… See the full description on the dataset page: https://huggingface.co/datasets/morrislab/rna-lifecycle-ietswaart.Face-Edit-Bench
Face-Edit-Bench
Face-Edit-Bench is a large-scale benchmark for evaluating text-guided facial image editing models. It is introduced in:
A Large-scale Evaluation of Text-guided Models for Facial Editing
ACMMM 2026
GitHub: https://github.com/rahul1801/Face-Edit-Bench
The benchmark evaluates models on three axes — identity preservation, edit fidelity, and unintended attribute changes — with fine-grained breakdowns by demographic subgroup.
Models Evaluated… See the full description on the dataset page: https://huggingface.co/datasets/rnair21/Face-Edit-Bench.1M_RNAseq_SRA_samples_annotations
1M RNAseq SRA samples annotations
800k_completed_metadata.csv: main metadata table for approximately 800k annotated RNA-seq/SRA samples.
800k_rnaseq_metappuccino_labels.db: SQLite database containing Metappuccino labels and associated structured metadata.
manual_dics/: folder containing manual normalization and cross-column remapping dictionaries.
FINAL_pass_unique_biopsy_site_values_mapped_to_uberon_with_cosine.tsv: maps raw biopsy_site values to normalized UBERON anatomical… See the full description on the dataset page: https://huggingface.co/datasets/chumphati/1M_RNAseq_SRA_samples_annotations.
