RNA sequence
operon-identification-long-read-rna-sequencing-protein-sequences
Dataset for operon identification from long-read RNA sequencing
A dataset of annotated operons across 5 distinct bacterial strains. The operons were annotated by running and analysing long-read RNA sequencing and identifying genes
located on the same transcripts.
The genome protein sequences have been extracted from GenBank. Each row contains whole bacterial genome represented by an ordered list
of protein sequences.
Usage
For a complete example on how to read and use… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/operon-identification-long-read-rna-sequencing-protein-sequences.rna-3d-sequences-0001This dataset is derived from the Stanford RNA 3D Folding Kaggle competition data, under CC BY 4.0 license. Attribution: Stanford University / Competition Organizers.
Paul_RNA_Sequence_Processed_DatasetPaul_RNA_Sequence_Unprocessed_Data
